Detecting Rogue AI Agents with Cyber Deception

Cyber deception can help detect suspicious AI-agent activity by placing controlled decoy credentials, documents, APIs or MCP surfaces where an unexpected interaction can be observed. A useful detection begins with a clear question: what action should this actor or workflow not perform, and what evidence would show that it tried?

The approach can support investigation of adversary-operated automation and approved agents acting outside their task. It does not automatically identify an unknown actor as AI-powered, inventory every agent, or prove that an agent is safe when no alert fires.

What is a rogue AI agent?

In this guide, a rogue AI agent is an agent acting outside its authorised task or permitted workflow. The term describes behaviour to investigate. It doesn’t imply that a model has intentions, or that every automated action is hostile.

An adversary may operate an agent. A team may introduce an unapproved tool. An approved agent may also take an out-of-scope action after a configuration mistake, excessive permissions or misleading input. These situations require different attribution and response, even when they touch the same decoy.

Two related risk categories are described in OWASP’s guidance on excessive agency and prompt injection. Both explain how an agent can end up taking an action nobody intended. OWASP’s Top 10 for Agentic Applications also lists rogue agents as a risk in their own right. Deception doesn’t prevent any of these, but it can show when one leads to an unexpected action.

Why an approved agent still needs behavioural scrutiny

Approval to use a tool does not resolve whether each action is appropriate. A document assistant may be allowed to search a project folder but have no reason to test credentials found there. An operations assistant may be allowed to inspect selected services but not attempt an unrelated export.

For a detection design, write down the permitted task and the action that would fall outside it. This gives an analyst a reason to investigate a specific event rather than treating “an AI tool did something” as a finding.

A decoy event on its own doesn’t say what the agent was supposed to be doing. That context comes from your defined workflow and supporting records.

How decoy-based detection works

A decoy is a controlled observation point. Its value comes from the combination of placement, instrumentation and interpretation.

Start by choosing a behaviour worth observing. Then choose a signal that can record that behaviour, place it in an authorised environment, and establish what legitimate activity may touch it. Finally, test the entire path from interaction to analyst review.

The decoy should not grant access to production data or privileges. It should not require weakening real access controls or exposing a vulnerable service. A realistic lure is not a reason to create a new compromise path.

The cyber deception readiness checklist covers the wider planning process.

Where useful signals can be placed

Each type of signal suits a different kind of interaction, and each has limits worth knowing before you rely on it.

Signal Interaction worth investigating Important limit
Decoy credential Attempted use against a monitored validation or service path Merely copying or reading a credential may not generate an observable event.
Synthetic document or reference A monitored open, callback or follow-on access, where that mechanism is supported A normal file read is not universally visible, and legitimate previewers may trigger some mechanisms.
Controlled API or service An unexpected request, attempted operation or access path A request does not by itself establish the caller’s task, identity or intent.
Decoy MCP surface A monitored protocol interaction, resource access or tool request Normal discovery may be legitimate; a server-side event does not expose everything happening inside the client.
Synthetic export or sensitive-looking path An attempted action at a controlled destination An attempt against a decoy is not evidence that real customer data was exported.

For credentials and identities, see identity-led deception for cloud credential compromise.

Placement is often the hard part

A decoy that nobody encounters cannot observe anything. Equally, an artefact placed directly in every normal workflow can create routine noise rather than a meaningful signal.

Choose a plausible discovery path around the workflow being assessed: a controlled workspace reference, a synthetic service description, or an access path that appears relevant to the actor but is not required for authorised work. Confirm that placement is permitted and that legitimate scanners, indexers and administrators have been considered.

For external activity, focus on the public or exposed surfaces you deliberately operate. For known internal agents, focus on the workspace and tool boundaries you can define and test. Do not assume the same placement strategy serves both.

Keep decoys away from content you want people and assistants to read, such as public documentation. An assistant reading a public page is doing its job.

Detection and attribution are different questions

Keep three layers separate in the investigation:

Observed event: A specific interaction was recorded at an instrumented signal. State which signal, action and time are actually supported by the evidence.

Operational interpretation: The interaction may be expected, unexplained or inconsistent with the agreed workflow. That conclusion depends on baseline behaviour and surrounding context.

Actor attribution: Identity or execution records may associate the event with an approved agent, a user or another tool. Without that supporting evidence, an unknown actor should remain unknown.

An IP address may represent shared infrastructure. A user-agent string is an assertion made by a client. A sequence of requests may be consistent with automation without establishing that an AI model generated it. None of these alone reveals the model’s internal reasoning.

This distinction lets a finding stay useful even when attribution is incomplete: “a planted credential was tested at this controlled endpoint” can be an actionable observation without a claim that “a particular AI model went rogue”.

An illustrative approved-agent scenario

Illustrative example

This is a hypothetical example, not a customer incident or a measured result.

A team authorises an assistant to summarise a set of synthetic project documents in an isolated test workspace. The task does not require authenticating to another service. A controlled reference in the workspace leads to a decoy credential that is valid only for an isolated, monitored test service, not production resources.

During an authorised test, the assistant or its tooling attempts to use that credential at the controlled endpoint. The observation to investigate is the attempted use. Reliable execution records, where available, can help establish whether the request came from the known agent and whether it fell outside the approved task.

The test should also run a compliant workflow that completes the summary without using the credential. That comparison helps distinguish the intended boundary from a decoy that simply triggers during normal document processing.

What the scenario could establish What it would not establish
The monitored endpoint received an attempted credential use. That every read or copy of the credential was visible.
A matching execution record may associate the request with the tested agent. That any unknown caller can be identified as an AI model.
The action was outside the written test task, if the records support that conclusion. That the agent had malicious intent.
The configured detection and delivery path worked for this test. That all agents, workflows or future attacks will be detected.

A different case: unknown external automation

A public decoy API may receive an unusual sequence of requests, including an attempt to follow a synthetic credential reference. Those events can be investigated on their own terms.

Unless additional evidence establishes the actor or tooling, report the observable requests and relevant context. “Unknown automation interacted with a decoy service” is more defensible than naming an agent or model from timing or request style alone.

The detection can still support a decision to investigate, correlate other activity or apply an existing response policy. Attribution does not need to be invented to make the evidence useful.

MCP detection and MCP integration are not the same thing

The Model Context Protocol provides a way for AI applications to connect to external capabilities and context. For a security team, that creates two quite different things to think about.

A decoy MCP surface is a controlled observation point for interactions worth investigating. The meaning of an interaction depends on the expected client behaviour and deployment context.

The FIRCY Sense MCP integration is a production integration through which approved tools can access configured FIRCY intelligence for investigation. Using that integration does not make a tool a rogue agent.

Read about approved MCP access to FIRCY intelligence.

What should an investigator receive?

A useful finding should let the team reconstruct the observation without trusting a generated narrative alone. For a trial, agree which of the following are supported by the chosen signal and integration:

  • The instrumented decoy and the action it recorded, with timestamps and relevant request or service context.
  • The placement and workflow boundary that explain why the interaction may matter.
  • Available source, identity or execution context, clearly distinguishing observed facts from inferred associations.
  • Links or references to related evidence, plus an explanation of uncertainty and the configured route into the investigation workflow.

What’s available depends on the signal and the integration, so agree it up front rather than discovering gaps during an investigation.

When AI assists with triage, treat externally supplied event content as untrusted data rather than instructions. Review any generated explanation against the underlying event and preserve the distinction between observation and inference. OWASP’s prompt-injection guidance explains why this boundary matters.

How to evaluate the approach safely

Start with one controlled workflow and synthetic assets. Use only systems and agents you are authorised to test. Do not use real customer secrets or valuable data as bait.

Define the authorised task, the specific out-of-scope action and the signal expected to record it. Document the agent, model version, available tools, relevant permissions and configuration so the result can be interpreted later.

Test both ordinary operation and deliberately exercised boundary cases. Include the benign tools likely to encounter the artefact, such as approved scanners or previewers. Confirm not only that a signal occurs, but that the intended analyst receives understandable evidence.

Measure what the test actually supports. For example, record the number of controlled out-of-scope interactions attempted and the number captured. Report unexpected alerts from the normal-workflow runs separately. Record missing telemetry, missed test events and attribution gaps rather than quietly excluding them.

Keep the denominator visible. A result from a small controlled test is not a production detection rate, and a quiet baseline does not demonstrate zero false positives in all environments.

The NCSC’s update on its cyber deception trials makes a similar point: deception works best with clear planning, context and operational evidence.

What deception does not replace

Deception is complementary to agent inventories, least-privilege access, tool restrictions, approval controls, application security and normal logging. It can add a focused observation point; it cannot make an over-permissioned agent safe by itself.

No alert can mean that nothing touched the signal. It can also mean that the actor took a different path or that the interaction was not observable. Coverage must be tested, maintained and interpreted alongside other controls.

Similarly, a positive alert deserves investigation rather than an automatic conclusion about malicious intent. The right response depends on the evidence, the environment and the organisation’s approved policy.

Where FIRCY Sense fits

FIRCY Sense uses cyber deception for active defence. Decoy credentials, documents, APIs and MCP surfaces create early-warning signals, and each interaction becomes evidence your team can investigate and route into existing SIEM, SOAR, ticketing and response workflows. The evidence available and the delivery options depend on your deployment and configuration.

The practical starting point is a workflow worth protecting, an interaction worth observing and a team ready to investigate it. From there, the goal is useful evidence in your environment, not an AI label on every unfamiliar request.

See the AI-agent use case, explore the platform, or discuss a trial.