Adversarial agent testing

Agent Mail Red-Team Sandbox

Run an email agent permission model against prompt-injection, external exfiltration, attachment release, mass-send and delegated-tool abuse scenarios.

email agent red teamprompt injection email agentgmail mcp red teamai mailbox sandbox
Red-team a mailbox agent before production accessThis sandbox models permissions and control boundaries only. It does not connect to a mailbox or execute any agent action.

Mailbox capabilities

Control boundaries

What this analysis does

Configure the actions an email agent can take and the controls surrounding those actions. The sandbox evaluates deterministic adversarial scenarios before any real mailbox is connected, separates exposed paths from controlled paths, and produces concrete boundaries for human approval, recipient restriction, untrusted-content handling and action logging.

An email agent should be tested against adversarial input before it touches a production mailbox. The relevant scenarios are known: instructions embedded in message content, attempts to route data outward, attachment release, mass sending and abuse of whatever tools the agent can reach.

Reviewed against current agent security practice on 29 August 2026

The message body is the attack surface

Anyone can send mail to the mailbox, so message content is untrusted input by definition. An agent that reads messages and acts on them is executing instructions supplied by whoever chose to email it, and no system prompt reliably prevents that.

Injected instructions hide in ordinary places: quoted text, signature blocks, HTML that is invisible when rendered, and attachment content. A test suite should include each of those positions rather than only obvious plain-text commands.

  • Any sender can place content in the mailbox.
  • Injection hides in quoted text, signatures and invisible markup.
  • Instruction-based defences are advisory rather than enforcing.

Exfiltration needs no send permission

A read-only agent can still leak. Any outbound channel serves — a URL fetched, a calendar invite, a document shared, a webhook called — and content read from the mailbox can be encoded into a request the agent is permitted to make.

Test whether the agent can be induced to include mailbox content in any outbound action, not only in an email. Teams that assess only the send path routinely miss the tool paths that provide equivalent capability.

Test the limits, then the recovery

Rate and volume limits should be verified adversarially. An agent induced to send to every contact in the mailbox is a serious incident, and the control that prevents it must be enforced by the system rather than requested in a prompt.

Recovery matters as much as prevention. Determine in advance how the agent is stopped mid-run, whether its actions are logged in enough detail to reconstruct what happened, and whether sent mail can be identified and recalled. An agent that cannot be audited cannot be safely deployed.

Example: injection in an invisible element

Evidence supplied
A scenario placing instructions in HTML that is not visible when the message is rendered.

How to read the result
The sandbox reports whether the agent acted on content the user could not see. An agent that follows those instructions will do so for any sender, since nothing distinguishes the test message from an attacker’s.

Known limits

  • Scenarios model known attack classes and cannot enumerate every technique.
  • Testing is performed against your described model rather than a live mailbox.
  • Passing these scenarios demonstrates resistance to the cases tested, not general safety.

Common questions

Does this connect to a real mailbox?

No. Scenarios are generated against the permission model you describe. Nothing is sent and no mailbox is accessed.

Which scenario should be tested first?

Prompt injection combined with an outbound channel, because it applies to read-only agents and is the most commonly overlooked path.

How often should agents be retested?

After any change to permissions, model or tool access. Each of those changes the reachable action set.

Primary references