What this analysis does
Select the mailbox and adjacent capabilities an AI agent would receive, then add the approval, destination, rate and audit controls that constrain those permissions. Mailybox builds a deterministic abuse-path model so teams can see which combinations create external exfiltration, destructive-action or high-fanout risk without connecting a live mailbox.
Granting an AI agent access to a mailbox is a security decision with a blast radius that is rarely calculated beforehand. Read access to a mailbox exposes everything in it, and send access lets the agent act with the account holder’s identity. Modelling those permissions makes the consequence explicit before access is granted.
Read access is broader than it sounds
A mailbox is an archive of credentials, contracts, personal data and password reset links. Read access grants all of it, including historical content the agent has no operational need for, and including material belonging to counterparties who never consented to that processing.
Scope granularity is the control that matters. An agent that only needs recent unread messages should not receive full archive access, and where the API supports label or folder restriction, using it materially reduces exposure.
- Read access covers the entire archive unless explicitly scoped.
- Mailboxes contain credentials and reset links, not only correspondence.
- Label or folder restriction is the most effective available limit.
Send access transfers identity
An agent that can send mail sends as the account holder. Recipients see a trusted internal sender, and the message carries the organization’s authentication. There is no technical marker distinguishing agent-generated mail from human-generated mail unless one is deliberately added.
The combination that creates the largest risk is read plus send. It enables an agent to read a thread and reply within it, which is precisely the capability that makes conversation hijacking effective when the agent is manipulated.
Prompt injection makes mailbox content untrusted input
Any message in the mailbox is attacker-controllable, because anyone can send mail to it. An agent that reads messages and acts on their contents is processing untrusted instructions, and an attacker who can email the mailbox can attempt to steer it directly.
The mitigation is architectural rather than instructional. Constrain what the agent can do irrespective of what it is told: no sending to external recipients without approval, no attachment release, no acting on instructions found in message content, and hard limits on volume.
Example: an assistant with read and send scope
Evidence supplied
A drafting assistant granted full mailbox read access and unrestricted send permission.
How to read the result
The model reports the blast radius as the entire archive plus the ability to send as the account holder, and identifies external send without approval as the control whose absence most increases exposure to prompt injection.
Known limits
- The model reflects the permissions you describe rather than inspecting a live grant.
- Provider scope definitions change and differ between platforms.
- A modelled permission set does not evaluate the agent implementation itself.
Common questions
Is read-only access safe?
Safer, not safe. Read access still exposes the full archive, and an agent with any outbound channel can exfiltrate what it reads.
What is the most important single control?
Requiring human approval before sending to external recipients. It contains the majority of realistic abuse paths.
Can prompt injection be prevented by instructions?
No. Instructions are advisory. Effective mitigation constrains capability so the agent cannot perform the harmful action regardless of what it is told.