What this analysis does
Choose failure stages and delivery conditions to generate deterministic test responses for retries, temporary failures, hard bounces, throttling and webhook disorder.
Most applications are tested against email delivery that succeeds. The failure paths — deferrals, hard rejections, timeouts, duplicate webhooks and out-of-order events — are where transactional systems actually break, usually in production and usually at volume.
Temporary failures are the ones that break applications
A permanent rejection is straightforward: the address is bad, the message stops, the record is updated. A temporary deferral is harder, because the correct behaviour is to retry with backoff while treating the message as still in flight. Applications that mark a deferral as failed generate duplicate sends when the original eventually delivers.
The opposite error is equally common. Treating a permanent rejection as retryable produces queues that grind against an address that will never accept mail, consuming reputation with every attempt. Distinguishing the two classes correctly is the single most valuable behaviour to test.
- 4xx responses require retry with backoff, not failure handling.
- 5xx responses require suppression, not retry.
- Misclassifying either direction produces duplicates or reputation damage.
Webhooks arrive late, twice and out of order
Delivery notifications are not a reliable ordered stream. A bounce webhook can arrive before the accepted event for the same message, the same event can be delivered more than once, and a provider outage can produce a burst of backdated events hours later. Handlers written against the happy path fail all three ways.
The defensive design is idempotent handling keyed on the provider message identifier, with state transitions that ignore events older than the current state. Testing that requires deliberately replaying duplicates and reordering the sequence, which is exactly what generated scenarios provide.
Test the user-visible consequence, not just the log
The failure that matters is rarely the SMTP error itself. It is the password reset that silently never arrives, the order confirmation that duplicates, or the interface that reports success while the message sits in a dead-letter queue. Those outcomes are what a resilience test should assert against.
Include the recovery path in scope. Systems frequently handle the failure correctly and then fail to recover when the provider returns, leaving a backlog that never drains or a suppression list that keeps a valid address blocked long after the underlying condition cleared.
Example: a deferral treated as a permanent failure
Evidence supplied
A scenario returning a 421 temporary response during a transactional send.
How to read the result
An application that marks the message failed and re-queues from scratch will deliver twice once the deferral clears. The scenario makes that behaviour observable in testing rather than in production.
Known limits
- Generated scenarios model provider behaviour patterns and are not a live connection to any provider.
- Real provider responses vary and change over time.
- Passing these scenarios demonstrates handling of the cases modelled, not resilience in general.
Common questions
Does this send real email?
No. It generates failure scenarios and payloads for testing your own application. No mail is transmitted.
Which failure should be tested first?
The temporary deferral, because misclassifying it as permanent is the most common cause of duplicate transactional mail.
How do I make webhook handling safe?
Key on the provider message identifier, make handlers idempotent, and ignore events that would move state backwards.