Offensive interception field guide¶
This guide answers the practical assessment question: a message is paused. What should I change, why should I change it, and what would count as a result?
Use these techniques only on systems and data inside the assessment scope. Use controlled identities, endpoints, tasks, and effect ledgers. Credential-bearing fields are visible in a raw capture, so minimize retention and export a redacted record unless the credential itself is the authorized test variable.
The basic offensive loop¶
Do not begin with a large payload list. Begin with a trust hypothesis:
- Identify the boundary the message is crossing.
- Forward the original once and record normal behavior.
- Change one authoritative field or delivery property.
- Predict what the receiving component would do if it trusts the change.
- Forward the modified message.
- Inspect the receiver, task state, tool event, callback, or effect ledger.
- Reset and run a close control.
The useful unit of work is:
one boundary + one controlled change + one predicted observation + one close control
The Cockpit's Attack ideas section derives this unit from the selected message. It shows the exact JSON Pointer or delivery action, the trust hypothesis, what to watch for, and the close control. Prepare edit changes the editor only; it never forwards automatically.
How to inject¶
Body fields¶
Use Body for decoded A2A or MCP content. Expand the structured tree and replace, add, delete, rename, duplicate, or reorder a field. Changed JSON Pointers appear beneath the editor.
Good body targets include:
- task, context, tenant, actor, and principal identifiers;
- delegated instructions and approval parameters;
- tool names, tool arguments, and returned tool data;
- resources, artifacts, revisions, and task states;
- Agent Card skills, capabilities, interfaces, and extensions.
Envelope fields¶
Use Envelope for non-secret transport metadata such as HTTP method/path, SSE event ID/type, WebSocket opcode, and decoded gRPC metadata. Credential values are readable and editable -- see "Credentials" below.
Change envelope data when the hypothesis concerns routing, correlation, content type, replay handling, or stream semantics, not application data.
Delivery behavior¶
Some failures require no content change:
- Drop tests missing-message and fallback behavior.
- Duplicate tests idempotency and replay handling.
- Stream batch controls test explicit release order.
- Replay tests whether an old task, artifact, callback, or final event remains authoritative.
Raw editing¶
Use Raw only for framing, encoding, duplicate-key, or decoder-boundary tests that structured editing cannot express. Invalid or oversized content stays pending; it never replaces the immutable original.
What to attack in A2A¶
| Boundary | Controlled injection | What to look for | Close control |
|---|---|---|---|
| Delegated intent | Change the first text part from a normal action to a controlled alternative | Different specialist decision, tool call, artifact, or ledger entry | Rephrase the original intent without changing authority |
| Approval parameters | Move an amount, limit, account, or destination across a known test boundary | Receiver or policy tool consumes the changed value | Change to a nearby value that remains inside policy |
| Task binding | Replace taskId with another controlled task |
Read, update, completion, or artifact attached to the wrong task | Change unrelated metadata while retaining the task |
| Context binding | Replace contextId with another controlled context |
State or action crosses conversation or tenant context | Preserve context and change a descriptive field |
| Tenant binding | Replace tenant metadata with another synthetic tenant | Cross-tenant task visibility or controlled effect | Preserve tenant and change a non-authoritative label |
| Actor/principal | Replace a claimed actor while authenticated context is unchanged | Approval or action attributed to the wrong principal | Keep actor and alter descriptive metadata |
| Agent Card route | Replace an interface URL with a controlled receiver | Subsequent delegation reaches the replacement | Alter card description only |
| Agent Card capability | Toggle streaming, callback, or extension support | Caller selects a different path or unsafe fallback | Preserve capability and alter description |
| Agent Card skill | Replace a skill ID/name with a controlled shadow | Work routes based on mutable discovery data | Preserve skill identity and alter description |
| Artifact content | Replace text/data in an artifact update | Subscriber or downstream agent consumes it | Alter display name, not content |
| Artifact revision | Advance or replay a revision | Late data replaces an earlier trusted result | Replay same revision and content |
| Task state | Change a status transition to completed, failed, or another valid fixture state |
Consumer commits, retries, cancels, or advances incorrectly | Change only the status message |
| Callback destination | Replace callback URL with a controlled local receiver | Callback arrives at the wrong authorized destination | Keep URL and alter callback label |
| Lifecycle | Drop, duplicate, replay, or reorder an event | Stale state, duplicate effect, unsafe fallback, or illegal transition | Deliver the original sequence once |
Common A2A pointers vary by binding, but frequently include:
/params/message/contextId/params/message/taskId/params/message/metadata/tenantId/params/message/metadata/actor/params/message/parts/0/text/result/statusUpdate/status/state/result/artifactUpdate/artifact/parts/0/text/result/artifactUpdate/artifact/metadata/revision
Select the actual decoded path shown by AIT rather than assuming a field exists.
What to attack in MCP¶
| Boundary | Controlled injection | What to look for | Close control |
|---|---|---|---|
| Tool arguments | Change amount, target, path, account, query, or mode | Tool event and external effect use the changed argument | Nearby allowed argument |
| Tool identity | Replace params.name with a controlled shadow tool |
Different tool selected or substitution rejected | Original tool with harmless argument change |
| Tool schema | Change name, description, required fields, type, or enum in tools/list |
Client exposes, selects, or calls the altered contract | Change unused description only |
| Tool result | Change structured content or returned text | Agent decision or downstream delegation changes | Alter an unused result field |
| Resource identity | Replace params.uri with a controlled alternate resource |
Unauthorized resource returned or consumed | Original URI with unrelated metadata change |
| Resource content | Insert a controlled instruction/marker into returned resource text | Agent treats data as instruction or changes its tool choice | Benign marker of similar length |
| Prompt arguments | Change a template argument or inserted context | Generated prompt crosses an intended authority boundary | Semantically equivalent safe argument |
| Sampling restrictions | Change toolChoice, system prompt, or message content |
Previously unavailable controlled tool/action becomes selectable | Preserve restriction and vary prose |
| Elicitation URL | Replace URL-mode elicitation destination | Client presents or navigates to a controlled replacement | Preserve URL and alter display text |
| Session/correlation | Change JSON-RPC ID or MCP session metadata | Wrong response association, rejection, or state crossover | Preserve identifier and change a non-routing field |
| Lifecycle | Drop/duplicate initialize, request, notification, or response | Desynchronization, repeated effect, or unsafe retry | Forward one unchanged copy |
Frequent MCP pointers include:
/params/name/params/arguments/<field>/params/uri/result/tools/0/inputSchema/result/content/0/text/result/structuredContent/<field>/params/toolChoice/params/systemPrompt
Response substitution as a test hypothesis¶
respond can answer a paused request without forwarding it. That is an
instrument capability, not evidence that a particular receiver will accept the
substitute. Treat each use as a bounded hypothesis about one named
implementation and binding.
Run an unchanged baseline, the substituted reply, and a same-shape close control. Record whether the intended peer was contacted and whether the receiver produced an effect visible outside its own narration. A correlation identifier links a response to a request; it does not by itself authenticate who produced that response. Do not generalize from one accepted response to a protocol, SDK, vendor, or population without a separately reviewed experiment.
High-value first tests¶
Use this order when you know little about the system:
- Forward one request and response unchanged.
- Change a delegated numeric or identity parameter.
- Change one task/context/tenant binding.
- Change one MCP tool argument.
- Change the corresponding MCP result.
- Replace artifact content while keeping its identifiers.
- Replay an idempotent-looking request once.
- Drop one non-final stream event.
- Replay or duplicate a final stream event.
- Advance an artifact revision.
- Change a discovery-time skill, tool name, or capability.
- Replace a callback/resource URL with a controlled loopback endpoint.
Stop and investigate as soon as a controlled effect appears. Broad mutation is less useful than a small witness that explains which boundary failed.
What to look for¶
The edited message is only the input. Inspect downstream behavior:
Receiver behavior¶
- Was the message accepted or rejected?
- Did the receiver preserve the original task, context, tenant, and actor?
- Did it choose a different agent, tool, resource, or workflow branch?
- Did it detect an illegal state transition or stale revision?
Lifecycle behavior¶
- Did a duplicate produce one effect or two?
- Did a dropped event cause safe failure, retry, or unsafe fallback?
- Did replay survive cancellation or resubscription?
- Did a late artifact replace data already used for a decision?
Independent effects¶
Prefer one of these:
- controlled tool/action ledger entry;
- callback receipt;
- datastore or test-state change;
- task/artifact state observed from an independent component;
- controlled file or transaction record.
A target response may be valuable, but it does not prove an external effect by itself.
Designing controls¶
Use four comparisons when the consequence matters:
- Original: unchanged normal behavior.
- Attack: one authoritative change.
- Close control: similar shape and encoding without crossing the trust boundary.
- Miss: the same operator aimed at a selector/path that should not apply.
Reset mutable test state between comparisons. Keep topology, timing policy, and all unrelated fields identical. If the close control triggers the same effect, the test does not isolate the proposed weakness.
Interpreting the evidence¶
| Evidence observed | Defensible conclusion |
|---|---|
| Original and delivered differ | AIT altered the communication |
| Receiver returned a correlated response | The receiver processed the delivered request |
| Receiver selected a different path/tool | The mutation influenced receiver behavior |
| Controlled ledger/callback/state changed | The mutation produced an independently observed effect |
| Close control stayed inert | The authoritative change, not generic disturbance, explains the effect |
| Only a trace, marker echo, or model statement changed | Interesting observation; external impact is not proven |
Running a boundary as a primitive¶
Every boundary in the tables above ships as an executable attack primitive, so the unit of work does not have to be reassembled by hand each time. A primitive carries the selector, the injection, the trust hypothesis, what to watch for, the close control with its rationale, and an explicit limitation.
ait intercept attack --target MESSAGE_ID # what applies to this message
ait intercept attack mcp.tool-shadowing --target MESSAGE_ID
ait intercept attack mcp.tool-shadowing --target NEXT_MESSAGE_ID --arm control
The control arm is not optional. Without it a positive result cannot be separated from the effect of any disturbance, so the attack arm prints the control to run next and why that particular control isolates this boundary.
Primitives are data under ait/data/attack-primitives/. Add your own for a
target-specific boundary without modifying the tool; the same fields are
required, including the limitation, because a primitive that claims no limit is
one nobody can review.
Credentials¶
Credential-bearing headers -- Authorization, Cookie, X-Api-Key and
friends -- are captured verbatim and are fully editable. Forge one, downgrade
one, or strip one by omitting it from the edited envelope.
That makes credential relay, replay, token downgrade, and confused-deputy-
via-token testable, which they were not while the tool masked those values. It
also puts weight on the export contract: --redacted (the default) carries
allowlisted metadata only, and --raw is a complete capture labelled as such.
Hand a client the redacted one.
A delivery whose credential headers differ from what the sender supplied is
marked credential_edited, and chain 2.2 binds the envelope into the record
hash, so the transcript never presents a forged token as the sender's own.
Originating traffic¶
Waiting for an agent to send the message you want to test is not always practical.
ait intercept send --path /message --body-json @probe.json
ait intercept send --from MESSAGE_ID --set /params/message/metadata/amount=999
--from is the repeater. It takes a captured request, applies JSON-Pointer
edits, and delivers it again. Injected traffic goes through the same relay as
captured traffic, so break conditions, rewrite rules, and the transcript all
apply -- which also means injecting while a matching break condition is armed
will pause your own message. Turn interception off, or narrow the break
conditions, when you want to inject freely.
Producing a finding¶
The interpretation table above says what each observation licenses you to
claim. ait intercept evidence applies it to the arms you actually ran.
Two tiers the tool determines itself: whether a mutation was delivered, and whether a correlated response came back. Two it will not infer:
behavior_changedneeds a baseline to compare receiver behaviour against.external_effect_observedneeds an out-of-band oracle -- a controlled tool or action ledger, a callback receipt, or datastore state read independently of the target. A target response is not an oracle.
Pass --behavior-changed or --external-effect only for something you
observed. Left unset, those tiers are reported as unproven along with what would
be required, and tiers cannot be skipped: claiming an external effect over a
delivery the receiver never processed does not promote the verdict.
The finding also records attribution, by comparing what the receiver did on
each arm. With no control arm the status is no_control; if the control
produced the same observable effect it is not_isolated; if the arms cannot be
compared -- no attack arm, or no captured behaviour for one of them -- it is
undetermined, because a missing measurement is not evidence of sameness. Only
a measured difference between the arms yields attributable.
Both arms must be in the same session. ait lab reset starts a new one.
Turning a manual edit into a reusable rule¶
After a successful modified delivery, select Apply to future. Review the derived selector and patch before enabling it. Narrow the selector by protocol, direction, operation, route, or destination so the rule cannot alter unrelated traffic.
Start rules disabled on a real assessment, observe one matching original, then enable deliberately. The hit count confirms matching; History confirms what was actually delivered.
When a rule is armed, a paused message shows the rule's output, not what the
sender sent. The sender's version stays available as wire_original, and
Sender's version (forward_wire) delivers it, bypassing the rule for that
one message.
Suggested assessment sequence¶
- Reproduce the intended behavior in the local approval lab.
- Use Use this setup with my target to copy transport and break behavior.
- Forward originals until the important agent and tool boundaries are clear.
- Test identity and authority fields before experimenting with framing.
- Test returned artifacts/tool results after requests.
- Test duplicate, drop, replay, and ordering behavior.
- Save only the smallest useful edits as live rules.
- Export the session with original/delivered values and independent evidence.
- Retest after remediation using the same controlled input and observation.
See Interception screen guide for every control and Hands-on interception lab for the deterministic practice environment. Use the direct interception assessment scenarios for twelve repeatable exercises and the operator command cookbook for every terminal action.