Skip to content

Offensive interception field guide

This guide answers the practical assessment question: a message is paused. What should I change, why should I change it, and what would count as a result?

Use these techniques only on systems and data inside the assessment scope. Use controlled identities, endpoints, tasks, and effect ledgers. Credential-bearing fields are visible in a raw capture, so minimize retention and export a redacted record unless the credential itself is the authorized test variable.

The basic offensive loop

Do not begin with a large payload list. Begin with a trust hypothesis:

  1. Identify the boundary the message is crossing.
  2. Forward the original once and record normal behavior.
  3. Change one authoritative field or delivery property.
  4. Predict what the receiving component would do if it trusts the change.
  5. Forward the modified message.
  6. Inspect the receiver, task state, tool event, callback, or effect ledger.
  7. Reset and run a close control.

The useful unit of work is:

one boundary + one controlled change + one predicted observation + one close control

The Cockpit's Attack ideas section derives this unit from the selected message. It shows the exact JSON Pointer or delivery action, the trust hypothesis, what to watch for, and the close control. Prepare edit changes the editor only; it never forwards automatically.

How to inject

Body fields

Use Body for decoded A2A or MCP content. Expand the structured tree and replace, add, delete, rename, duplicate, or reorder a field. Changed JSON Pointers appear beneath the editor.

Good body targets include:

  • task, context, tenant, actor, and principal identifiers;
  • delegated instructions and approval parameters;
  • tool names, tool arguments, and returned tool data;
  • resources, artifacts, revisions, and task states;
  • Agent Card skills, capabilities, interfaces, and extensions.

Envelope fields

Use Envelope for non-secret transport metadata such as HTTP method/path, SSE event ID/type, WebSocket opcode, and decoded gRPC metadata. Credential values are readable and editable -- see "Credentials" below.

Change envelope data when the hypothesis concerns routing, correlation, content type, replay handling, or stream semantics, not application data.

Delivery behavior

Some failures require no content change:

  • Drop tests missing-message and fallback behavior.
  • Duplicate tests idempotency and replay handling.
  • Stream batch controls test explicit release order.
  • Replay tests whether an old task, artifact, callback, or final event remains authoritative.

Raw editing

Use Raw only for framing, encoding, duplicate-key, or decoder-boundary tests that structured editing cannot express. Invalid or oversized content stays pending; it never replaces the immutable original.

What to attack in A2A

Boundary Controlled injection What to look for Close control
Delegated intent Change the first text part from a normal action to a controlled alternative Different specialist decision, tool call, artifact, or ledger entry Rephrase the original intent without changing authority
Approval parameters Move an amount, limit, account, or destination across a known test boundary Receiver or policy tool consumes the changed value Change to a nearby value that remains inside policy
Task binding Replace taskId with another controlled task Read, update, completion, or artifact attached to the wrong task Change unrelated metadata while retaining the task
Context binding Replace contextId with another controlled context State or action crosses conversation or tenant context Preserve context and change a descriptive field
Tenant binding Replace tenant metadata with another synthetic tenant Cross-tenant task visibility or controlled effect Preserve tenant and change a non-authoritative label
Actor/principal Replace a claimed actor while authenticated context is unchanged Approval or action attributed to the wrong principal Keep actor and alter descriptive metadata
Agent Card route Replace an interface URL with a controlled receiver Subsequent delegation reaches the replacement Alter card description only
Agent Card capability Toggle streaming, callback, or extension support Caller selects a different path or unsafe fallback Preserve capability and alter description
Agent Card skill Replace a skill ID/name with a controlled shadow Work routes based on mutable discovery data Preserve skill identity and alter description
Artifact content Replace text/data in an artifact update Subscriber or downstream agent consumes it Alter display name, not content
Artifact revision Advance or replay a revision Late data replaces an earlier trusted result Replay same revision and content
Task state Change a status transition to completed, failed, or another valid fixture state Consumer commits, retries, cancels, or advances incorrectly Change only the status message
Callback destination Replace callback URL with a controlled local receiver Callback arrives at the wrong authorized destination Keep URL and alter callback label
Lifecycle Drop, duplicate, replay, or reorder an event Stale state, duplicate effect, unsafe fallback, or illegal transition Deliver the original sequence once

Common A2A pointers vary by binding, but frequently include:

  • /params/message/contextId
  • /params/message/taskId
  • /params/message/metadata/tenantId
  • /params/message/metadata/actor
  • /params/message/parts/0/text
  • /result/statusUpdate/status/state
  • /result/artifactUpdate/artifact/parts/0/text
  • /result/artifactUpdate/artifact/metadata/revision

Select the actual decoded path shown by AIT rather than assuming a field exists.

What to attack in MCP

Boundary Controlled injection What to look for Close control
Tool arguments Change amount, target, path, account, query, or mode Tool event and external effect use the changed argument Nearby allowed argument
Tool identity Replace params.name with a controlled shadow tool Different tool selected or substitution rejected Original tool with harmless argument change
Tool schema Change name, description, required fields, type, or enum in tools/list Client exposes, selects, or calls the altered contract Change unused description only
Tool result Change structured content or returned text Agent decision or downstream delegation changes Alter an unused result field
Resource identity Replace params.uri with a controlled alternate resource Unauthorized resource returned or consumed Original URI with unrelated metadata change
Resource content Insert a controlled instruction/marker into returned resource text Agent treats data as instruction or changes its tool choice Benign marker of similar length
Prompt arguments Change a template argument or inserted context Generated prompt crosses an intended authority boundary Semantically equivalent safe argument
Sampling restrictions Change toolChoice, system prompt, or message content Previously unavailable controlled tool/action becomes selectable Preserve restriction and vary prose
Elicitation URL Replace URL-mode elicitation destination Client presents or navigates to a controlled replacement Preserve URL and alter display text
Session/correlation Change JSON-RPC ID or MCP session metadata Wrong response association, rejection, or state crossover Preserve identifier and change a non-routing field
Lifecycle Drop/duplicate initialize, request, notification, or response Desynchronization, repeated effect, or unsafe retry Forward one unchanged copy

Frequent MCP pointers include:

  • /params/name
  • /params/arguments/<field>
  • /params/uri
  • /result/tools/0/inputSchema
  • /result/content/0/text
  • /result/structuredContent/<field>
  • /params/toolChoice
  • /params/systemPrompt

Response substitution as a test hypothesis

respond can answer a paused request without forwarding it. That is an instrument capability, not evidence that a particular receiver will accept the substitute. Treat each use as a bounded hypothesis about one named implementation and binding.

Run an unchanged baseline, the substituted reply, and a same-shape close control. Record whether the intended peer was contacted and whether the receiver produced an effect visible outside its own narration. A correlation identifier links a response to a request; it does not by itself authenticate who produced that response. Do not generalize from one accepted response to a protocol, SDK, vendor, or population without a separately reviewed experiment.

High-value first tests

Use this order when you know little about the system:

  1. Forward one request and response unchanged.
  2. Change a delegated numeric or identity parameter.
  3. Change one task/context/tenant binding.
  4. Change one MCP tool argument.
  5. Change the corresponding MCP result.
  6. Replace artifact content while keeping its identifiers.
  7. Replay an idempotent-looking request once.
  8. Drop one non-final stream event.
  9. Replay or duplicate a final stream event.
  10. Advance an artifact revision.
  11. Change a discovery-time skill, tool name, or capability.
  12. Replace a callback/resource URL with a controlled loopback endpoint.

Stop and investigate as soon as a controlled effect appears. Broad mutation is less useful than a small witness that explains which boundary failed.

What to look for

The edited message is only the input. Inspect downstream behavior:

Receiver behavior

  • Was the message accepted or rejected?
  • Did the receiver preserve the original task, context, tenant, and actor?
  • Did it choose a different agent, tool, resource, or workflow branch?
  • Did it detect an illegal state transition or stale revision?

Lifecycle behavior

  • Did a duplicate produce one effect or two?
  • Did a dropped event cause safe failure, retry, or unsafe fallback?
  • Did replay survive cancellation or resubscription?
  • Did a late artifact replace data already used for a decision?

Independent effects

Prefer one of these:

  • controlled tool/action ledger entry;
  • callback receipt;
  • datastore or test-state change;
  • task/artifact state observed from an independent component;
  • controlled file or transaction record.

A target response may be valuable, but it does not prove an external effect by itself.

Designing controls

Use four comparisons when the consequence matters:

  • Original: unchanged normal behavior.
  • Attack: one authoritative change.
  • Close control: similar shape and encoding without crossing the trust boundary.
  • Miss: the same operator aimed at a selector/path that should not apply.

Reset mutable test state between comparisons. Keep topology, timing policy, and all unrelated fields identical. If the close control triggers the same effect, the test does not isolate the proposed weakness.

Interpreting the evidence

Evidence observed Defensible conclusion
Original and delivered differ AIT altered the communication
Receiver returned a correlated response The receiver processed the delivered request
Receiver selected a different path/tool The mutation influenced receiver behavior
Controlled ledger/callback/state changed The mutation produced an independently observed effect
Close control stayed inert The authoritative change, not generic disturbance, explains the effect
Only a trace, marker echo, or model statement changed Interesting observation; external impact is not proven

Running a boundary as a primitive

Every boundary in the tables above ships as an executable attack primitive, so the unit of work does not have to be reassembled by hand each time. A primitive carries the selector, the injection, the trust hypothesis, what to watch for, the close control with its rationale, and an explicit limitation.

ait intercept attack --target MESSAGE_ID          # what applies to this message
ait intercept attack mcp.tool-shadowing --target MESSAGE_ID
ait intercept attack mcp.tool-shadowing --target NEXT_MESSAGE_ID --arm control

The control arm is not optional. Without it a positive result cannot be separated from the effect of any disturbance, so the attack arm prints the control to run next and why that particular control isolates this boundary.

Primitives are data under ait/data/attack-primitives/. Add your own for a target-specific boundary without modifying the tool; the same fields are required, including the limitation, because a primitive that claims no limit is one nobody can review.

Credentials

Credential-bearing headers -- Authorization, Cookie, X-Api-Key and friends -- are captured verbatim and are fully editable. Forge one, downgrade one, or strip one by omitting it from the edited envelope.

That makes credential relay, replay, token downgrade, and confused-deputy- via-token testable, which they were not while the tool masked those values. It also puts weight on the export contract: --redacted (the default) carries allowlisted metadata only, and --raw is a complete capture labelled as such. Hand a client the redacted one.

A delivery whose credential headers differ from what the sender supplied is marked credential_edited, and chain 2.2 binds the envelope into the record hash, so the transcript never presents a forged token as the sender's own.

Originating traffic

Waiting for an agent to send the message you want to test is not always practical.

ait intercept send --path /message --body-json @probe.json
ait intercept send --from MESSAGE_ID --set /params/message/metadata/amount=999

--from is the repeater. It takes a captured request, applies JSON-Pointer edits, and delivers it again. Injected traffic goes through the same relay as captured traffic, so break conditions, rewrite rules, and the transcript all apply -- which also means injecting while a matching break condition is armed will pause your own message. Turn interception off, or narrow the break conditions, when you want to inject freely.

Producing a finding

The interpretation table above says what each observation licenses you to claim. ait intercept evidence applies it to the arms you actually ran.

ait intercept evidence --primitive mcp.tool-shadowing --out finding.json

Two tiers the tool determines itself: whether a mutation was delivered, and whether a correlated response came back. Two it will not infer:

  • behavior_changed needs a baseline to compare receiver behaviour against.
  • external_effect_observed needs an out-of-band oracle -- a controlled tool or action ledger, a callback receipt, or datastore state read independently of the target. A target response is not an oracle.

Pass --behavior-changed or --external-effect only for something you observed. Left unset, those tiers are reported as unproven along with what would be required, and tiers cannot be skipped: claiming an external effect over a delivery the receiver never processed does not promote the verdict.

The finding also records attribution, by comparing what the receiver did on each arm. With no control arm the status is no_control; if the control produced the same observable effect it is not_isolated; if the arms cannot be compared -- no attack arm, or no captured behaviour for one of them -- it is undetermined, because a missing measurement is not evidence of sameness. Only a measured difference between the arms yields attributable.

Both arms must be in the same session. ait lab reset starts a new one.

Turning a manual edit into a reusable rule

After a successful modified delivery, select Apply to future. Review the derived selector and patch before enabling it. Narrow the selector by protocol, direction, operation, route, or destination so the rule cannot alter unrelated traffic.

Start rules disabled on a real assessment, observe one matching original, then enable deliberately. The hit count confirms matching; History confirms what was actually delivered.

When a rule is armed, a paused message shows the rule's output, not what the sender sent. The sender's version stays available as wire_original, and Sender's version (forward_wire) delivers it, bypassing the rule for that one message.

Suggested assessment sequence

  1. Reproduce the intended behavior in the local approval lab.
  2. Use Use this setup with my target to copy transport and break behavior.
  3. Forward originals until the important agent and tool boundaries are clear.
  4. Test identity and authority fields before experimenting with framing.
  5. Test returned artifacts/tool results after requests.
  6. Test duplicate, drop, replay, and ordering behavior.
  7. Save only the smallest useful edits as live rules.
  8. Export the session with original/delivered values and independent evidence.
  9. Retest after remediation using the same controlled input and observation.

See Interception screen guide for every control and Hands-on interception lab for the deterministic practice environment. Use the direct interception assessment scenarios for twelve repeatable exercises and the operator command cookbook for every terminal action.