Skip to content

Practising the techniques

Nine exercises, twenty-four techniques, no Docker and no provider key. Each exercise stands up real A2A and MCP agents using the official SDKs, so you are practising against a protocol implementation rather than a mock.

Work through them in order. Each builds on the decision loop of the one before.

The loop, once

Every technique is the same four moves:

  1. Pause the message that carries the trust decision.
  2. Change one thing.
  3. Look at what the receiver did: in the effect ledger, not in the reply.
  4. Run the control: the near-miss that should leave the receiver alone.

Step 4 is the one people skip and the one that makes a result mean anything. If the control moves the target too, you learned that any disturbance moves it, not that your change did.

ait lab start --exercise delegated-a2a-message
ait lab trigger LAB_ID
ait intercept pending
ait intercept attack a2a.approval-amount --target MESSAGE_ID
ait lab status LAB_ID          # what the receiver actually did

Then trigger the exercise again and run the same primitive with --arm control. Both arms must be in one session -- ait lab reset starts a new one and splits them, leaving neither scoreable.


1. Change a delegated parameter

Exercise: delegated-a2a-message: a gateway delegates a refund approval.

Technique Boundary What you are testing
a2a.approval-amount Approval parameters Is the amount trusted because it arrived from a peer agent?
a2a.delegated-instruction Delegated intent Does the instruction text steer the specialist independently of the metadata?
credential.forge Actor and authority Does the receiver check who is asking, or only that someone asked?

Start here. It is the whole loop with one field, and the effect ledger shows the approved amount so the result is unambiguous.

The credential technique is where the tool stops being a proxy and starts being an attack instrument. It forges, downgrades and strips Authorization, Cookie, and X-Api-Key on the envelope.

The gateway authenticates to the specialist, so there is a real credential to act on. Watch the ledger, not the reply:

attack   presented_principal=unverified:operator-forged-token  credential_verified=false
control  presented_principal=gateway-agent                     credential_verified=true

The specialist acted on a credential it could not verify. The token mapping is trivial and public on purpose. This shows whether a receiver checks what it was handed, not whether a token can be cracked.

2. Cross an identity boundary

Exercise: a2a-identity-crossover

Technique Boundary
a2a.context-crossover Context binding
a2a.tenant-crossover Tenant binding

Two agents share a conversation. The question is whether the receiver rebinds work to whichever context or tenant the message claims, or checks it against what it already knows.

The control here is instructive: change the context to one that also belongs to the sender. If the receiver accepts that too, you have not shown a crossover: you have shown it accepts context changes.

3. Move a route

Exercises: a2a-agent-card-substitution, a2a-callback-replay

Technique Boundary
a2a.card-endpoint-substitution Agent Card route
a2a.callback-substitution Callback destination

Discovery data is mutable and rarely pinned. Rewrite where the next delegation goes, or where the completion callback lands.

Start a shadow first. Both techniques redirect traffic, and a redirect into nothing proves only that routing moved:

ait intercept shadow start a2a-agent
ait intercept attack a2a.card-endpoint-substitution --target MESSAGE_ID
ait intercept shadow observations

The delegation arriving at your endpoint with its task ID, arguments, and headers turns "the route changed" into "here is what I received." It is also the out-of-band oracle that lets the finding reach external_effect_observed.

Against the card exercise this is the full path, and it is covered by tests/test_card_substitution_into_shadow.py: the A2A client fetches the card through the intercept, you rewrite where it points, and the client's next delegation lands on your endpoint carrying its real payload: SendMessage, the task and context ids, "approve refund 25", and the account and amount metadata.

4. Alter what a tool returns

Exercise: mcp-tool-result

Technique Boundary
mcp.tool-result-alteration Tool result

The client asked a policy server whether something is risky. You answer instead. This is the smallest complete demonstration that an agent's decision is only as good as the bytes it received.

5. Poison discovery

Exercise: mcp-tool-schema

Technique Boundary
mcp.tool-schema-poisoning Tool schema
mcp.tool-shadowing Tool identity

The client picks its call arguments from the schema you deliver. Raise a declared ceiling and watch it ask for more than policy allows.

Tool shadowing is the sharper one: insert a tool ahead of the real one whose description claims precedence. This is why the rewrite vocabulary has positional insert rather than only set: overwriting tools/0 renames the real tool, which is a different and more obvious attack.

Both of these steer by text, so read the ledger's decision_reason, not just its numbers:

tool_called: assess_risk_v2   steered: true
decision_reason: selected 'assess_risk_v2' over 'approval_policy' because its
                 description claimed precedence

6. Tamper with the request

Exercise: mcp-request-tampering

Technique Boundary
mcp.resource-uri-substitution Resource identity
mcp.tool-argument-escalation Tool arguments

Every other MCP exercise pauses a response. This one pauses what the client asks for, which is a different trust boundary with a different owner: a server that validates its arguments can still be undone by a poisoned catalogue, and one that serves whatever is requested can be undone by a rewritten URI even when every response is authentic.

7. Poison retrieved content

Exercise: mcp-resource-poisoning

Technique Boundary
mcp.resource-content-poisoning Resource content

Guidance the agent retrieves and then acts on. Change the guidance, not the action, and see whether the action follows.

8. Work asynchronously

Exercise: asynchronous-artifact

Technique Boundary
a2a.artifact-content Artifact content
a2a.task-state-forge Task state

Streaming artifacts and task lifecycle. The interesting property is timing: an artifact arrives in revisions, and a task reports states you can rewrite. Try metadata.revision gte 2 as a where clause to act only on the second version.

9. Chain two steps

Exercise: mcp-tool-schema, with the shipped chain pack.

PACK=agentic-redteam/seam/rules/chains/mcp-tool-schema-induced-call
for f in "$PACK"/attack/*.yaml; do ait intercept transform arm "$f"; done
ait lab trigger LAB_ID
ait intercept status                 # chain armed, and which steps delivered

Poison the catalogue, then act on the call it induces, and only once the poison actually reached the wire. Full walkthrough with the control arm in Chained tool-schema attack.


What the lab's consumer is susceptible to

Six techniques work by putting words in front of something that reads them. The lab's consumers read them, but they are not models. They are fixtures with four susceptibilities, declared in every decision they record:

  • it ranks tools by what their descriptions claim rather than by pinned identity;
  • it reads a tool description as instruction and obeys pass KEY=VALUE directives;
  • it reads retrieved resource text as instruction rather than as data;
  • it reads a delegated message or a consumed artifact as instruction when the text carries a directive, and lets it stand in for the authoritative policy check.

The last one is the sharpest to watch: on the attack arm the ledger records decided_by: delegated_instruction and policy_consulted: false, so you can see the specialist obeyed the delegated text instead of the policy tool it was supposed to consult.

A lone status token is still data. This lab's own artifacts carry the word approved as a value, so only a sentence built around a directive verb is read as an instruction -- otherwise a2a.artifact-content's control arm, which renames the artifact and leaves its text alone, would be indistinguishable from the attack.

Each is a deliberately vulnerable fixture policy. Practising against it shows that, for a consumer with this declared policy, the injection is well-formed and reaches the intended decision point.

What it does not tell you is whether any particular client has that weakness. A deterministic reader makes instruction injection look perfectly reliable; the real technique is model-mediated and probabilistic. Every affected primitive says so in its limitation, and a test enforces that. Treat a positive here as evidence your injection is well-formed and reaches the decision point, then measure what an external client does with it under a separate protocol.

What you cannot practise here

Four techniques ship with no exercise. mcp.elicitation-url-substitution and mcp.sampling-restriction-removal both target server-initiated requests, but the packaged MCP server runs stateless with JSON responses. Originating one needs a bidirectional session plus a client that declares the matching capability. Changing the fixture's transport mode to host two techniques would put four working exercises at risk, so they stay disclaimed and say exactly what they need.

mcp.prompt-arguments joins them: the packaged MCP server registers no prompts, so nothing here produces a prompts/get to intercept.

transport.correlation-id is unanchored for a different reason worth understanding: breaking correlation strands the sender by design, so an exercise anchored to it could never run to completion and the harness would report a working technique as broken.

These four primitives remain catalogue hypotheses. Their selectors address the required wire direction, but no public exercise or external-target run validates their effect.

ait doctor fails if any primitive ships without a close control or a stated limitation, and tests/test_primitives_run_in_the_lab.py fails if a primitive claims an exercise it cannot actually match. A technique nobody can practise is tracked as a gap rather than presented as a capability.

Reading the result

After any technique:

ait intercept evidence --out finding.json

The verdict is the strongest tier the evidence supports, and the tiers cannot be skipped:

Tier Earned by
mutation_delivered the transcript shows original and delivered differ
receiver_processed a correlated response came back
behavior_changed you observed a different path, out of band
external_effect_observed a controlled endpoint or ledger recorded the effect

The tool determines the first two and refuses to infer the other two. A target's own response is not an oracle: that is the rule the whole design rests on, and the reason a shadow endpoint exists.