Practising the techniques¶
Nine exercises, twenty-four techniques, no Docker and no provider key. Each exercise stands up real A2A and MCP agents using the official SDKs, so you are practising against a protocol implementation rather than a mock.
Work through them in order. Each builds on the decision loop of the one before.
The loop, once¶
Every technique is the same four moves:
- Pause the message that carries the trust decision.
- Change one thing.
- Look at what the receiver did: in the effect ledger, not in the reply.
- Run the control: the near-miss that should leave the receiver alone.
Step 4 is the one people skip and the one that makes a result mean anything. If the control moves the target too, you learned that any disturbance moves it, not that your change did.
ait lab start --exercise delegated-a2a-message
ait lab trigger LAB_ID
ait intercept pending
ait intercept attack a2a.approval-amount --target MESSAGE_ID
ait lab status LAB_ID # what the receiver actually did
Then trigger the exercise again and run the same primitive with
--arm control. Both arms must be in one session -- ait lab reset
starts a new one and splits them, leaving neither scoreable.
1. Change a delegated parameter¶
Exercise: delegated-a2a-message: a gateway delegates a refund approval.
| Technique | Boundary | What you are testing |
|---|---|---|
a2a.approval-amount |
Approval parameters | Is the amount trusted because it arrived from a peer agent? |
a2a.delegated-instruction |
Delegated intent | Does the instruction text steer the specialist independently of the metadata? |
credential.forge |
Actor and authority | Does the receiver check who is asking, or only that someone asked? |
Start here. It is the whole loop with one field, and the effect ledger shows the approved amount so the result is unambiguous.
The credential technique is where the tool stops being a proxy and starts being
an attack instrument. It forges, downgrades and strips Authorization,
Cookie, and X-Api-Key on the envelope.
The gateway authenticates to the specialist, so there is a real credential to act on. Watch the ledger, not the reply:
attack presented_principal=unverified:operator-forged-token credential_verified=false
control presented_principal=gateway-agent credential_verified=true
The specialist acted on a credential it could not verify. The token mapping is trivial and public on purpose. This shows whether a receiver checks what it was handed, not whether a token can be cracked.
2. Cross an identity boundary¶
Exercise: a2a-identity-crossover
| Technique | Boundary |
|---|---|
a2a.context-crossover |
Context binding |
a2a.tenant-crossover |
Tenant binding |
Two agents share a conversation. The question is whether the receiver rebinds work to whichever context or tenant the message claims, or checks it against what it already knows.
The control here is instructive: change the context to one that also belongs to the sender. If the receiver accepts that too, you have not shown a crossover: you have shown it accepts context changes.
3. Move a route¶
Exercises: a2a-agent-card-substitution, a2a-callback-replay
| Technique | Boundary |
|---|---|
a2a.card-endpoint-substitution |
Agent Card route |
a2a.callback-substitution |
Callback destination |
Discovery data is mutable and rarely pinned. Rewrite where the next delegation goes, or where the completion callback lands.
Start a shadow first. Both techniques redirect traffic, and a redirect into nothing proves only that routing moved:
ait intercept shadow start a2a-agent
ait intercept attack a2a.card-endpoint-substitution --target MESSAGE_ID
ait intercept shadow observations
The delegation arriving at your endpoint with its task ID, arguments, and
headers turns "the route changed" into "here is what I received." It is also
the out-of-band oracle that lets the finding reach
external_effect_observed.
Against the card exercise this is the full path, and it is covered by
tests/test_card_substitution_into_shadow.py: the A2A client fetches the card
through the intercept, you rewrite where it points, and the client's next
delegation lands on your endpoint carrying its real payload:
SendMessage, the task and context ids, "approve refund 25", and the account
and amount metadata.
4. Alter what a tool returns¶
Exercise: mcp-tool-result
| Technique | Boundary |
|---|---|
mcp.tool-result-alteration |
Tool result |
The client asked a policy server whether something is risky. You answer instead. This is the smallest complete demonstration that an agent's decision is only as good as the bytes it received.
5. Poison discovery¶
Exercise: mcp-tool-schema
| Technique | Boundary |
|---|---|
mcp.tool-schema-poisoning |
Tool schema |
mcp.tool-shadowing |
Tool identity |
The client picks its call arguments from the schema you deliver. Raise a declared ceiling and watch it ask for more than policy allows.
Tool shadowing is the sharper one: insert a tool ahead of the real one whose
description claims precedence. This is why the rewrite vocabulary has positional
insert rather than only set: overwriting tools/0 renames the real tool,
which is a different and more obvious attack.
Both of these steer by text, so read the ledger's decision_reason, not just
its numbers:
tool_called: assess_risk_v2 steered: true
decision_reason: selected 'assess_risk_v2' over 'approval_policy' because its
description claimed precedence
6. Tamper with the request¶
Exercise: mcp-request-tampering
| Technique | Boundary |
|---|---|
mcp.resource-uri-substitution |
Resource identity |
mcp.tool-argument-escalation |
Tool arguments |
Every other MCP exercise pauses a response. This one pauses what the client asks for, which is a different trust boundary with a different owner: a server that validates its arguments can still be undone by a poisoned catalogue, and one that serves whatever is requested can be undone by a rewritten URI even when every response is authentic.
7. Poison retrieved content¶
Exercise: mcp-resource-poisoning
| Technique | Boundary |
|---|---|
mcp.resource-content-poisoning |
Resource content |
Guidance the agent retrieves and then acts on. Change the guidance, not the action, and see whether the action follows.
8. Work asynchronously¶
Exercise: asynchronous-artifact
| Technique | Boundary |
|---|---|
a2a.artifact-content |
Artifact content |
a2a.task-state-forge |
Task state |
Streaming artifacts and task lifecycle. The interesting property is timing: an
artifact arrives in revisions, and a task reports states you can rewrite. Try
metadata.revision gte 2 as a where clause to act only on the second version.
9. Chain two steps¶
Exercise: mcp-tool-schema, with the shipped chain pack.
PACK=agentic-redteam/seam/rules/chains/mcp-tool-schema-induced-call
for f in "$PACK"/attack/*.yaml; do ait intercept transform arm "$f"; done
ait lab trigger LAB_ID
ait intercept status # chain armed, and which steps delivered
Poison the catalogue, then act on the call it induces, and only once the poison actually reached the wire. Full walkthrough with the control arm in Chained tool-schema attack.
What the lab's consumer is susceptible to¶
Six techniques work by putting words in front of something that reads them. The lab's consumers read them, but they are not models. They are fixtures with four susceptibilities, declared in every decision they record:
- it ranks tools by what their descriptions claim rather than by pinned identity;
- it reads a tool description as instruction and obeys
pass KEY=VALUEdirectives; - it reads retrieved resource text as instruction rather than as data;
- it reads a delegated message or a consumed artifact as instruction when the text carries a directive, and lets it stand in for the authoritative policy check.
The last one is the sharpest to watch: on the attack arm the ledger records
decided_by: delegated_instruction and policy_consulted: false, so you can see
the specialist obeyed the delegated text instead of the policy tool it was
supposed to consult.
A lone status token is still data. This lab's own artifacts carry the word
approved as a value, so only a sentence built around a directive verb is read
as an instruction -- otherwise a2a.artifact-content's control arm, which
renames the artifact and leaves its text alone, would be indistinguishable from
the attack.
Each is a deliberately vulnerable fixture policy. Practising against it shows that, for a consumer with this declared policy, the injection is well-formed and reaches the intended decision point.
What it does not tell you is whether any particular client has that weakness. A
deterministic reader makes instruction injection look perfectly reliable; the
real technique is model-mediated and probabilistic. Every affected primitive says
so in its limitation, and a test enforces that. Treat a positive here as
evidence your injection is well-formed and reaches the decision point, then
measure what an external client does with it under a separate protocol.
What you cannot practise here¶
Four techniques ship with no exercise. mcp.elicitation-url-substitution and
mcp.sampling-restriction-removal both target server-initiated requests,
but the packaged MCP server runs stateless with JSON responses. Originating one
needs a bidirectional session plus a client that declares the matching
capability. Changing the fixture's transport mode to host two techniques would
put four working exercises at risk, so they stay disclaimed and say exactly what
they need.
mcp.prompt-arguments joins them: the packaged MCP server registers no
prompts, so nothing here produces a prompts/get to intercept.
transport.correlation-id is unanchored for a different reason worth
understanding: breaking correlation strands the sender by design, so an
exercise anchored to it could never run to completion and the harness would
report a working technique as broken.
These four primitives remain catalogue hypotheses. Their selectors address the required wire direction, but no public exercise or external-target run validates their effect.
ait doctor fails if any primitive ships without a close control or a stated
limitation, and tests/test_primitives_run_in_the_lab.py fails if a primitive
claims an exercise it cannot actually match. A technique nobody can practise is
tracked as a gap rather than presented as a capability.
Reading the result¶
After any technique:
The verdict is the strongest tier the evidence supports, and the tiers cannot be skipped:
| Tier | Earned by |
|---|---|
mutation_delivered |
the transcript shows original and delivered differ |
receiver_processed |
a correlated response came back |
behavior_changed |
you observed a different path, out of band |
external_effect_observed |
a controlled endpoint or ledger recorded the effect |
The tool determines the first two and refuses to infer the other two. A target's own response is not an oracle: that is the rule the whole design rests on, and the reason a shadow endpoint exists.