A guided path through the lab¶
9 stages about 4 hours local by default evidence first
Nine exercises in an order that builds. Each stage adds one idea, and the whole path is about four hours if you do the controls properly, which is the part that matters and the part people skip.
Everything runs locally. No Docker, provider key, or external network is required; the exercises use loopback sockets and local pipes.
Before you start:
The sequence is intentional. Stages 1 and 2 establish the evidence discipline used to interpret every later technique.
Progress map¶
| Stage | Boundary or skill | Exercise | Budget | Leave with |
|---|---|---|---|---|
| 1 | observe the loop | delegated A2A message | 20 min | a complete unchanged delivery timeline |
| 2 | isolate one field | delegated A2A message | 30 min | one scoreable attack/control pair |
| 3 | identity binding | task and tenant crossover | 25 min | receiver and access-ledger proof |
| 4 | route authority | Agent Card substitution | 30 min | a discovery-to-shadow timeline |
| 5 | MCP sources of truth | tool result + tool schema | 45 min | response and discovery proof chains |
| 6 | request authority | MCP request tampering | 30 min | server-side receipt evidence |
| 7 | event lifecycle | asynchronous artifact | 35 min | release order and idempotency evidence |
| 8 | dependent mutations | tool-schema-induced call | 40 min | chain evidence with causation caveat |
| 9 | handoff | best completed exercise | 20 min | redacted export and bounded finding |
Keep one short operator log as you work: lab ID, session ID, message IDs, attack edit, control edit, strongest evidence tier, and one sentence naming what remains unproven. This prevents a later screenshot from becoming detached from the run that produced it.
Stage 1: See the loop once¶
Goal: understand what "in path" means, and where the truth lives.
ait lab start --exercise delegated-a2a-message
ait lab trigger LAB_ID
ait intercept pending
ait intercept show MESSAGE_ID
Read the message. A gateway agent is delegating a refund approval to a
specialist. Note that metadata.amount is 25 and that nothing in the message
proves the gateway was allowed to ask for it.
Forward it unchanged, then look at where the answer actually shows up:
The idea: the specialist's reply is narration. A separate controlled process records the specialist's effect request as a ledger row. From here on, read both sources and do not let either stand in for the other.
Check yourself: what in that message would you have to change to make the specialist approve more money? Predict it before Stage 2.
Exit criterion: you can identify the sender, AIT boundary, receiver, and effect ledger without relying on the UI layout.
Stage 2: Change one thing, then prove it was your change¶
Goal: the attack/control discipline, which is the whole evidence model.
ait lab trigger LAB_ID
ait intercept edit MESSAGE_ID --arm attack --set /params/message/metadata/amount=750
ait lab status LAB_ID
The ledger should show 750. Now the part that makes it mean something:
ait lab trigger LAB_ID
ait intercept edit NEXT_ID --arm control --set /params/message/metadata/amount=26
ait lab status LAB_ID
ait intercept evidence --primitive a2a.approval-amount
Both arms in one session. ait lab reset starts a new session and splits
them, which leaves neither scoreable.
The idea: a control is a near-miss: same field, same visibility, a value
that does not cross the boundary. If the control moves the target too, you have
shown the target reacts to change, not to authority. The finding reports that
as not_isolated, and that is a real result, not a failure.
Check yourself: run the control with amount=750 instead. What attribution
do you get, and why is it right?
Exit criterion: both arms appear in one session and the evidence result distinguishes “changed” from “isolated.”
Stage 3: Cross a boundary instead of moving a number¶
Goal: identity is data on the wire, and data can be rewritten.
ait lab start --exercise a2a-identity-crossover
ait intercept attack a2a.tenant-crossover --target MESSAGE_ID
The idea: the receiver rebinds work to whichever tenant the message claims. The control here is the instructive one: change the tenant to another tenant the sender legitimately owns. If that is accepted too, you have not shown a crossover; you have shown it accepts tenant changes.
Exit criterion: the access ledger, not the modified message, tells you which tenant record was consumed.
Stage 4: Move where the traffic goes, and catch it¶
Goal: proving what an attacker receives, not just that routing moved.
ait lab start --exercise a2a-agent-card-substitution
ait intercept shadow start a2a-agent
ait intercept attack a2a.card-endpoint-substitution --target MESSAGE_ID
ait intercept shadow observations
The idea: without the shadow, the finding is "the URL changed". With it, the finding is "the delegation arrived at my endpoint carrying this task id, these arguments, and this bearer token." That record is out-of-band by construction: your endpoint wrote it, not the target, which is what the top evidence tier requires.
Check yourself: stop the shadow and run it again. Notice the primitive refuses rather than redirecting into nothing.
Exit criterion: you can join the card response to the next request and to a shadow receipt while showing that the primary stayed untouched.
Stage 5: Attack the agent's sources of truth¶
Goal: the MCP boundaries, where the target reads rather than receives.
ait lab start --exercise mcp-tool-result
ait intercept attack mcp.tool-result-alteration --target MESSAGE_ID
Then discovery, which is the sharper one:
ait lab start --exercise mcp-tool-schema
ait intercept attack mcp.tool-shadowing --target MESSAGE_ID
ait lab status LAB_ID
Read decision_reason in the ledger, not just the numbers:
tool_called: assess_risk_v2 steered: true
decision_reason: selected 'assess_risk_v2' over 'approval_policy' because its
description claimed precedence
The idea: the consumer chose the inserted tool because its description
claimed precedence. And the caveat that matters: the lab's consumer is
deterministic. A real client's selection is model-mediated and probabilistic, so
a positive here proves your injection is well-formed and reaches the decision
point, not that an external client would obey it. Every one of these primitives says
so in its limitation. Read them.
Exit criterion: you can show the difference between altering an existing tool result and altering discovery information used before a call.
Stage 6: Attack the request, not the response¶
Goal: a different trust boundary with a different owner.
ait lab start --exercise mcp-request-tampering
ait intercept attack mcp.resource-uri-substitution --target MESSAGE_ID
The idea: response tampering asks whether an agent trusts what it is told. Request tampering asks whether a server trusts what it is asked. A server that validates arguments carefully can still be undone by a poisoned catalogue, and one that serves whatever URI it is handed can be undone even when every response is authentic.
Exit criterion: the MCP server receipt shows the URI or arguments it actually received, and the control changes a nearby non-authoritative field.
Stage 7: Work asynchronously¶
Goal: streams, ordering, and idempotency.
ait lab start --exercise asynchronous-artifact
ait intercept replay MESSAGE_ID --copies 3
ait lab status LAB_ID
The idea: duplicate delivery is an attack on its own. Three ledger rows from one message is an idempotency finding, and it needs no content change at all. Try dropping the first artifact revision and forwarding the second.
Exit criterion: requested order, actual release order, subscriber order, and duplicate-processing count are all accounted for.
Stage 8: Chain two steps¶
Goal: an attack whose second step depends on the first.
ait intercept transform arm agentic-redteam/seam/rules/chains/mcp-tool-schema-induced-call/attack/
ait lab start --exercise mcp-tool-schema
ait intercept status
Poison the schema, then let the induced call carry the escalated argument.
The idea: the chain guard records that AIT delivered step one before allowing step two. That is bookkeeping, not causation: the target may have produced the second message regardless. Every chained finding carries that caveat. To show inducement you need the control pack, which runs the same chain with an in-policy payload.
Exit criterion: the chain record proves ordering, the control pack tests inducement, and your note does not confuse the two.
Stage 9: Turn it into something you could hand over¶
Goal: the deliverable.
ait intercept evidence --primitive mcp.tool-shadowing --out finding.json
ait intercept export SESSION_ID --format jsonl --redacted --out session.jsonl
ait intercept stop SESSION_ID --pending forward
Read finding.json properly. Look at unproven: the tiers you did not reach
and the reason. Look at attribution. Look at limitation.
The idea: a finding that claims less than you hoped is doing its job. The failure mode that destroys a security tool's credibility is claiming impact it did not observe.
Exit criterion: another operator can trace the finding back to the run, understand the attack/control difference, and see every unproven tier.
Verify the chain, and note that you can still score after teardown:
agentic-redteam/seam/seam transcript verify \
--schema agentic-redteam/schema/transcript.schema.json \
--transcript .ait/runs/RUN_ID/transcript.json
After the path¶
Re-run Stage 2 with --a2a-binding rest and --a2a-binding grpc. Then re-run
Stage 5 with --mcp-transport stdio. The bindings differ in ways that
have caught real defects, and stdio is the transport most real MCP deployments
use.
Then read How the lab is built for what the lab cannot teach you: placement, TLS, model-mediated behaviour, and scale. Read Taking it to a real target for what changes outside the controlled lab.
Training completion checklist¶
- I can distinguish capture, delivery, processing, behavior, and effect.
- I keep attack and close control in one session.
- I choose a proof source the intercepted receiver does not narrate.
- I separate identity, routing, content, request, and lifecycle claims.
- I know when deterministic consumer behavior does not predict a model.
- I can stop a finding at the evidence tier actually reached.
- I can export a redacted run another operator can verify.