How the lab is built¶
The default lab is a controlled, process-backed protocol fixture. It starts separate programs using the packaged A2A and MCP SDKs, sends complete messages over sockets or pipes, places the production Seam interception engine on one selected hop, and records receiver and effect state in separate stores.
Its identities, tasks, policy, values, and effects are synthetic by design. That makes the fixture disposable and repeatable. It does not make a lab result external validity, and it does not show that an unrelated agent or model will make the same decision.
The protocol and process paths are real. The roles and business effect are controlled. Separate process ownership prevents the transcript, receiver state, and effect observation from becoming one self-confirming record.
The four processes¶
ait lab start launches up to four components plus the interception session:
| Component | What it is | Why it exists |
|---|---|---|
| gateway | An A2A agent that delegates work | The sender. It is the thing whose message you intercept. |
| specialist | An A2A agent that receives the delegation and acts | The receiver. Its state is what an attack has to move. |
| mcp | An MCP server exposing a policy tool and a policy resource | The authority the specialist consults. Poisoning it is a different boundary from poisoning the delegation. |
| executor | The effect ledger | Where the receiver writes down what it actually did. |
Some exercises add a fifth: a shadow specialist for Agent Card substitution, or a callback receiver for push-notification replay.
AIT sits between two of these rather than modifying their program code. The sender is explicitly pointed at the listener or wrapper, so this is not covert placement. It rehearses the same relay or wrapper boundary an authorized external assessment can use when that sender is configurable.
Launch and ownership lifecycle¶
ait lab start creates a supervised run rather than a loose collection of
background commands. Startup is ordered so that a sender cannot emit traffic
before the proof sources and interception listener are ready.
run + ports Start
ledger + receiver Start AIT
listener Start
sender Verify
readiness
| Phase | AIT verifies | Failure behavior |
|---|---|---|
| allocate | run identity, loopback ports, exercise configuration | no child process starts if allocation fails |
| receiver first | health/readiness of ledger, specialist, MCP server, or shadow | sender remains stopped |
| interception | listener binding and configured upstream | traffic is not triggered until the boundary is usable |
| sender last | sender process and target URL | failed child is reported with its process receipt |
| stop | pending-message disposition and owned process termination | the run remains inspectable even after teardown |
The run owns only the processes it launched. AIT records process receipts and uses them during cleanup; it does not sweep unrelated listeners or kill a process merely because it uses a familiar port.
What is real and what is controlled¶
The distinction matters when interpreting a result.
| Layer | Real in the default lab | Controlled for repeatability |
|---|---|---|
| process boundary | separate supervised operating-system processes | all run on loopback |
| protocol implementation | packaged a2a-sdk and mcp clients/servers |
fixed exercise operations and payloads |
| transport | sockets for A2A/MCP HTTP; pipes for MCP stdio | ephemeral local endpoints |
| interception engine | the same direct interception path used outside the lab | preconfigured break conditions |
| consumer decision | the receiver process parses the representation Seam actually delivers | deterministic, declared rules rather than an undeclared model |
| effect observation | a separate controlled service records a receiver request or resulting state | synthetic accounts, tenants, tasks, values, and effects |
--fixture selects the older handwritten topology for compatibility tests. It
is not the recommended learning mode. Optional reasoning exercises replace the
deterministic decision with a configured model call and therefore require a
different measurement contract.
Why the effect ledger is separate¶
When the specialist approves a payment, it does not just answer the gateway. It posts to the executor, and that separate process records the request as a row. The row is read out of band from the intercepted reply:
ait lab status LAB_ID
effect ledger: {'kind': 'approval', 'value': 750.0, 'account': 'acct-demo', ...}
This is the distinction the evidence model rests on. A target's reply is narration; it can disagree with what the receiver actually did. A ledger row is a controlled observation authored outside the interception decision. It can show that the receiver attempted the fixture effect. It is not independent of the research owner, and it is not evidence that the same effect would occur in an external deployment.
When you read ait lab status, read the ledger, not the reply.
A run stops at the strongest tier it actually supports. Missing proof is reported as unproven rather than inferred from the target's narration.
One message, six operator states¶
The traffic record changes state as the operator works it:
- captured: the listener received a real request, response, or event;
- pending: a break condition stopped delivery;
- decided: forward, edit, drop, replay, answer, or batch order was chosen;
- delivered: the downstream transport accepted the chosen bytes;
- observed: receiver state or a ledger receipt was collected;
- compared: attack and close control were scored in one session.
These states answer different questions. A message can be successfully edited but never delivered, delivered but rejected by the receiver, processed without changing behavior, or change behavior without an out-of-band effect source. The evidence view keeps those outcomes separate.
The consumers read, and they say what they are susceptible to¶
Six techniques work by putting words in front of something that reads them: poisoned tool descriptions, shadow tools, injected resource text, rewritten delegated instructions. For those to demonstrate anything, something in the lab has to actually read the text and act on it.
The lab's consumers do. They are not models; they are fixtures with four declared susceptibilities, and every decision they record states them:
- they rank tools by what descriptions claim rather than by pinned identity;
- they obey
pass KEY=VALUEdirectives found in a tool description; - they treat retrieved resource text as instruction rather than as data;
- they treat a delegated message or consumed artifact as instruction when it carries a directive, and let it stand in for the policy check.
These are deliberately vulnerable fixture policies selected to exercise the
interception and evidence path. A positive result proves the payload is
well-formed for this fixture and reaches its declared decision point. It does
not establish prevalence in shipping clients or predict whether a
model-mediated client will comply. Every affected primitive carries that bound
in its limitation.
A lone status token stays data. The lab's own artifacts carry approved as a
value, so only a sentence built around a directive verb is read as an
instruction. Without that line, the control arm of a2a.artifact-content,
which renames the artifact and leaves the text alone, would be indistinguishable
from the attack.
The nine exercises¶
| Exercise | Protocol | What it stands up | Proof source |
|---|---|---|---|
delegated-a2a-message |
A2A JSON-RPC | gateway → specialist → MCP policy | specialist state and approval ledger |
a2a-identity-crossover |
A2A JSON-RPC | tenant-partitioned specialist | tenant-partitioned state and access ledger |
a2a-agent-card-substitution |
A2A JSON-RPC | primary and shadow specialists | primary-versus-shadow state and routing ledger |
a2a-callback-replay |
A2A JSON-RPC | callback receiver | callback receipt and idempotency ledger |
asynchronous-artifact |
SSE | streaming artifact revisions | subscriber history and idempotency ledger |
mcp-tool-result |
MCP HTTP | MCP policy server | client decision and effect ledger |
mcp-tool-schema |
MCP HTTP | MCP discovery surface | call receipt and decision ledger |
mcp-resource-poisoning |
MCP HTTP | MCP resource surface | resource receipt, later call, effect ledger |
mcp-request-tampering |
MCP HTTP | request-side MCP boundary | server receipt and effect ledger |
delegated-a2a-message also runs on --a2a-binding rest and --a2a-binding
grpc, and mcp-tool-result on --mcp-transport stdio. Use them: the bindings
differ in ways that have caught real defects, and stdio is the transport most
real MCP deployments use.
Run artifacts and their authority¶
The lab run record uses the ait.intercept-lab/v3 contract. The exact storage
path is workspace-dependent, so use ait lab status and the Cockpit links
instead of assuming a directory. A complete run can include:
| Artifact | Produced by | Use it for |
|---|---|---|
| exercise configuration | AIT launcher | protocol, binding, transport, synthetic identities |
| process receipts | supervisor | command ownership, readiness, exit, SDK/runtime version |
| traffic transcript | interception session | original bytes, operator decisions, delivered bytes, correlation |
| receiver state | specialist, gateway, subscriber, or MCP server | what the consumer parsed and retained |
| effect/access/routing ledger | controlled observation service | out-of-band record of the fixture action or destination |
| evidence result | AIT evidence scorer | attack/control attribution, achieved tier, limitations |
No single artifact substitutes for the others. The transcript is authoritative about the exchange; receiver state is authoritative about local consumption; the ledger is authoritative about its controlled observation. The final claim is their intersection.
Reset, and why arms must share a session¶
ait lab reset returns the components to a known state between comparisons. It
does not preserve the interception session. It starts a new one, so an
attack run before a reset and a control run after it land in different sessions
and cannot be scored against each other.
Run both arms in one session instead. Trigger the exercise twice:
ait lab trigger LAB_ID
ait intercept edit MESSAGE_ID --arm attack --set /params/message/metadata/amount=750
ait lab trigger LAB_ID
ait intercept edit NEXT_ID --arm control --set /params/message/metadata/amount=26
ait intercept evidence --primitive a2a.approval-amount
Reset when you want a clean ledger for a fresh comparison, not between the two arms of one.
What the lab cannot teach you¶
Being explicit about this is part of the lab's job.
- Placement. Every component is on loopback and expects to be pointed at AIT. A real engagement's hard part is getting in path at all, and the lab skips it entirely.
- TLS deployment. The default fixture uses loopback HTTP or pipes, so it does not exercise certificate issuance, private CAs, mTLS, DNS, or trust-store changes. AIT can use an explicit HTTPS listener, upstream CA material, and client certificates, but those are external placement decisions.
- Model-mediated behaviour. The consumers are deterministic. Real instruction injection is probabilistic, and a single positive against a real client means much less than a positive here.
- Scale. Exercises produce a handful of messages. A real session runs for hours across thousands, and nothing here rehearses that.
Two techniques have no exercise at all: mcp.elicitation-url-substitution and
mcp.sampling-restriction-removal because both need server-initiated requests
and the packaged MCP server runs stateless. They say so in their limitation,
and ait doctor fails if a primitive ships without one.
Diagnose a lab that does not move¶
Work from transport to effect. Do not start by changing the payload again.
| Symptom | Inspect first | Likely distinction |
|---|---|---|
| no pending message | trigger status, break operation, selected protocol | traffic never reached the intended interception rule |
| request pending forever | operator decision and downstream readiness | capture succeeded; delivery did not happen |
| delivered, receiver unchanged | receiver parse/error state and exact field path | wire mutation was accepted by AIT but not consumed |
| receiver changed, ledger empty | ledger readiness and receiver-to-ledger call | processing occurred without an out-of-band effect observation |
| attack and control both move | arm labels, shared session, control design | behavior is not isolated to the tested property |
| stale state after a run | lab ID and reset receipt | evidence from two runs may be mixed |
ait lab status LAB_ID
ait intercept pending
ait intercept show MESSAGE_ID
ait intercept evidence --primitive PRIMITIVE_ID
Preserve the failed state long enough to inspect it. Reset is useful only after you have captured the reason the current run stopped short.
Beyond the lab¶
Three committed tests run interception against code the lab did not produce. They are listed weakest to strongest, because each closes something the one before it leaves open:
tests/test_against_an_independent_a2a_agent.pydrives an A2A agent that imports nothing fromait, enforces its own ceiling, and calls the valuepayoutrather thanamount. Independent of the tool, but it runs the same Pythona2a-sdkthe lab does, so an assumption baked into that one library is invisible to it.tests/test_against_a_third_party_server.pydrives Anthropic's published@modelcontextprotocol/server-everything, a TypeScript package on npm that has never heard of AIT, over stdio. It asserts an operator's rewritten argument is what the server executes. Different authors and a different language, but MCP.tests/test_against_a_cross_implementation_agent.pyis the strongest: A2A, served by@a2a-js/sdk. A different codebase in a different ecosystem, naming its own field (settlement) and setting its own ceiling. An edit that lands here lands because it is correct on the wire, not because both sides agreed on a spelling. Writing it found two things the Python fixtures could not: ataskIdthe server has not issued is rejected outright, and this SDK carries message parts in a protobuf-derived shape internally.
Read those when you want to know what the tool does against something that was
not built around it. The first runs in the default suite; the other two need
node and npm and are run by third-party-interception.yml, which fails if
either skips.