Skip to content

How the lab is built

The default lab is a controlled, process-backed protocol fixture. It starts separate programs using the packaged A2A and MCP SDKs, sends complete messages over sockets or pipes, places the production Seam interception engine on one selected hop, and records receiver and effect state in separate stores.

Its identities, tasks, policy, values, and effects are synthetic by design. That makes the fixture disposable and repeatable. It does not make a lab result external validity, and it does not show that an unrelated agent or model will make the same decision.

The controlled lab starts a gateway sender, AIT boundary, specialist receiver, MCP policy source, effect ledger, and private run directory as separate responsibilities. The controlled lab starts a gateway sender, AIT boundary, specialist receiver, MCP policy source, effect ledger, and private run directory as separate responsibilities. A mobile view of the lab's gateway, interception boundary, specialist, MCP policy source, controlled effect ledger, and run directory. A mobile view of the lab's gateway, interception boundary, specialist, MCP policy source, controlled effect ledger, and run directory.

The protocol and process paths are real. The roles and business effect are controlled. Separate process ownership prevents the transcript, receiver state, and effect observation from becoming one self-confirming record.

The four processes

ait lab start launches up to four components plus the interception session:

Component What it is Why it exists
gateway An A2A agent that delegates work The sender. It is the thing whose message you intercept.
specialist An A2A agent that receives the delegation and acts The receiver. Its state is what an attack has to move.
mcp An MCP server exposing a policy tool and a policy resource The authority the specialist consults. Poisoning it is a different boundary from poisoning the delegation.
executor The effect ledger Where the receiver writes down what it actually did.

Some exercises add a fifth: a shadow specialist for Agent Card substitution, or a callback receiver for push-notification replay.

AIT sits between two of these rather than modifying their program code. The sender is explicitly pointed at the listener or wrapper, so this is not covert placement. It rehearses the same relay or wrapper boundary an authorized external assessment can use when that sender is configurable.

Launch and ownership lifecycle

ait lab start creates a supervised run rather than a loose collection of background commands. Startup is ordered so that a sender cannot emit traffic before the proof sources and interception listener are ready.

Allocate
run + ports
Start
ledger + receiver
Start AIT
listener
Start
sender
Verify
readiness
Phase AIT verifies Failure behavior
allocate run identity, loopback ports, exercise configuration no child process starts if allocation fails
receiver first health/readiness of ledger, specialist, MCP server, or shadow sender remains stopped
interception listener binding and configured upstream traffic is not triggered until the boundary is usable
sender last sender process and target URL failed child is reported with its process receipt
stop pending-message disposition and owned process termination the run remains inspectable even after teardown

The run owns only the processes it launched. AIT records process receipts and uses them during cleanup; it does not sweep unrelated listeners or kill a process merely because it uses a familiar port.

What is real and what is controlled

The distinction matters when interpreting a result.

Layer Real in the default lab Controlled for repeatability
process boundary separate supervised operating-system processes all run on loopback
protocol implementation packaged a2a-sdk and mcp clients/servers fixed exercise operations and payloads
transport sockets for A2A/MCP HTTP; pipes for MCP stdio ephemeral local endpoints
interception engine the same direct interception path used outside the lab preconfigured break conditions
consumer decision the receiver process parses the representation Seam actually delivers deterministic, declared rules rather than an undeclared model
effect observation a separate controlled service records a receiver request or resulting state synthetic accounts, tenants, tasks, values, and effects

--fixture selects the older handwritten topology for compatibility tests. It is not the recommended learning mode. Optional reasoning exercises replace the deterministic decision with a configured model call and therefore require a different measurement contract.

Why the effect ledger is separate

When the specialist approves a payment, it does not just answer the gateway. It posts to the executor, and that separate process records the request as a row. The row is read out of band from the intercepted reply:

ait lab status LAB_ID
  effect ledger: {'kind': 'approval', 'value': 750.0, 'account': 'acct-demo', ...}

This is the distinction the evidence model rests on. A target's reply is narration; it can disagree with what the receiver actually did. A ledger row is a controlled observation authored outside the interception decision. It can show that the receiver attempted the fixture effect. It is not independent of the research owner, and it is not evidence that the same effect would occur in an external deployment.

When you read ait lab status, read the ledger, not the reply.

Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. A mobile evidence ladder showing the source required for each of five claim tiers. A mobile evidence ladder showing the source required for each of five claim tiers.

A run stops at the strongest tier it actually supports. Missing proof is reported as unproven rather than inferred from the target's narration.

One message, six operator states

The lab workflow captures, pauses, decides, delivers, observes, and compares each tested change. The lab workflow captures, pauses, decides, delivers, observes, and compares each tested change. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare workflow. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare workflow.

The traffic record changes state as the operator works it:

  1. captured: the listener received a real request, response, or event;
  2. pending: a break condition stopped delivery;
  3. decided: forward, edit, drop, replay, answer, or batch order was chosen;
  4. delivered: the downstream transport accepted the chosen bytes;
  5. observed: receiver state or a ledger receipt was collected;
  6. compared: attack and close control were scored in one session.

These states answer different questions. A message can be successfully edited but never delivered, delivered but rejected by the receiver, processed without changing behavior, or change behavior without an out-of-band effect source. The evidence view keeps those outcomes separate.

The consumers read, and they say what they are susceptible to

Six techniques work by putting words in front of something that reads them: poisoned tool descriptions, shadow tools, injected resource text, rewritten delegated instructions. For those to demonstrate anything, something in the lab has to actually read the text and act on it.

The lab's consumers do. They are not models; they are fixtures with four declared susceptibilities, and every decision they record states them:

  • they rank tools by what descriptions claim rather than by pinned identity;
  • they obey pass KEY=VALUE directives found in a tool description;
  • they treat retrieved resource text as instruction rather than as data;
  • they treat a delegated message or consumed artifact as instruction when it carries a directive, and let it stand in for the policy check.

These are deliberately vulnerable fixture policies selected to exercise the interception and evidence path. A positive result proves the payload is well-formed for this fixture and reaches its declared decision point. It does not establish prevalence in shipping clients or predict whether a model-mediated client will comply. Every affected primitive carries that bound in its limitation.

A lone status token stays data. The lab's own artifacts carry approved as a value, so only a sentence built around a directive verb is read as an instruction. Without that line, the control arm of a2a.artifact-content, which renames the artifact and leaves the text alone, would be indistinguishable from the attack.

The nine exercises

Exercise Protocol What it stands up Proof source
delegated-a2a-message A2A JSON-RPC gateway → specialist → MCP policy specialist state and approval ledger
a2a-identity-crossover A2A JSON-RPC tenant-partitioned specialist tenant-partitioned state and access ledger
a2a-agent-card-substitution A2A JSON-RPC primary and shadow specialists primary-versus-shadow state and routing ledger
a2a-callback-replay A2A JSON-RPC callback receiver callback receipt and idempotency ledger
asynchronous-artifact SSE streaming artifact revisions subscriber history and idempotency ledger
mcp-tool-result MCP HTTP MCP policy server client decision and effect ledger
mcp-tool-schema MCP HTTP MCP discovery surface call receipt and decision ledger
mcp-resource-poisoning MCP HTTP MCP resource surface resource receipt, later call, effect ledger
mcp-request-tampering MCP HTTP request-side MCP boundary server receipt and effect ledger

delegated-a2a-message also runs on --a2a-binding rest and --a2a-binding grpc, and mcp-tool-result on --mcp-transport stdio. Use them: the bindings differ in ways that have caught real defects, and stdio is the transport most real MCP deployments use.

Run artifacts and their authority

The lab run record uses the ait.intercept-lab/v3 contract. The exact storage path is workspace-dependent, so use ait lab status and the Cockpit links instead of assuming a directory. A complete run can include:

Artifact Produced by Use it for
exercise configuration AIT launcher protocol, binding, transport, synthetic identities
process receipts supervisor command ownership, readiness, exit, SDK/runtime version
traffic transcript interception session original bytes, operator decisions, delivered bytes, correlation
receiver state specialist, gateway, subscriber, or MCP server what the consumer parsed and retained
effect/access/routing ledger controlled observation service out-of-band record of the fixture action or destination
evidence result AIT evidence scorer attack/control attribution, achieved tier, limitations

No single artifact substitutes for the others. The transcript is authoritative about the exchange; receiver state is authoritative about local consumption; the ledger is authoritative about its controlled observation. The final claim is their intersection.

Reset, and why arms must share a session

ait lab reset returns the components to a known state between comparisons. It does not preserve the interception session. It starts a new one, so an attack run before a reset and a control run after it land in different sessions and cannot be scored against each other.

Run both arms in one session instead. Trigger the exercise twice:

ait lab trigger LAB_ID
ait intercept edit MESSAGE_ID --arm attack  --set /params/message/metadata/amount=750
ait lab trigger LAB_ID
ait intercept edit NEXT_ID    --arm control --set /params/message/metadata/amount=26
ait intercept evidence --primitive a2a.approval-amount

Reset when you want a clean ledger for a fresh comparison, not between the two arms of one.

What the lab cannot teach you

Being explicit about this is part of the lab's job.

  • Placement. Every component is on loopback and expects to be pointed at AIT. A real engagement's hard part is getting in path at all, and the lab skips it entirely.
  • TLS deployment. The default fixture uses loopback HTTP or pipes, so it does not exercise certificate issuance, private CAs, mTLS, DNS, or trust-store changes. AIT can use an explicit HTTPS listener, upstream CA material, and client certificates, but those are external placement decisions.
  • Model-mediated behaviour. The consumers are deterministic. Real instruction injection is probabilistic, and a single positive against a real client means much less than a positive here.
  • Scale. Exercises produce a handful of messages. A real session runs for hours across thousands, and nothing here rehearses that.

Two techniques have no exercise at all: mcp.elicitation-url-substitution and mcp.sampling-restriction-removal because both need server-initiated requests and the packaged MCP server runs stateless. They say so in their limitation, and ait doctor fails if a primitive ships without one.

Diagnose a lab that does not move

Work from transport to effect. Do not start by changing the payload again.

Symptom Inspect first Likely distinction
no pending message trigger status, break operation, selected protocol traffic never reached the intended interception rule
request pending forever operator decision and downstream readiness capture succeeded; delivery did not happen
delivered, receiver unchanged receiver parse/error state and exact field path wire mutation was accepted by AIT but not consumed
receiver changed, ledger empty ledger readiness and receiver-to-ledger call processing occurred without an out-of-band effect observation
attack and control both move arm labels, shared session, control design behavior is not isolated to the tested property
stale state after a run lab ID and reset receipt evidence from two runs may be mixed
ait lab status LAB_ID
ait intercept pending
ait intercept show MESSAGE_ID
ait intercept evidence --primitive PRIMITIVE_ID

Preserve the failed state long enough to inspect it. Reset is useful only after you have captured the reason the current run stopped short.

Beyond the lab

Three committed tests run interception against code the lab did not produce. They are listed weakest to strongest, because each closes something the one before it leaves open:

  • tests/test_against_an_independent_a2a_agent.py drives an A2A agent that imports nothing from ait, enforces its own ceiling, and calls the value payout rather than amount. Independent of the tool, but it runs the same Python a2a-sdk the lab does, so an assumption baked into that one library is invisible to it.
  • tests/test_against_a_third_party_server.py drives Anthropic's published @modelcontextprotocol/server-everything, a TypeScript package on npm that has never heard of AIT, over stdio. It asserts an operator's rewritten argument is what the server executes. Different authors and a different language, but MCP.
  • tests/test_against_a_cross_implementation_agent.py is the strongest: A2A, served by @a2a-js/sdk. A different codebase in a different ecosystem, naming its own field (settlement) and setting its own ceiling. An edit that lands here lands because it is correct on the wire, not because both sides agreed on a spelling. Writing it found two things the Python fixtures could not: a taskId the server has not issued is rejected outright, and this SDK carries message parts in a protobuf-derived shape internally.

Read those when you want to know what the tool does against something that was not built around it. The first runs in the default suite; the other two need node and npm and are run by third-party-interception.yml, which fails if either skips.