Skip to content

Controlled research lab

The lab is a process-backed fixture for learning and testing one agent boundary at a time. Nine public exercises use deterministic receivers. Optional private research exercises add configured model calls and repeated paired trials.

The controlled lab starts a gateway sender, AIT boundary, specialist receiver, MCP policy source, effect ledger, and private run directory as separate responsibilities. The controlled lab starts a gateway sender, AIT boundary, specialist receiver, MCP policy source, effect ledger, and private run directory as separate responsibilities. A mobile view of the lab's gateway, interception boundary, specialist, MCP policy source, controlled effect ledger, and run directory. A mobile view of the lab's gateway, interception boundary, specialist, MCP policy source, controlled effect ledger, and run directory.

SDK-backed protocol messages cross the same checked-in Seam execution path used by external-target workflows. Identities, tasks, policy, and effects are synthetic so the state can be reset and compared.

ait doctor
ait lab start

The default launch starts separate local processes, places the same interception engine used in assessments between them, opens the Cockpit, and waits for you to decide what reaches the receiver. Docker, a provider key, a pre-existing workspace, and external network access are not required.

The lab architecture names which mechanisms are shared with external assessments and which fixture inputs remain controlled.

Pick a route through the lab

I have 20 minutes Run delegated A2A message, change one value, forward a close control, and read the ledger. Run the 20-minute route
I want the curriculum Build from observation to identity, routing, MCP, asynchronous delivery, chaining, and handoff. Follow the curriculum
I need one technique Choose an exercise by protocol boundary and use its exact pause, change, control, and proof recipe. Choose a deterministic exercise
I am studying model behavior Start with the evidence boundary, then design paired trials against your own configured consumer. Read the evidence guide

What counts as success

Changing a message is not enough. Each deterministic exercise exposes a proof source outside the intercepted reply and gives you a close control. Read the result as a ladder: stop at the highest level the evidence actually supports.

Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. A mobile evidence ladder showing the source required for each of five claim tiers. A mobile evidence ladder showing the source required for each of five claim tiers.

The receiver's reply is narration. Receiver state shows what it consumed. The effect, routing, access, callback, or idempotency ledger shows what happened in the controlled environment. The attack/control pair tells you whether the tested property, not merely any edit, moved the outcome.

Deterministic exercise matrix

Exercise Boundary you control Attack edit Close control Out-of-band proof
Delegated A2A message A2A request approval amount message text specialist state + approval ledger
Task and tenant crossover A2A request tenant or task identity descriptive label tenant state + access ledger
Agent Card substitution A2A discovery response interface URL card description primary/shadow state + routing ledger
Callback substitution and replay A2A callback configuration destination or replay count one unchanged delivery callback + idempotency ledger
Asynchronous artifact A2A stream events drop, reorder, duplicate release once in order subscriber history + idempotency ledger
MCP tool result MCP response structured risk descriptive text client decision + effect ledger
MCP tool schema MCP discovery response input maximum tool description server call receipt + decision ledger
MCP request tampering MCP request resource URI or arguments _meta or unused note server receipt + effect ledger
MCP resource poisoning MCP resource response policy-like content benign marker resource receipt + later call + effect ledger

Every detailed guide answers six questions: what is running, which message to pause, what to change, what not to change in the control, where to read the result, and what the result does not establish.

Model-backed experiments

Model-backed experiments add a model call and paired trials. They are experiments in consumer behavior, not better versions of the deterministic exercises.

Use them after the deterministic path because the interpretation changes:

  • deterministic exercises prove that a mutation is well-formed, reaches the decision point, and can move the controlled fixture;
  • model-backed exercises measure whether a particular configured consumer follows the delivered content under repeated attack and control trials;
  • neither result generalizes to an unrelated deployment without a new measurement.

The public lab stops at this design boundary. Project-specific model experiments and their measurements remain in the unpublished research corpus.

Before you leave the lab

You should be able to explain, from the evidence rather than the interface:

  1. which process produced the original message;
  2. exactly which boundary AIT controlled;
  3. what the receiver consumed after release;
  4. which controlled out-of-band observation recorded the effect;
  5. how the close control differs from the attack;
  6. what remains unproven.

Then read Taking it to a real target. Placement, TLS, target ownership, evidence sources, and cleanup all become engagement decisions outside the local lab.