Skip to content

A research platform for the boundary between agents

AIT lets a researcher place an explicit relay or wrapper on an A2A, MCP, streaming, stdio, gRPC, or framework-mediated exchange. It keeps three questions separate: what crossed the boundary, what the receiver did, and whether a close control did the same thing.

Start with the system architecture, the controlled lab design, or the R/R/R research method.

Traffic moves from a sender through an explicit interception boundary to a receiver, while a separate evidence plane records the transcript, receiver state, a controlled out-of-band observation, and a paired comparison. Traffic moves from a sender through an explicit interception boundary to a receiver, while a separate evidence plane records the transcript, receiver state, a controlled out-of-band observation, and a paired comparison. A mobile view of the sender, controlled boundary, receiver, and separate evidence sources. A mobile view of the sender, controlled boundary, receiver, and separate evidence sources.

AIT is the middle plane, not the whole experiment. It can prove what it delivered. Receiver state, a controlled observation source, and a close control are what allow a stronger claim.

Research platform status

AIT is developed as a research platform, not distributed as a packaged public download. This site documents its architecture, controlled lab, and source-based workflows for authorized research and project collaborators.

Architecture before operation

Every AIT experiment is designed around three questions.

Question Owning evidence Why it stays separate
What was sent? Seam's chained before-and-after transcript A changed editor view is not proof that those bytes were delivered.
What did the receiver do? receiver state, correlated processing, or a controlled endpoint The receiver's explanation may disagree with its behaviour.
Did the tested change cause it? attack and close-control observations from the same design One successful run cannot distinguish the intervention from normal behaviour.

The system architecture follows one message across those planes, names component ownership, and documents transport placement, credential handling, storage, supervision, and architectural limits.

What is real, and what is controlled

The default lab is intentionally mixed. It uses SDK-backed protocol messages and the same checked-in Seam execution path as external-target workflows inside a controlled, resettable fixture.

Layer Real mechanism Controlled variable
process and transport separate operating-system processes, sockets, pipes, A2A and MCP SDKs loopback placement and fixed endpoints
interception the same Seam engine, queue, rules, decisions, and transcripts used with external targets preselected operation and disposable run state
receiver actual fixture code parses the bytes that arrive and writes its own state synthetic identities, tasks, policy, and values
observation a separate service or state source records the downstream action resettable effects chosen for the exercise
model-mediated experiment live provider call when that experiment explicitly requires one pinned prompt, task, arm construction, and scoring contract

This supports repeatable claims inside the fixture. It does not turn a controlled target into external validity. The lab architecture states the boundary precisely, and real-target placement shows what an external experiment must supply.

The research loop

A closed research loop captures traffic, pauses at a named operation, makes one decision, delivers it, observes the receiver and oracle, and compares the result with a close control. A closed research loop captures traffic, pauses at a named operation, makes one decision, delivers it, observes the receiver and oracle, and compares the result with a close control. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare research loop. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare research loop.

The mutation is the intervention, not the finding. Each claim stops at the strongest observation the run actually produced.

The same loop applies to a local protocol fixture, a live model experiment, and an authorized external target:

  1. capture an exchange from the system under test;
  2. pause at the operation named in the hypothesis;
  3. change one tested property or deliver its close control;
  4. record exactly what crossed the boundary;
  5. observe receiver state and the preregistered oracle;
  6. compare the paired outcomes and report the limit.

Read the platform in this order

  1. System architectureHow the sender, boundary, receiver, evidence plane, and four platform components fit together.
  2. Controlled lab designWhich mechanisms are production paths, which identities and effects are synthetic, and what the fixture can establish.
  3. R/R/R methodologyHow hypotheses, controls, replication, raw records, and bounded claims are constructed.
  4. Authorized target placementWhat changes when the sender, receiver, TLS, routing, task, and oracle come from an external system.

Authorized research only

Use AIT only on systems you own or are explicitly authorized to assess. The controlled lab binds locally by default. Any external target requires deliberate placement, documented scope, a retention plan, and its own evidence design.