Skip to content

System architecture

AIT is a local research control plane for experiments at agent interaction boundaries. It does not need to run inside either agent. A researcher places an explicit relay or wrapper on one selected hop, observes complete protocol messages, decides what crosses that boundary, and evaluates the receiver using evidence collected outside the interception decision itself.

AIT separates sender and receiver traffic, the controlled interception boundary, and an evidence plane with receiver state, a controlled out-of-band observation, and a paired comparison. AIT separates sender and receiver traffic, the controlled interception boundary, and an evidence plane with receiver state, a controlled out-of-band observation, and a paired comparison. A mobile view of the sender, controlled boundary, receiver, and separate sources for delivery, receiver state, out-of-band observation, and comparison. A mobile view of the sender, controlled boundary, receiver, and separate sources for delivery, receiver state, out-of-band observation, and comparison.

The central architectural rule is separation. The boundary records delivery. Receiver state and a controlled out-of-band source establish what followed. A close control determines whether the tested change explains the difference.

The three planes

Plane What crosses it What it can establish
Traffic and target A2A, MCP, SSE, WebSocket, stdio, gRPC, or framework-mediated messages between the actual sender and receiver selected for the experiment What the target accepted, rejected, returned, or recorded in its own state
Interception and decision Complete decoded frames, operator decisions, rule evaluations, and the exact bytes released downstream What was observed at the boundary and what was delivered
Evidence and comparison Transcript references, receiver observations, external oracle readings, baselines, and close-control results Whether downstream behaviour or an external effect differed between arms

These planes remain separate even when AIT starts all of them for a controlled lab. A transcript cannot promote itself into an impact claim, and a target's own explanation cannot substitute for an external observation.

Runtime ownership

AIT owns experiment coordination, Seam owns protocol delivery evidence, meshmapper owns deterministic trust-path mapping, and Assay owns controlled comparison and adjudication. AIT owns experiment coordination, Seam owns protocol delivery evidence, meshmapper owns deterministic trust-path mapping, and Assay owns controlled comparison and adjudication. A mobile view of AIT, Seam, meshmapper, and Assay, with the evidence claim each component can and cannot establish. A mobile view of AIT, Seam, meshmapper, and Assay, with the evidence claim each component can and cannot establish.

No component is allowed to answer every research question. The split keeps candidate paths, delivered mutations, observed effects, and comparative findings from collapsing into one claim.

Component Owns Does not establish by itself
AIT workspaces, connections, sessions, supervised processes, labs, shadows, jobs, operator decisions, the local API, and the Cockpit that a delivered mutation affected the target
Seam transport termination or wrapping, complete-frame decoding, break conditions, transform rules, mutation, delivery, correlation, and the chained transcript that the receiver used a changed field or produced an external effect
meshmapper deterministic communication graphs, trust-boundary candidates, and paths worth testing that a mapped path is exploitable
Assay baselines, external oracles, paired evidence, attribution, statistical comparisons, findings, and reports where the interception boundary should be placed

General infrastructure testing stays behind adapter boundaries. AIT can invoke or import established security tools; it does not reimplement every scanner.

One message through the system

A message moves from a source into Seam, where it is decoded and held, then to the researcher for a decision, and finally to the selected receiver. A comparison run follows the same receiver path. A message moves from a source into Seam, where it is decoded and held, then to the researcher for a decision, and finally to the selected receiver. A comparison run follows the same receiver path. A mobile view of a complete message moving through emit, decode, hold, decide, release, and receiver-processing states, followed by a comparison run. A mobile view of a complete message moving through emit, decode, hold, decide, release, and receiver-processing states, followed by a comparison run.

A held item is a complete protocol message rather than a packet fragment. The operator can inspect the decoded structure while the sender's connection remains open.

The direct interception path is:

  1. Place. The sender is explicitly configured to use a Seam listener or wrapper. Traffic that is not pointed at AIT is not observed.
  2. Decode. Seam reassembles the transport unit and exposes a typed message with protocol, direction, operation, body, envelope, and correlation fields.
  3. Evaluate. Enabled transform rules run first. Break conditions then decide whether the message continues or enters the bounded pending queue.
  4. Decide. The researcher forwards the presented message, bypasses a rule and forwards the sender's wire value, edits it, drops it, replays it, or answers at the boundary.
  5. Deliver. Seam re-encodes the selected representation for the same transport and sends it to the configured upstream or back to the caller.
  6. Record. The original frame, presented frame, delivered frame, rule evidence, changed paths, decision, correlation, and response are appended to the transcript.

The pending queue is bounded. A researcher can also set a timeout and an explicit timeout decision so unattended interception does not become an unbounded outage.

Transport placement

AIT is explicit infrastructure, not a transparent network interceptor.

Binding Placement Boundary AIT controls
A2A JSON-RPC and REST sender points to an HTTP or HTTPS Seam listener with an explicit upstream request, response, task, message, and artifact envelopes
A2A gRPC sender points to the protocol-specific gRPC bridge supported A2A service messages; this is not general-purpose gRPC interception
MCP Streamable HTTP MCP client points to a Seam HTTP listener initialization, sessions, tools, resources, prompts, notifications, and SSE responses
MCP stdio the client launches the wrapper command printed by AIT framed stdin and stdout traffic without a shell
SSE and WebSocket client points to the explicit relay complete events or messages, including ordering and replay decisions
framework-mediated calls an adapter selects the first-party driver and its typed artifacts the framework surface declared by that driver

For network listeners, the data plane can bind to loopback or to the exact routable address the researcher supplies. The control API remains on loopback. The upstream may use a public CA, a private CA, or mTLS. A listener may also present a supplied certificate when the sender requires HTTPS.

AIT does not generate a trusted interception certificate, install a trust root, alter DNS, perform ARP redirection, or silently repoint a client. If the sender cannot be configured to use the listener or wrapper, AIT cannot occupy that hop.

Data and credential handling

Raw traffic is evidence and may contain credentials. AIT therefore does not claim that captured credentials disappear before the editor or transcript. Instead:

  • control endpoints bind locally and use per-session authorization;
  • run artifacts are private by default and written with restrictive modes;
  • immutable message records retain the before and after values needed to audit a decision;
  • the UI can mask sensitive paths while retaining the underlying evidence;
  • export is redacted by default, while --raw is explicit and labelled raw.

This is a research tradeoff, not a general logging recommendation. The researcher owns the retention and disclosure plan before placing the relay.

Evidence architecture

Execution state and evidence state are independent. A process can complete successfully while proving nothing about the hypothesis. The evidence model therefore records a highest supported tier rather than inferring later tiers:

Tier Required source
nothing_delivered decision and delivery record show that the selected frame did not cross the boundary
mutation_delivered chained transcript shows the altered representation delivered downstream
receiver_processed correlated response, receiver state, or a controlled endpoint shows downstream processing
behavior_changed attack and close-control observations differ in the preregistered direction
external_effect_observed a target-appropriate system outside the interception path records the downstream effect

An external oracle can be a controlled datastore row, file transition, callback receipt, service API reading, or another effect source the interception path does not write. Its baseline is recorded before the intervention. A baseline that is already positive voids that probe rather than manufacturing a finding.

See Reading evidence and R/R/R methodology for the claim rules above this architecture.

Controlled lab versus external target

The default lab is a process-backed protocol fixture. It uses SDK-backed A2A and MCP messages, real sockets or pipes, the same checked-in Seam execution path as external-target workflows, and separate receiver and observation processes. It controls identities, tasks, policies, values, and effects so trials can be reset and repeated.

That combination is useful for validating transport handling, rule placement, causal attribution inside the fixture, and the evidence pipeline. It is not external validity. It does not show that an unrelated deployment, model, or policy will respond the same way. External experiments must bring their own target, task, baseline, close control, and oracle.

The lab architecture names every synthetic and real layer. Taking AIT to a real target covers authorized placement and evidence design outside the fixture.

Storage architecture

workspace/
└── .ait/
    ├── workspace.json
    ├── platform.db                 # SQLite/WAL domain and job state
    ├── objects/                    # SHA-256 content-addressed blobs
    ├── media/                      # manifests; bytes remain in object storage
    ├── runs/<uuidv7>/              # one exclusive execution directory
    ├── jobs/                       # bounded stdout, stderr, and process state
    ├── events/                     # resumable event chunks
    └── server.token                # random local API token

Run directories do not share writable artifacts. Raw traffic remains in immutable JSONL or transcript artifacts. SQLite traffic indexes contain offsets and correlation keys rather than payload copies. Experiment tables may use Parquet and DuckDB when the research dependencies are installed.

Supervision, API, and Cockpit

The process supervisor creates process groups, records bounded logs, performs cooperative cancellation, escalates when required, and recovers durable terminal state. Stdio cleanup owns the wrapper, target, and descendants it started. It does not kill unrelated processes merely because they use a known port.

FastAPI exposes the local /api/v1 control surface and resumable events. The React Cockpit is generated against the same contract. Local browser bootstrap uses a short-lived, single-use token carried in the URL fragment, exchanges it for an HttpOnly same-origin cookie, and removes it from browser history.

Target handoff copies protocol placement, break conditions, and editor preferences into a fresh connection. It does not copy controlled lab identities, fixture state, observation rows, or credentials. Preparing, testing, and starting a target connection remain separate operator actions.

Architectural limits

  • AIT observes only traffic deliberately routed through its listener or wrapper.
  • A2A gRPC support is scoped to the A2A service, not arbitrary protobuf APIs.
  • A modified message proves delivery only after the transcript records it.
  • A correlated target response can support processing, but not an external effect by itself.
  • A controlled lab result is bounded to that fixture until an external target reproduces it.
  • Mapping a trust path does not show that the path is exploitable.
  • The platform does not choose an attack or authorize an assessment.

For the state machines behind these paths, continue to the domain model. For the operator-facing control surface, see the direct interception guide.