Skip to content

Replicable, reproducible, robust research

Agent systems are stochastic, stateful, and assembled from components with different owners. One successful screenshot cannot establish a security property. AIT therefore treats every intervention as an experiment with a named boundary, an observable receiver, a close control, and records another reviewer can inspect without trusting the operator's recollection.

R/R/R is the platform's research contract:

  • Replicable: another operator can rebuild the design and run the same intervention against the stated target.
  • Reproducible: another reviewer can recompute the reported result from the deposited records without contacting the model or target again.
  • Robust: the result is tested across the variation the claim says it survives, rather than generalized from one configuration.

These are separate properties. A perfectly reproducible synthetic fixture can still lack external validity. A live external run can still be irreproducible if its request, response, provider, target state, and oracle were not retained.

One closed research loop

A closed research loop captures traffic, pauses at a named operation, makes one decision, delivers it, observes the receiver and a controlled out-of-band source, and compares the result with a close control. A closed research loop captures traffic, pauses at a named operation, makes one decision, delivers it, observes the receiver and a controlled out-of-band source, and compares the result with a close control. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare research loop. A mobile view of the six-step capture, pause, decide, deliver, observe, and compare research loop.

The mutation is the intervention, not the finding. A run becomes evidence only after delivery, receiver state, an out-of-band observation, and the close control have been reconciled.

Every experiment follows the same logical sequence:

  1. Capture an exchange produced by the system under test.
  2. Pause at the operation declared by the hypothesis.
  3. Decide whether to preserve it or apply one tested change.
  4. Deliver the selected representation to the selected receiver.
  5. Observe the transcript, receiver state, and preregistered effect source.
  6. Compare the intervention with its close control.

The implementation may be a local protocol fixture, a configured model call, or an authorized external service. The logical contract does not change.

Start with a falsifiable claim

A useful hypothesis names five things before execution:

Field Question it answers Example shape
boundary Which exact request, response, event, or wrapper is controlled? A2A message/send response
intervention What single property differs in the attack arm? delegated verdict value
receiver Which program consumes the delivered bytes? coordinator process
effect What receiver or external state should change? attempted merge or settlement row
close control What matched change should leave the tested property absent? same reply shape with a negative verdict

“The attack works” is not a hypothesis. “Replacing a delegated verdict changes the coordinator's attempted action relative to a matched negative verdict” is testable because each noun resolves to a record.

Pair the intervention with a close control

Attack and close-control runs start from the same seed, receiver, operation, and observation window, then differ only by the preregistered tested change. Attack and close-control runs start from the same seed, receiver, operation, and observation window, then differ only by the preregistered tested change. A mobile paired-trial view showing the shared seed, attack run, close-control run, and comparison result. A mobile paired-trial view showing the shared seed, attack run, close-control run, and comparison result.

The control need not be an untouched message. A matched harmless change is often stronger because it separates the content of the intervention from the mere fact that the exchange was intercepted and edited.

A pair is interpretable only when the design holds constant what the claim needs held constant. Typical requirements are:

  • the same target identity and version;
  • the same seed or reset state;
  • the same receiver and downstream services;
  • the same operation and interception placement;
  • the same observation window;
  • the same prompt and model configuration when a model is in the loop;
  • one preregistered difference between attack and control.

Stochastic scheduling and provider execution are observations, not automatically design defects. They must be recorded and included in the claim boundary. An additional designed difference, such as changing both the reply verdict and the target repository, confounds the pair.

Keep delivery, processing, behaviour, and effect separate

Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. Five evidence tiers separate nothing delivered, changed bytes delivered, receiver processing, behaviour changed relative to a close control, and an effect observed outside the interception path. A mobile evidence ladder showing the source required for each of five claim tiers. A mobile evidence ladder showing the source required for each of five claim tiers.

Each tier adds a claim and therefore needs a source. One source may support several tiers, but no tier is inferred merely because a lower one was observed.

The evidence tiers answer different questions:

Tier Question Minimum supporting source
nothing_delivered Did the selected representation cross the boundary? decision and delivery record show it did not
mutation_delivered Did the changed bytes cross? chained before-and-after transcript
receiver_processed Did the receiver process that delivery? correlated receiver response, state, or controlled endpoint
behavior_changed Did the outcome differ from the close control? paired receiver or oracle observations
external_effect_observed Did a source outside the interception decision record the downstream effect? target API, service state, callback, datastore, file transition, or controlled action ledger

The receiver's explanation is not allowed to stand in for its behaviour. A controlled observation service is out of band with respect to the interception decision, but it is not magically independent of the research owner. Its implementation and reset behaviour remain part of the fixture's validity.

Distinguish real mechanisms from controlled variables

“Real” is not a binary label for an experiment. Record it layer by layer.

Layer Questions to record
transport Was this a production protocol implementation, a protocol-compatible fixture, or a handwritten stub?
process Were sender, receiver, interceptor, and observation source separate programs or one in-memory test?
model Was a live provider or local model called? What requested model and serving path answered?
task Did the task come from an external workflow or was it authored for the experiment?
policy Who wrote the policy and where was it enforced?
effect Did the target's own service state change, or did a controlled ledger record a synthetic effect?
placement Was the sender explicitly pointed at the boundary, or was existing deployment routing changed under authorization?

AIT's default lab uses production interception code, SDK-backed messages, separate processes, sockets, and pipes. Its roles, tasks, policies, values, and effects are controlled. External fixtures should replace those authored layers one at a time and state what remains controlled.

Synthetic components are acceptable for instrument validation, negative controls, fault injection, and repeatable operator training. They cannot, by themselves, support claims about prevalence, a third-party product, or a model family. Those claims need a target owned outside the fixture and an oracle tied to that target's own state.

Validate the instrument before scoring the hypothesis

An experiment should have an unscored validation phase. Its purpose is to show that the apparatus can produce and reject the observations the confirmatory test depends on.

Validate at least:

  1. the positive gate can produce the allowed effect;
  2. the negative or close-control gate does not produce it;
  3. the intended message is the one intercepted;
  4. the attack and control payloads differ only where declared;
  5. the receiver state and external observation agree on their own domains;
  6. reset removes state from the previous trial;
  7. a missing or malformed observation remains missing rather than becoming a success or failure by coercion;
  8. the run records the code version and whether the worktree was clean at launch.

If a gate fails because the instrument measured the wrong operation, the batch is invalid. Preserve the record, explain the defect, revise the protocol before another validation batch, and do not recycle those trials as confirmatory data.

Account for missing data without inventing outcomes

An unreadable model reply, exhausted retry, unavailable service, or missing ledger event is data about the run. It is not automatically evidence that the attack failed or the defence held.

The protocol must decide before execution whether a condition is:

  • a scored negative outcome;
  • an excluded infrastructure failure;
  • a dropped pair because one arm cannot be interpreted;
  • a failed validation gate that voids the batch.

Exclusions require an independently checkable reason, such as a recorded service error. Refusal, no tool call, and server-side rejection are usually behavioural outcomes and should remain in the denominator when the hypothesis is about whether the action occurred.

Preserve provenance at the time it matters

Every scored record should identify the environment at launch, not reconstruct it after the last trial:

  • commit or immutable source identifier;
  • clean, dirty, or unknown worktree state;
  • experiment and protocol revision;
  • package and runtime versions;
  • target, model, and serving-provider identity;
  • container image or artifact digests where applicable;
  • exact request, intervention, close control, and observation configuration;
  • per-trial raw reply or service event needed to rescore the outcome;
  • transcript and oracle references with hashes.

A run that names a commit while executing later uncommitted code has ambiguous provenance. A record that omits cleanliness is unknown, not clean.

Make publication a derivation, not a second dataset

Published counts, tables, and figures should be generated from one designated canonical record set. A checker should match the dimensions that identify a claim, such as target, primitive, protocol revision, arm, and sample size. A historical run may show spread, but it must not silently vouch for a current number that the canonical run contradicts.

Useful publication checks include:

  • every outward count resolves to a current clean record;
  • no missing value is converted into a numeric outcome;
  • trial witnesses remain available per arm and trial rather than one example overwritten for a whole cell;
  • source, generated page, figure, and briefing agree;
  • stale historical records cannot satisfy a current claim;
  • a dirty or unknown run cannot become canonical without an explicit reason;
  • the checker has a positive control that demonstrates it can report a real contradiction.

The public documentation is a separate boundary. Unpublished papers, run records, preregistrations, and submission material do not enter it by directory discovery. The site is staged from an exact reviewed allowlist.

Choose robustness dimensions from the claim

Robustness is not “run more.” It is variation aligned with the proposed generalization.

Proposed claim Relevant variation
transport-independent repeat across the specific bindings named in the claim
model-independent repeat across model families and serving paths, not only sizes from one family
deployment-relevant use a target and effect source not maintained by the instrument project
stable over time repeat clean runs at distinct times and retain upstream provenance
payload-family effect sweep preregistered phrasings with one negative control and retain per-variant results
protocol-level mechanism demonstrate the result without relying on persuasion or a model-specific response

When outcomes vary, report the variation. Do not promote a midpoint from one run into a property of the target. Stable failure, stable success, and unstable rates are different findings.

Component responsibilities

  • Seam records complete delivered frames, rule decisions, correlation, and a hash-chained transcript.
  • meshmapper produces deterministic communication graphs and candidate trust paths. Those hypotheses remain unproven until a test crosses the selected boundary.
  • Assay attaches baselines and controlled or external oracles, compares paired observations, and packages bounded findings.
  • AIT coordinates the run, session, process ownership, operator decision, and artifact references without letting coordination stand in for evidence.

Final review questions

Before calling a result ready, answer these in plain language:

  1. What exact hop did the researcher control?
  2. How did the message reach that hop?
  3. Which bytes did the selected receiver actually get?
  4. Which source shows what the receiver processed?
  5. Which source records the downstream effect?
  6. What differs between attack and close control?
  7. What missing data was dropped, excluded, or scored negative, and why?
  8. Which layers are controlled, and which are owned outside the fixture?
  9. Can another reviewer recompute the reported value from retained records?
  10. What is the strongest sentence the evidence supports, and what stronger sentence must remain unwritten?

If any answer depends on “the model said so” or “the UI looked right,” the research loop is not closed.