Authoring your own attack primitives¶
The shipped catalogue is pinned to the lab. mcp.tool-argument-escalation
rewrites /params/arguments/account, because that is what the lab's policy
server calls the field. Point it at a real payments tool whose argument is
beneficiary_iban and it fails with JSON Pointer does not resolve.
That is expected, and there are two ways past it. Override a shipped primitive for a one-off, or write your own when you will run it more than once.
Override a shipped primitive¶
--set POINTER=VALUE on ait intercept attack, repeatable. A pointer the
primitive already patches replaces that patch's value and keeps its operation.
Any other pointer becomes an extra edit on the same arm.
ait intercept attack mcp.tool-argument-escalation --target MSG \
--set /params/arguments/beneficiary_iban='"DE99OPERATOR"'
ait intercept attack a2a.approval-amount --target MSG --set /params/message/metadata/amount=90000
An override changes what is written, never which arm is running. The arm keeps the hypothesis, observation, control rationale and limitation the catalogue declared, so the finding still says what the technique was and what it cannot show.
Write your own¶
Drop a JSON file in <workspace>/.ait/primitives/. Anything there is added to
the catalogue. It does not replace the shipped set, and a duplicate ID is
refused rather than silently shadowing.
{
"primitives": [
{
"id": "acme.approval-ceiling",
"title": "Push an approval past the client's own ceiling",
"boundary": "Approval parameters",
"family": "content",
"selector": {
"protocol": "a2a",
"direction": "request",
"requires_paths": ["/params/message/metadata/amount"]
},
"hypothesis": "The specialist trusts the delegated amount without re-checking policy.",
"observation": "The approval ledger records an amount above the documented ceiling.",
"control_rationale": "Send an amount just under the ceiling: the same field, a compliant value.",
"limitation": "Shows the specialist acted on the delegated amount. It does not show a human reviewer would have missed it.",
"attack_patches": [
{"operation": "replace", "path": "/params/message/metadata/amount", "value": 9000}
],
"control_patches": [
{"operation": "replace", "path": "/params/message/metadata/amount", "value": 24}
]
}
]
}
Then it behaves like any shipped primitive:
ait intercept attack --target MSG
ait intercept attack acme.approval-ceiling --target MSG
ait intercept attack acme.approval-ceiling --target NEXT_MSG --arm control
ait intercept evidence --primitive acme.approval-ceiling
The fields that carry the weight¶
selector decides which paused messages the primitive offers itself for.
requires_paths is what keeps it from appearing against traffic it cannot
patch: get it wrong and the primitive lists as applicable and then fails.
control_patches is not optional in practice. A control is a near-miss: the
same field, the same visibility to the receiver, a value that does not cross the
boundary. Without one, ait intercept evidence reports no_control and the
finding cannot be attributed to anything. Attribution is computed by comparing
what the receiver did on each arm, so the control has to be something the
receiver will actually process.
limitation is what stops the finding overclaiming, and it is the field
reviewers read first. Say what the result does not establish. "Shows the agent
consumed the altered result. Whether a human reviewer would have caught it is
out of scope" is the shape.
Paths and the A2A bindings¶
You do not write a path per binding. The three A2A bindings differ in exactly one
structural way: JSON-RPC wraps the payload in params (requests) or result
(responses), REST and gRPC carry it at the document root:
Paths are interpreted as payload-relative. A declared params or result prefix
is stripped and the prefix the message actually uses is applied, so
/params/message/metadata/amount and /message/metadata/amount both work on all
three. Write whichever reads better.
This matters more than it sounds: before it, seven of the nine shipped A2A
primitives could not address REST or gRPC traffic, and intercept attack --target
MSG reported "none apply" on those targets rather than saying why.
Envelope paths (target: "envelope") address transport headers, which have no
wrapper, and are never rebased.
operation is replace, add or delete, with RFC 6901/6902 semantics:
replace requires the path to exist, add creates it, add at an array index
inserts, and - appends. Use add for a field the message does not already
carry: replace will refuse, which is deliberate, because a replace that
silently creates cannot catch a mistyped pointer.
Attacks that withhold or repeat instead of editing¶
Not every boundary is a content edit. The Lifecycle boundary is about sequence
-- what happens when an event never arrives, or arrives twice -- and a primitive
expresses that with action rather than patches:
{
"id": "a2a.stream-completion-drop",
"boundary": "Lifecycle",
"family": "lifecycle",
"action": "drop",
"control_action": "forward_original",
"attack_patches": [],
"control_patches": []
}
action is forward_modified (the default), drop, or replay with copies.
control_action is what the control arm does, and an ordering primitive must
set it. Without it the arm inherits action, so a drop primitive drops on both
arms and its control compares nothing to nothing. For an ordering attack the
field guide's close control is "deliver the original sequence once", which is
forward_original: attribution then compares the receiver's behaviour against
the run where nothing was withheld or repeated.
There is no diff to assert on these. changed_paths is empty by construction --
withholding an event changes no field, and repeating one changes no field twice
-- so tests/test_primitives_run_in_the_lab.py judges them on the decision
recorded, the copies delivered, and the exercise reaching a terminal state.
Both terminal states are legitimate for an ordering arm. A consumer that completes without ever receiving its terminal event has told you something, and so has one that hangs; that is the finding, not a broken run.
Anchoring to a lab exercise¶
lab_exercise names an exercise the primitive can be practised against.
tests/test_primitives_run_in_the_lab.py then runs both arms against it and
asserts the patch lands and the exercise still completes.
If there is no exercise for your boundary, leave it out and say so in
limitation. Shipping an unpractisable primitive silently is worse than
shipping one that names its own gap.