Multimodal Injection to Computer Action¶
Methodology, not an executable workflow. This page defines a trust boundary, close control, and evidence requirement. The operator must implement it against a specific authorized target; no command here claims to execute the full chain.
R/R/R boundary¶
| Part | Status |
|---|---|
| Media and control inputs | deposited bytes with retained digests |
| Model and computer-use runtime | must be the exact authorized target to support an external claim; a JSONL adapter is controlled instrumentation |
| Effect evidence | application state, callback, or action record outside the model response |
| Claim limit | a synthetic desktop proves the experiment can be inspected and reset, not that a shipping computer-use agent is vulnerable |
Deposit the image or document and its semantically close control before target contact. Invoke the agent, then pass the decision to computer_jsonl. Record media digest, evaluator version, decision span, accessibility/window identity, pointer or keyboard actions, and the terminal controlled callback or action log.
Use a fresh process/context and synthetic application state. Screenshots and response evaluators are behavioral evidence, not impact. Ablate the media interpretation and computer-action stages. Mitigation candidates include document-layer stripping, cross-modal policy checks, action confirmation, and destination constraints.