Skip to content

Multimodal Injection to Computer Action

Methodology, not an executable workflow. This page defines a trust boundary, close control, and evidence requirement. The operator must implement it against a specific authorized target; no command here claims to execute the full chain.

R/R/R boundary

Part Status
Media and control inputs deposited bytes with retained digests
Model and computer-use runtime must be the exact authorized target to support an external claim; a JSONL adapter is controlled instrumentation
Effect evidence application state, callback, or action record outside the model response
Claim limit a synthetic desktop proves the experiment can be inspected and reset, not that a shipping computer-use agent is vulnerable

Deposit the image or document and its semantically close control before target contact. Invoke the agent, then pass the decision to computer_jsonl. Record media digest, evaluator version, decision span, accessibility/window identity, pointer or keyboard actions, and the terminal controlled callback or action log.

Use a fresh process/context and synthetic application state. Screenshots and response evaluators are behavioral evidence, not impact. Ablate the media interpretation and computer-action stages. Mitigation candidates include document-layer stripping, cross-modal policy checks, action confirmation, and destination constraints.