Multimodal and Provider Playbook¶
First-party packs: multimodal.input-trust and model.provider-boundary.
Deposit every input before target contact. A manifest records digest, MIME type, structure, transformations, layers, semantic labels, control relationship, provenance and CAS references.
R/R/R boundary¶
| Part | Status |
|---|---|
| Input bytes and transformations | deposited and digest-addressed |
| Model or provider path | exact authorized endpoint and model identity for an external result; local compatible endpoints are instrument checks |
| Effect evidence | target-owned state or a controlled action source outside response scoring |
| Claim limit | refusal, text matching, latency, and token use are behavior metrics; none alone establishes external impact |
Multimodal variants¶
multimodal.image-instruction-injectionmultimodal.audio-instruction-injectionmultimodal.document-hidden-layer-injectionmultimodal.cross-modal-conflict
Each needs a byte-deterministic attack input and close control. Bind them separately through media_manifest_refs and control_media_manifest_refs. The OpenAI-compatible runtime resolves only deposited ait-media://N references. Other providers implement provider_jsonl and receive the case-appropriate deposited manifest/input references.
Provider variants¶
model.policy-boundary-jailbreakmodel.provider-system-prompt-extractionmodel.privacy-canary-extractionmodel.resource-amplification
Use synthetic policies, prompts and canaries. Do not treat arbitrary text matches as real-secret extraction. Record endpoint, model identity, request/response digests, latency, token use, retries and evaluator version.
openai_http supports Responses and Chat Completions, bounded JSON/SSE, request/response limits, input/output token limits, budgets and zero retries by default. Credentials are env:NAME references. Explicit retries record every attempt.
Evaluators measure refusal, leakage, policy violation, latency, tokens and amplification. These are behavioral metrics and cannot populate impact_validated. For transfer research use multimodal_transfer or model_provider_policy; missing evaluator references keep the result descriptive.