Skip to content

Multimodal and Provider Playbook

First-party packs: multimodal.input-trust and model.provider-boundary.

Deposit every input before target contact. A manifest records digest, MIME type, structure, transformations, layers, semantic labels, control relationship, provenance and CAS references.

R/R/R boundary

Part Status
Input bytes and transformations deposited and digest-addressed
Model or provider path exact authorized endpoint and model identity for an external result; local compatible endpoints are instrument checks
Effect evidence target-owned state or a controlled action source outside response scoring
Claim limit refusal, text matching, latency, and token use are behavior metrics; none alone establishes external impact

Multimodal variants

  • multimodal.image-instruction-injection
  • multimodal.audio-instruction-injection
  • multimodal.document-hidden-layer-injection
  • multimodal.cross-modal-conflict

Each needs a byte-deterministic attack input and close control. Bind them separately through media_manifest_refs and control_media_manifest_refs. The OpenAI-compatible runtime resolves only deposited ait-media://N references. Other providers implement provider_jsonl and receive the case-appropriate deposited manifest/input references.

Provider variants

  • model.policy-boundary-jailbreak
  • model.provider-system-prompt-extraction
  • model.privacy-canary-extraction
  • model.resource-amplification

Use synthetic policies, prompts and canaries. Do not treat arbitrary text matches as real-secret extraction. Record endpoint, model identity, request/response digests, latency, token use, retries and evaluator version.

openai_http supports Responses and Chat Completions, bounded JSON/SSE, request/response limits, input/output token limits, budgets and zero retries by default. Credentials are env:NAME references. Explicit retries record every attempt.

Evaluators measure refusal, leakage, policy violation, latency, tokens and amplification. These are behavioral metrics and cannot populate impact_validated. For transfer research use multimodal_transfer or model_provider_policy; missing evaluator references keep the result descriptive.