Skip to content

Mutation Engine

Crucible uses structure-aware mutation to preserve enough syntax to exercise semantic checks while still breaking the invariants that matter to parsers, loaders, and compute paths. It does not assume that structured mutation always wins. Byte mutation and structure-aware mutation test different parts of the same surface and should be measured in matched campaigns.

Two complementary modes

Mode Strong at Blind spot
Byte mutation Truncation, corrupt framing, parser error paths, and malformed structures a serializer cannot emit Often spends substantial budget on inputs rejected before semantic code
Structure-aware mutation Cross-field mismatches, hostile dimensions and offsets, enum boundaries, and protocol-valid state transitions Can preserve a gate whose failure path is itself vulnerable

Structure-aware output is not guaranteed to parse, reach a particular line, or fault. Some strategies deliberately corrupt magic, counts, lengths, or padding. The claim is narrower: the engine can target named fields and relationships instead of relying only on chance byte placement.

Format-specific engines

One schema does not describe every attack surface, so Crucible keeps separate engines behind a shared campaign workflow.

Surface Package Structural knowledge
GGUF pkg/mutator header, metadata, tensor descriptors, alignment, data, and cross-field invariants
Declarative GGUF synthesis pkg/mutator/synth validated mutation rules authored from a format specification
ggml-rpc pkg/mutator/rpc compiler-measured wire profiles, messages, tensor fields, graph requests, and stateful sequences
ONNX/protobuf pkg/mutator/protobuf wire types, varints, and length-delimited fields
TFLite, ExecuTorch, MNN pkg/mutator/flatbuffer vtables, offsets, vectors, and table fields
SafeTensors pkg/mutator/safetensors header length, JSON tensor metadata, offsets, shapes, and dtypes
NumPy .npy pkg/mutator/npy header dictionary, shape, dtype, order, and payload relationship
Tokenizer artifacts pkg/mutator/tokenizer vocabulary and tokenizer-specific structured fields

The generated format and harness topology is the inventory check. A new mutator package must appear there instead of being silently omitted from the public architecture.

GGUF selection contract

For a parsed GGUF seed, pkg/mutator.Mutator:

  1. selects one to three strategies;
  2. selects a category by the weights below;
  3. selects uniformly among registered strategies in that category;
  4. mutates the in-memory gguf.File;
  5. synchronizes header counts unless a strategy explicitly preserves a mismatch; and
  6. serializes the result.
flowchart LR
    A[GGUF seed bytes] --> B[Parse]
    B --> C[Select 1-3 strategies]
    C --> D[Mutate typed fields]
    D --> E[Preserve intentional mismatches]
    E --> F[Serialize]
    F --> G[Target harness]
    B -->|not GGUF| H[Delegate to libFuzzer byte mutation]

The current category weights are code constants:

Category Weight Examples
Metadata, including model-loader strategies 35% keys, types, arrays, architecture, vocabulary
Tensor information 35% dimensions, types, names, offsets
Header 10% magic, version, declared counts
Consistency 10% count, size, offset, and alignment disagreement
Alignment 5% padding and alignment values
Data 5% truncation, overlap, zero length, special floats

There are currently 46 registered GGUF strategies. The canonical names and per-category counts live in Mutation Strategies; tests should be updated with the code when that inventory changes.

The custom-mutator fallback

LLVMFuzzerCustomMutator replaces libFuzzer's default mutator. Returning the input unchanged is not a fallback. When the GGUF engine cannot parse or mutate an input, Crucible now delegates to the runtime's LLVMFuzzerMutate implementation. The archive carries a weak stub only so ordinary Go builds can link without libFuzzer.

This creates a build-time obligation: a real fuzz harness must prove that the runtime's strong symbol won. Crucible's mutator-linkage gate checks that condition, and the harness build scripts must force-link the custom entry point. A campaign should not infer linkage from the archive merely being present on the link line.

Reproducibility boundary

A non-zero explicit seed makes selection deterministic for the same engine version, starting bytes, and call sequence. It does not make a whole fuzz campaign byte-identical: libFuzzer scheduling, corpus evolution, target build, environment, and concurrency also affect the run. Bank those inputs when comparing arms.

Measuring effectiveness

The relevant question is empirical: under matched seeds, corpus, target build, limits, and runtime, does a structured arm improve semantic reach or finding yield over a byte-only arm?

The retained SafeTensors pilot found two unique identities in the structured arm and one in the byte arm. It is useful mechanism evidence: the structured-only identity lived behind the JSON parse gate. With one run per arm, it is not an uplift estimate. Replicated campaigns must bank empty runs, record censored time-to-event observations correctly, and keep pkg/experiment in its actual role: scoring completed runs, not launching or capturing them.

See Measuring Mutation Effectiveness for the experiment contract.

What a strategy name proves

A name such as tensorinfo.dim_product_overflow records what the mutator attempted. It does not prove the target overflowed, allocated incorrectly, or became exploitable. Those claims require a replayed execution, process facts, the target's raw diagnostic, and source inspection under the triage contract.