Mutation Engine¶
Crucible uses structure-aware mutation to preserve enough syntax to exercise semantic checks while still breaking the invariants that matter to parsers, loaders, and compute paths. It does not assume that structured mutation always wins. Byte mutation and structure-aware mutation test different parts of the same surface and should be measured in matched campaigns.
Two complementary modes¶
| Mode | Strong at | Blind spot |
|---|---|---|
| Byte mutation | Truncation, corrupt framing, parser error paths, and malformed structures a serializer cannot emit | Often spends substantial budget on inputs rejected before semantic code |
| Structure-aware mutation | Cross-field mismatches, hostile dimensions and offsets, enum boundaries, and protocol-valid state transitions | Can preserve a gate whose failure path is itself vulnerable |
Structure-aware output is not guaranteed to parse, reach a particular line, or fault. Some strategies deliberately corrupt magic, counts, lengths, or padding. The claim is narrower: the engine can target named fields and relationships instead of relying only on chance byte placement.
Format-specific engines¶
One schema does not describe every attack surface, so Crucible keeps separate engines behind a shared campaign workflow.
| Surface | Package | Structural knowledge |
|---|---|---|
| GGUF | pkg/mutator | header, metadata, tensor descriptors, alignment, data, and cross-field invariants |
| Declarative GGUF synthesis | pkg/mutator/synth | validated mutation rules authored from a format specification |
| ggml-rpc | pkg/mutator/rpc | compiler-measured wire profiles, messages, tensor fields, graph requests, and stateful sequences |
| ONNX/protobuf | pkg/mutator/protobuf | wire types, varints, and length-delimited fields |
| TFLite, ExecuTorch, MNN | pkg/mutator/flatbuffer | vtables, offsets, vectors, and table fields |
| SafeTensors | pkg/mutator/safetensors | header length, JSON tensor metadata, offsets, shapes, and dtypes |
NumPy .npy | pkg/mutator/npy | header dictionary, shape, dtype, order, and payload relationship |
| Tokenizer artifacts | pkg/mutator/tokenizer | vocabulary and tokenizer-specific structured fields |
The generated format and harness topology is the inventory check. A new mutator package must appear there instead of being silently omitted from the public architecture.
GGUF selection contract¶
For a parsed GGUF seed, pkg/mutator.Mutator:
- selects one to three strategies;
- selects a category by the weights below;
- selects uniformly among registered strategies in that category;
- mutates the in-memory
gguf.File; - synchronizes header counts unless a strategy explicitly preserves a mismatch; and
- serializes the result.
flowchart LR
A[GGUF seed bytes] --> B[Parse]
B --> C[Select 1-3 strategies]
C --> D[Mutate typed fields]
D --> E[Preserve intentional mismatches]
E --> F[Serialize]
F --> G[Target harness]
B -->|not GGUF| H[Delegate to libFuzzer byte mutation] The current category weights are code constants:
| Category | Weight | Examples |
|---|---|---|
| Metadata, including model-loader strategies | 35% | keys, types, arrays, architecture, vocabulary |
| Tensor information | 35% | dimensions, types, names, offsets |
| Header | 10% | magic, version, declared counts |
| Consistency | 10% | count, size, offset, and alignment disagreement |
| Alignment | 5% | padding and alignment values |
| Data | 5% | truncation, overlap, zero length, special floats |
There are currently 46 registered GGUF strategies. The canonical names and per-category counts live in Mutation Strategies; tests should be updated with the code when that inventory changes.
The custom-mutator fallback¶
LLVMFuzzerCustomMutator replaces libFuzzer's default mutator. Returning the input unchanged is not a fallback. When the GGUF engine cannot parse or mutate an input, Crucible now delegates to the runtime's LLVMFuzzerMutate implementation. The archive carries a weak stub only so ordinary Go builds can link without libFuzzer.
This creates a build-time obligation: a real fuzz harness must prove that the runtime's strong symbol won. Crucible's mutator-linkage gate checks that condition, and the harness build scripts must force-link the custom entry point. A campaign should not infer linkage from the archive merely being present on the link line.
Reproducibility boundary¶
A non-zero explicit seed makes selection deterministic for the same engine version, starting bytes, and call sequence. It does not make a whole fuzz campaign byte-identical: libFuzzer scheduling, corpus evolution, target build, environment, and concurrency also affect the run. Bank those inputs when comparing arms.
Measuring effectiveness¶
The relevant question is empirical: under matched seeds, corpus, target build, limits, and runtime, does a structured arm improve semantic reach or finding yield over a byte-only arm?
The retained SafeTensors pilot found two unique identities in the structured arm and one in the byte arm. It is useful mechanism evidence: the structured-only identity lived behind the JSON parse gate. With one run per arm, it is not an uplift estimate. Replicated campaigns must bank empty runs, record censored time-to-event observations correctly, and keep pkg/experiment in its actual role: scoring completed runs, not launching or capturing them.
See Measuring Mutation Effectiveness for the experiment contract.
What a strategy name proves¶
A name such as tensorinfo.dim_product_overflow records what the mutator attempted. It does not prove the target overflowed, allocated incorrectly, or became exploitable. Those claims require a replayed execution, process facts, the target's raw diagnostic, and source inspection under the triage contract.