Skip to content

Crucible

Structure-aware fuzzing for the machine learning model supply chain and inference stack

Quick Start Architecture


Crucible is a structure-aware fuzzer for the code that turns an untrusted model file into a running model. That path runs through hand-written parsers, loaders, and quantization and compute kernels inside tools such as llama.cpp, whisper.cpp, stable-diffusion.cpp, ONNX Runtime, and TensorFlow Lite. Crucible targets semantic fields to improve reach beyond shallow format gates, then replays and source-reviews each surviving observation before it becomes a finding.

18 CVEs reserved/assigned
17 Public CVE records
211 Research findings
71 Named projects

Latest CVE row-level record check: 2026-10-04; earlier rows retain their own dates. Research breadth is the audited 2026-09-07 snapshot: a finding is a documented research case, and a named project is a distinct primary target repository. The snapshot includes independent rediscoveries and robustness observations at differing evidence levels; its totals do not imply novelty, vendor acceptance, or a current inventory.

Why Crucible?

The model file is untrusted input, and it reaches memory-unsafe code

Loading a model through runtimes such as Ollama, LM Studio, llama.cpp, whisper.cpp, stable-diffusion.cpp, ONNX Runtime, or TensorFlow Lite can place untrusted model bytes in native parsers, loaders, and kernels. Networked deployments add protocol inputs such as ggml-rpc messages.

The public record includes out-of-bounds reads and writes, arithmetic defects, reachable aborts, and resource-lifetime failures reached through crafted model or protocol inputs. Their severity varies with the real entry point and deployment boundary; a crash class alone cannot supply a CVSS vector. They keep appearing because the code is hand-written, invariants are implicit, and complex artifacts cross trust boundaries before inference begins.

Generic fuzzers stall on the format. Point AFL++ or libFuzzer at a binary model format and most of their inputs die on the first few bytes. Crucible instead:

  • Parses valid seed files into typed structures (headers, metadata key-value pairs, tensor info blocks, alignment, data regions)
  • Mutates at the semantic level: corrupting cross-field invariants, injecting type-confused metadata, overflowing dimension products, and forging the quantization fields that drive kernel addressing
  • Serializes back to binary while preserving the particular valid and invalid relationships the strategy intended

The surface, from parse to run

Security work in ML has concentrated on model behavior (adversarial examples, prompt injection) and the web layer (SSRF, path traversal). The layer in between, the binary parsers and the quantization and compute kernels that every inference stack depends on, has had far less systematic attention. That is the surface Crucible targets, and it splits into two stages:

  1. Load time. Parsers and loaders read the file's structure. A crafted GGUF, ONNX, SafeTensors, or TFLite file with corrupted offsets, dimension products, or metadata drives out-of-bounds reads and writes before the model ever runs. Most inference stacks are now hardening this stage with up-front verification, and Crucible tracks who has and who has not.
  2. Run time. After a file parses cleanly, quantization metadata can flow into pointer arithmetic in dequantization and GEMM kernels. A structurally valid but semantically hostile value may only fault when a natural consumer runs. Up-front verification helps, but it must validate the same invariants the downstream kernels use.

One useful review technique across both stages is a differential: a sibling path validates the same value another path uses raw. The GPU kernel may bound an index the CPU kernel does not; one branch of a loader may have a check its twin omits. That contrast is evidence for the intended invariant, not a substitute for replay, reachability, and source review.

Impact

The publicly validated cases include memory-safety and denial-of-service bugs in real parsers. The counted findings above span a broader research inventory; check the evidence and public-record state before treating any individual case as live or assigned. The dated research metrics methodology explains the counting units and limits. The generated CVE catalog lists records with a public source and omits identifiers lacking a public record. Public findings and fixes links to selected upstream reports and remediations; ecosystem prior art maps the broader target surface.

One bug, made visible

CVE-2026-14647 mechanism: mismatched ONNX ranks drive a heap out-of-bounds read

The 255-byte ONNX reproducer is small; the violated invariant is smaller. Shape inference derives three kernel dimensions from the weight tensor but allocates two dilation values from the input rank, then uses one length to index the other. The worked visual separates the witnessed primitive, upstream fix, and score provenance.

What ships

make build produces three binaries, and those three are the public tool:

Binary Role
crucible Environment checks, corpus work, campaigns, triage, analysis
crucible-gen Synthetic GGUF seed generation
crucible-triage Crash triage and report generation

Other directories under cmd/ hold mutator archives, experiment drivers and fixture tools that are developer-only: not built by make build, not documented as commands, and carrying no stability expectation. Installation defines each support label used across these pages.

Quick start

No target needed. Generate a synthetic corpus, read its structure, and mutate it, all from a clean checkout:

make build
crucible-gen --output ./corpus --count 5 --seed 42
crucible gguf inspect corpus/seed_001.gguf
crucible mutate corpus/seed_001.gguf ./mutated.gguf --seed 7 --show-mutations

The quick start walks through that with real output.

With a target. The harness build is target-specific; these flows need a local upstream checkout:

make build
make generate
TARGET_COMMIT=$(git -C /path/to/llama.cpp rev-parse HEAD)
make -C targets/llamacpp build-fuzz LLAMA_CPP=/path/to/llama.cpp \
  LLAMA_CPP_VERSION="$TARGET_COMMIT"
make -C harness/libfuzzer LLAMA_CPP=/path/to/llama.cpp crucible-libfuzzer
./crucible harness-smoke --harness ./harness/libfuzzer/crucible-libfuzzer \
  --seed ./corpus/generated/seed_000.gguf
./crucible run --harness ./harness/libfuzzer/crucible-libfuzzer \
  --corpus ./corpus/generated --output ./crashes
./crucible triage --artifact-kind input --crashes ./crashes \
  --harness ./harness/libfuzzer/crucible-libfuzzer --output ./reports
make build
make generate
make harness-afl LLAMA_CPP=/path/to/llama.cpp
make run-afl LLAMA_CPP=/path/to/llama.cpp
# Replay AFL artifacts through the corresponding file-argument harness.
make build
make generate
make fuzz-go
# Review Go fuzz failures through the Go test replay path.

Prerequisites

You need a clean local clone of the target project for the C and C++ harnesses (for example llama.cpp). Set the target path and current commit when building. Use a fresh target checkout or a build directory with a matching Crucible commit marker; the recipe rejects unmarked existing build objects. The GGUF seed shown above is created by make generate and checks parser-harness lifecycle; it is not a full model-loader control. The Go native harness requires only the Go toolchain.

How it works

Structure-Aware Mutation Engine

Crucible is not limited to random byte flips. Its format-specific engines can operate on a parsed structure, modifying header fields, injecting malformed metadata, corrupting tensor dimensions, forging quantization fields, and breaking cross-field invariants, while byte mutation remains available when structure-aware mutation cannot apply.

Weighted Strategy Selection

Not all mutation categories receive equal probability. The current GGUF profile weights metadata and tensor-info mutations most heavily; the table describes configured selection mass, not a universal measurement of bug density:

Category Weight Focus
Metadata 35% String handling, type confusion, model-loader targeting
Tensor Info 35% Dimension overflows, offset manipulation, type fuzzing
Header 10% Version, counts, magic corruption
Consistency 10% Cross-field mismatches, where the worst bugs hide
Alignment 5% Padding and stride calculation bugs
Data 5% Truncation, overlap, size mismatches

Two Crash Identities, Two Jobs

The triage engine parses sanitizer output and records an Exact build-local identity for deduplication plus a normalized Stable identity for cross-build regression tracking. Thousands of duplicate artifacts collapse locally without pretending that line numbers or build paths are portable across machines.

The Differential as a Review Technique

When one code path validates a file-controlled value and a sibling path does not, the contrast supplies an invariant to test. Crucible surfaces the observation; replay and source review determine whether the differential explains a defect and whether it applies to the proposed fix.

Evidence-Ready Triage

Triage output includes process outcome, evidence class, current identities, target attribution, raw stack evidence, and a reproducer reference. Severity is unrated and affected versions remain unknown until an operator validates them. The workflow requires replay on a clean HEAD build and a source read before an observation can be named as a finding; historical records must be audited rather than assumed to satisfy that contract.

Active Research

Additional findings are under coordinated disclosure with upstream maintainers and vendors. Specifics are published only after each disclosure completes.


Crucible is a Halo Forge Labs project.