Fuzzing llama.cpp¶
llama.cpp exposes several distinct attack surfaces: GGUF parsing and writing, model loading, tokenizers, grammar and template handling, quantization, server inputs, and ggml-rpc. Name the surface before building; a result from one harness is not coverage of the others.
1. Start from a clean current tree¶
git clone https://github.com/ggml-org/llama.cpp.git ~/src/llama.cpp
git -C ~/src/llama.cpp fetch origin
git -C ~/src/llama.cpp switch --detach origin/HEAD
./crucible provenance --strict ~/src/llama.cpp
./crucible recon llama.cpp "GGUF model loader"
Historical pins are useful for regression experiments, but a verdict on a historical commit is not a verdict on current upstream. Keep the pinned and HEAD builds separate.
2. Choose and build a harness¶
List the available targets:
For the general model loader, build the instrumented target at the recorded commit first. The target recipe's default pin is historical, so pass the current commit explicitly:
TARGET_COMMIT=$(git -C "$HOME/src/llama.cpp" rev-parse HEAD)
make -C targets/llamacpp build-fuzz \
LLAMA_CPP="$HOME/src/llama.cpp" \
LLAMA_CPP_VERSION="$TARGET_COMMIT"
make -C harness/libfuzzer \
LLAMA_CPP="$HOME/src/llama.cpp" \
crucible-libfuzzer-model
Start from a fresh target checkout or a build-fuzz directory stamped by this recipe for the recorded commit. An unmarked existing build is rejected so ignored objects from an earlier revision cannot silently supply the harness.
For a structure-aware variant, build the corresponding mutator archive/harness and verify linkage:
The symbol must be defined, not weak-undefined. A missing custom mutator silently turns the intended structure-aware arm into byte mutation.
RPC builds¶
RPC harnesses require llama.cpp built with RPC enabled and linked against its RPC implementation. The wire profile belongs to the pinned build; do not copy a struct layout from a different commit.
When running libFuzzer RPC harnesses, pass:
Without those flags, libFuzzer can terminate on its own allocation limit before the target's bad_alloc handling runs, manufacturing a fuzzer OOM that says nothing about the server.
3. Prepare controls and corpus¶
Keep three populations separate:
- a known-good control artifact;
- a retained positive for the selected surface;
- a pristine launch corpus that does not contain the positive.
For GGUF:
This generates built-in structural GGUF seeds; it does not provide a known-good full model or a retained crashing input for the model-loader harness. Supply those separately for the target revision being tested. Real models can add production structure, but record their source and hash and keep the libFuzzer launch directory flat.
4. Prove the harness is meaningful¶
./crucible harness-smoke \
--harness ./harness/libfuzzer/crucible-libfuzzer-model \
--seed /path/to/known-good-model.gguf
./crucible preflight \
--harness ./harness/libfuzzer/crucible-libfuzzer-model \
--control /path/to/known-good-model.gguf \
--known-crash /path/to/retained-model-crash.gguf \
--corpus ./corpus/llama-model-seeds
Replace the /path/to/ examples with actual local files. Exit 1 means a control failed; exit 2 means visibility was not established. Without a retained positive, preflight exits 2. You may still explore, but a quiet campaign cannot support a negative finding claim.
5. Run and preserve the command¶
./crucible run \
--harness ./harness/libfuzzer/crucible-libfuzzer-model \
--corpus ./corpus/llama-model-seeds \
--output ./crashes/llama-model \
--dict corpus/gguf.dict \
--jobs 8 \
--sift \
--supervise \
--json \
-- -max_total_time=3600 -seed=42 -print_final_stats=1
Use the engine's actual telemetry for executions, coverage, and memory. Do not promise a generic executions-per-second or time-to-first-crash number: both are surface, corpus, sanitizer, hardware, and build dependent.
6. Inspect saturation¶
./crucible saturation \
--harness ./harness/libfuzzer/crucible-libfuzzer-model \
--artifacts ./crashes/llama-model
If one Exact identity dominates, design a narrow reversible escape from the target condition and verify it:
The verifier must recover the same escaped bug when disabled and must retain a different known bug. Otherwise the filter may simply have blinded the harness.
7. Triage observations¶
./crucible triage \
--artifact-kind input \
--crashes ./crashes/llama-model \
--harness ./harness/libfuzzer/crucible-libfuzzer-model \
--target gguf-model-loader \
--output ./reports/llama-model \
--json
Full triage replays binary inputs and groups by Exact identity. It records Stable identity for cross-build tracking and the legacy stack hash for historical oracle compatibility. Generated reports are unrated.
For every unique observation:
- replay the control and crafted input;
- read the complete sanitizer output;
- read the target source and callers at the attributed site;
- determine whether the primitive is read, write, assertion, arithmetic, resource exhaustion, or a harness/runtime artifact;
- search existing llama.cpp issues and advisories;
- replay on a clean current HEAD build;
- bank raw output and all provenance hashes.
8. Validate a candidate fix¶
A fix must close the same primitive, preserve genuine models, execute the new guard, cover sibling sites, and retain visibility of a different positive. “The PoC no longer crashes” is necessary but not sufficient.
See Responsible Disclosure before sending any report or PR. A public guard diff can disclose the vulnerable condition even when the PoC is withheld.