Skip to content

Fuzzing llama.cpp

llama.cpp exposes several distinct attack surfaces: GGUF parsing and writing, model loading, tokenizers, grammar and template handling, quantization, server inputs, and ggml-rpc. Name the surface before building; a result from one harness is not coverage of the others.

1. Start from a clean current tree

git clone https://github.com/ggml-org/llama.cpp.git ~/src/llama.cpp
git -C ~/src/llama.cpp fetch origin
git -C ~/src/llama.cpp switch --detach origin/HEAD

./crucible provenance --strict ~/src/llama.cpp
./crucible recon llama.cpp "GGUF model loader"

Historical pins are useful for regression experiments, but a verdict on a historical commit is not a verdict on current upstream. Keep the pinned and HEAD builds separate.

2. Choose and build a harness

List the available targets:

make -C harness/libfuzzer help 2>/dev/null || \
  sed -n '1,220p' harness/libfuzzer/Makefile

For the general model loader, build the instrumented target at the recorded commit first. The target recipe's default pin is historical, so pass the current commit explicitly:

TARGET_COMMIT=$(git -C "$HOME/src/llama.cpp" rev-parse HEAD)
make -C targets/llamacpp build-fuzz \
  LLAMA_CPP="$HOME/src/llama.cpp" \
  LLAMA_CPP_VERSION="$TARGET_COMMIT"
make -C harness/libfuzzer \
  LLAMA_CPP="$HOME/src/llama.cpp" \
  crucible-libfuzzer-model

Start from a fresh target checkout or a build-fuzz directory stamped by this recipe for the recorded commit. An unmarked existing build is rejected so ignored objects from an earlier revision cannot silently supply the harness.

For a structure-aware variant, build the corresponding mutator archive/harness and verify linkage:

make verify-mutators
nm ./harness/libfuzzer/<mutator-harness> | grep LLVMFuzzerCustomMutator

The symbol must be defined, not weak-undefined. A missing custom mutator silently turns the intended structure-aware arm into byte mutation.

RPC builds

RPC harnesses require llama.cpp built with RPC enabled and linked against its RPC implementation. The wire profile belongs to the pinned build; do not copy a struct layout from a different commit.

When running libFuzzer RPC harnesses, pass:

-rss_limit_mb=0 -malloc_limit_mb=0

Without those flags, libFuzzer can terminate on its own allocation limit before the target's bad_alloc handling runs, manufacturing a fuzzer OOM that says nothing about the server.

3. Prepare controls and corpus

Keep three populations separate:

  • a known-good control artifact;
  • a retained positive for the selected surface;
  • a pristine launch corpus that does not contain the positive.

For GGUF:

./crucible generate \
  --output ./corpus/llama-model-seeds \
  --count 100 \
  --seed 42

This generates built-in structural GGUF seeds; it does not provide a known-good full model or a retained crashing input for the model-loader harness. Supply those separately for the target revision being tested. Real models can add production structure, but record their source and hash and keep the libFuzzer launch directory flat.

4. Prove the harness is meaningful

./crucible harness-smoke \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --seed /path/to/known-good-model.gguf

./crucible preflight \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --control /path/to/known-good-model.gguf \
  --known-crash /path/to/retained-model-crash.gguf \
  --corpus ./corpus/llama-model-seeds

Replace the /path/to/ examples with actual local files. Exit 1 means a control failed; exit 2 means visibility was not established. Without a retained positive, preflight exits 2. You may still explore, but a quiet campaign cannot support a negative finding claim.

5. Run and preserve the command

./crucible run \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --corpus ./corpus/llama-model-seeds \
  --output ./crashes/llama-model \
  --dict corpus/gguf.dict \
  --jobs 8 \
  --sift \
  --supervise \
  --json \
  -- -max_total_time=3600 -seed=42 -print_final_stats=1

Use the engine's actual telemetry for executions, coverage, and memory. Do not promise a generic executions-per-second or time-to-first-crash number: both are surface, corpus, sanitizer, hardware, and build dependent.

6. Inspect saturation

./crucible saturation \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --artifacts ./crashes/llama-model

If one Exact identity dominates, design a narrow reversible escape from the target condition and verify it:

./crucible escape-verify --help

The verifier must recover the same escaped bug when disabled and must retain a different known bug. Otherwise the filter may simply have blinded the harness.

7. Triage observations

./crucible triage \
  --artifact-kind input \
  --crashes ./crashes/llama-model \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --target gguf-model-loader \
  --output ./reports/llama-model \
  --json

Full triage replays binary inputs and groups by Exact identity. It records Stable identity for cross-build tracking and the legacy stack hash for historical oracle compatibility. Generated reports are unrated.

For every unique observation:

  1. replay the control and crafted input;
  2. read the complete sanitizer output;
  3. read the target source and callers at the attributed site;
  4. determine whether the primitive is read, write, assertion, arithmetic, resource exhaustion, or a harness/runtime artifact;
  5. search existing llama.cpp issues and advisories;
  6. replay on a clean current HEAD build;
  7. bank raw output and all provenance hashes.

8. Validate a candidate fix

A fix must close the same primitive, preserve genuine models, execute the new guard, cover sibling sites, and retain visibility of a different positive. “The PoC no longer crashes” is necessary but not sufficient.

See Responsible Disclosure before sending any report or PR. A public guard diff can disclose the vulnerable condition even when the PoC is withheld.