Architecture¶
This page describes the high-level design of Crucible, the relationships between its packages, and the data flow from seed files through mutation and execution to evidence-backed human triage.
High-Level Data Flow¶
Crucible is two halves. The discovery half turns seeds into classified crashes. The evidence half turns a classified crash into a claim that survives someone else checking it. Both are part of the tool; a crash without evidence is not a finding, and a fix without evidence is not a fix.
flowchart LR
subgraph DISC["Discovery"]
direction TB
Seeds["Seeds<br/>minimal GGUF files"]
Corpus["Corpus<br/>generated + downloaded"]
Mutator["Mutator<br/>structure-aware"]
Harness["Harness<br/>libFuzzer / AFL++ / Go"]
Target["Target parser"]
Crashes["Execution artifacts<br/>diagnostics + process facts"]
Triager["Observation + identity<br/>Exact dedup, Stable tracking"]
Seeds --> Corpus --> Mutator --> Harness --> Target
Target -->|crash| Crashes --> Triager
Target -->|no crash| Harness
end
subgraph EVID["Evidence"]
direction TB
Tier["Evidence tier<br/>detorch / real-artifact / end-to-end"]
Reach["Reachability<br/>compiled? linked? reached?"]
Patch["Candidate fix"]
Verify["Patch verification"]
Tier --> Reach --> Patch --> Verify
end
Triager --> Tier
Verify --> Gate{"Operator gates<br/>severity + disclosure"}
Gate --> Out["Public issue / PR /<br/>advisory / CVE"]
style DISC fill:#3b3028,stroke:#b08a5a,color:#fff
style EVID fill:#3c432d,stroke:#9a9b62,color:#fff
style Gate fill:#542f2b,stroke:#b96f5c,color:#fff The two human gates -- exploitability tier and disclosure -- sit at the end by design and are never automated. Everything upstream of them may surface and count; nothing upstream of them may decide.
Artifact counts are not bug counts¶
A campaign writes thousands of crash artifacts for a handful of distinct bugs. The pipeline therefore never treats artifact volume as a bug count. Exact identities create build-local replay buckets; Stable identities track likely continuity across builds. Neither proves that two inputs share a source-level root cause, so every surviving observation is replayed and read by hand before it is named.
Package Relationships¶
Crucible is organized as format/mutation, execution/triage, protocol, telemetry, experiment, and evidence packages under pkg/, with command-line entry points under cmd/ and C-archive builds for custom mutators:
graph TD
subgraph "cmd/"
crucible["cmd/crucible\n(main orchestrator)"]
gen["cmd/crucible-gen\n(corpus generator)"]
triage_cmd["cmd/crucible-triage\n(crash triager)"]
mutator["cmd/crucible-mutator\n(C-archive for libFuzzer)"]
rpcmutator["cmd/crucible-rpc-mutator\n(RPC C-archive for libFuzzer)"]
end
subgraph "pkg/"
gguf["pkg/gguf\n(format types + reader/writer)"]
mut["pkg/mutator\n(GGUF mutation engine)"]
rpcmut["pkg/mutator/rpc\n(RPC mutation engine)"]
corpus["pkg/corpus\n(seed generation + management)"]
triage["pkg/triage\n(crash dedup + reporting)"]
coverage["pkg/coverage\n(LLVM coverage collection + reporting)"]
experiment["pkg/experiment\n(replicated-run scoring contract)"]
rpc["pkg/rpc\n(profiled ggml-rpc execution)"]
mutstats["pkg/mutstats\n(strategy telemetry)"]
end
subgraph "harness/"
libfuzzer["harness/libfuzzer\n(C harness)"]
aflpp["harness/aflpp\n(C harness)"]
gofuzz["harness/go\n(Go native fuzz tests)"]
end
crucible --> gguf
crucible --> mut
crucible --> corpus
gen --> gguf
gen --> corpus
triage_cmd --> triage
mut --> gguf
rpcmut --> gguf
corpus --> gguf
corpus --> mut
gofuzz --> gguf
gofuzz --> mut
libfuzzer -.->|"links against"| gguf
aflpp -.->|"links against"| gguf Dependency direction
The format packages remain below their mutators, while execution services are reused by CLI consumers. Consult go list -deps for the current exact graph; this diagram describes ownership, not a promise that packages such as triage have no internal dependencies.
The Validation Layer¶
The tools under tools/ answer different evidence questions. live and patchcheck assess target builds and candidate fixes; ledger generates canonical counts from repository records. They are additive to discovery and do not change pkg/ or the harnesses.
graph TD
subgraph "tools/"
LIVE["<b>tools/live</b><br/>crucible-live<br/><i>is this code compiled,<br/>linked, referenced?</i>"]
PC["<b>tools/patchcheck</b><br/>crucible-patchcheck<br/><i>does this patch close the bug<br/>without breaking real input?</i>"]
LED["<b>tools/ledger</b><br/>build-ledger<br/><i>canonical counts,<br/>generated not typed</i>"]
end
subgraph "inputs"
BIN["Built artifact<br/>(binary, compile_commands,<br/>project file)"]
PATCH["Candidate patch"]
GEN["Genuine artifacts<br/>(real models / indexes)"]
TREE["findings/ + reports/"]
end
BIN --> LIVE
LIVE --> PC
PATCH --> PC
GEN --> PC
TREE --> LED
LED --> OUT["reports/generated/<br/>COUNTS.md + LEDGER.json"]
style LIVE fill:#44372b,stroke:#c49a6c,color:#fff
style PC fill:#3c432d,stroke:#9a9b62,color:#fff
style LED fill:#4b362c,stroke:#a86f4c,color:#fff live reports PRESENT / ABSENT / UNKNOWN, while patchcheck reports PASS / FAIL / UNKNOWN per check. UNKNOWN is a normal and correct answer; a weak negative must not silently become a verdict. Those assessment scripts leave the decision to the operator. build-ledger.py --check is different: it exits nonzero when its source binding or generated output has drifted, and the repository uses it as a correctness gate.
See Evidence and Validation for the evidence ladder, what each tier does and does not establish, and why the remedy carries a higher bar than the finding.
The correctness gates¶
The assessment scripts above report evidence for an operator to judge. Correctness gates refuse to proceed when a required condition fails; their exit codes are part of the interface.
| gate | refuses when |
|---|---|
crucible run --sift | a seed already crashes the harness, so libFuzzer would die during corpus replay |
crucible run --supervise | the harness thrashes, or no new sanitizer signature has appeared for N restarts |
crucible provenance --strict | the target checkout has modified tracked files, so the witness does not describe that commit |
tools/check-evidence.py | a named bundle is missing or a selected sanitizer fragment appears nowhere in its named bundle set; exact frames, lines, artifacts, and builds remain unchecked |
tools/ledger/build-ledger.py --check | committed source bindings or generated finding counts have drifted |
tools/check-findings-yaml.py | the authoritative state file does not parse, or has duplicate mapping keys |
tools/check-locked-artifacts.py | a locked reference manifest has been edited |
tools/diagrams/check.sh | a diagram no longer matches the code it claims to describe |
The rule they share is CLAUDE.md #8: a tool whose failure branch returns something that looks like an answer will silently corrupt every result downstream. So "cannot tell" gets its own exit code (2), separate from both success and failure. check-evidence.py exits 2 when the evidence tree is absent on this machine, because "I cannot see the artifacts" is not "the artifacts are fine".
The generated diagram Correctness gates renders this pipeline from the code, including each gate's exit contract, so it cannot drift.
Operational detail and the failure each gate retires: Campaign operations.
Machine-readable results¶
Many result-bearing commands expose --json; check each command's help for support. The core triage and run envelopes carry a crucible build stamp and report incomplete observations explicitly. Other commands use their own schemas, and preflight reports through text and exit status. Consumers must distinguish "we looked and found nothing" from "we could not look". Field reference: JSON result shapes.
Directory Structure¶
crucible/
├── cmd/
│ ├── crucible/ # Main orchestrator binary
│ │ ├── main.go
│ │ ├── sift.go # --sift: move already-crashing seeds out before launch
│ │ ├── supervise.go # --supervise: restart past crashes, thrash + saturation guards
│ │ ├── provenance.go # commit AND dirty state of the target checkout
│ │ ├── json_results.go # the --json wire shapes (docs/reference/json-results.md)
│ │ ├── harness_smoke.go # clean / crash / did-not-run before a campaign
│ │ └── mutate_stats_ab.go # A/B a mutation weighting against a measured baseline
│ ├── crucible-gen/ # Corpus generation tool
│ ├── crucible-mutator/ # Custom mutator (C-archive for libFuzzer)
│ ├── crucible-mutator-traced/ # Same, with per-strategy tracing for yield measurement
│ ├── crucible-rpc-mutator/ # ggml RPC wire-format mutator
│ ├── crucible-flatbuffer-mutator/ # FlatBuffers structure-aware mutator (TFLite, MNN)
│ ├── crucible-protobuf-mutator/ # Protobuf structure-aware mutator (ONNX, TF)
│ ├── crucible-safetensors-mutator/ # safetensors structure-aware mutator (fastllm, mlx, candle)
│ └── crucible-triage/ # Crash triage and reporting
│ ├── main.go
│ └── watch.go # Watch-mode with checkpoint persistence
├── pkg/
│ ├── gguf/ # GGUF format implementation
│ │ ├── format.go # Types: Header, MetadataKV, TensorInfo, File
│ │ ├── format_test.go # Format unit tests
│ │ ├── reader.go # Binary deserialization (Unmarshal)
│ │ ├── reader_test.go # Reader unit tests
│ │ └── writer.go # Binary serialization (Marshal)
│ ├── mutator/ # Structure-aware mutation engine
│ │ ├── mutator.go # Orchestrator: weighted category selection
│ │ ├── mutator_test.go # Mutation engine tests
│ │ ├── header.go # 5 header mutation strategies
│ │ ├── metadata.go # 13 metadata mutation strategies
│ │ ├── tensorinfo.go # 8 tensor info mutation strategies
│ │ ├── alignment.go # 3 alignment mutation strategies
│ │ ├── data.go # 6 tensor data mutation strategies
│ │ ├── consistency.go # 6 cross-field consistency strategies
│ │ ├── model_loader.go # 5 model-loader strategies (weighted under metadata)
│ │ ├── synth/ # Synthesised whole-file generation
│ │ ├── rpc/ # ggml RPC wire messages, incl. rpc_tensor field fuzzing
│ │ ├── flatbuffer/ # FlatBuffers: vtable, offset and vector-length mutation
│ │ ├── protobuf/ # Protobuf: wire-type, varint and length-delimited mutation
│ │ └── safetensors/ # safetensors: header length, data_offsets, shape, dtype
│ ├── corpus/ # Corpus generation and management
│ │ ├── corpus.go # Corpus loading and enumeration
│ │ ├── corpus_test.go # Corpus unit tests
│ │ ├── generate.go # Seed file generation
│ │ ├── minimize.go # Corpus minimization
│ │ └── minimize_test.go # Minimization tests
│ └── triage/ # Crash analysis and reporting
│ ├── triage.go # Compatibility crash taxonomy and Exact-keyed Triager
│ ├── triage_test.go # Triage unit tests
│ ├── cwe.go # CWE identifier mapping for each crash type
│ ├── stackhash.go # Frozen legacy five-frame oracle hash
│ ├── identity.go # Exact build-local and Stable cross-build identities
│ ├── observation.go # Admission + text evidence + process facts
│ ├── stackhash_test.go # Stack hash tests
│ ├── minimize.go # Crash reproducer minimization (recursive WalkDir)
│ ├── minimize_test.go # Minimize tests
│ ├── replay.go # Crash replay against harness binaries
│ ├── replay_test.go # Replay unit tests
│ ├── report.go # Unrated internal report generation
│ ├── report_target_test.go # Report target detection tests
│ ├── sarif.go # SARIF 2.1.0 output with target tags
│ ├── sarif_test.go # SARIF output tests
│ └── fixture_test.go # Shared test fixtures
│ ├── coverage/ # LLVM coverage collection
│ │ └── coverage.go # Replay corpus, merge profraw, HTML report
│ ├── mutstats/ # Measured per-strategy yield, and A/B against a baseline
│ ├── experiment/ # Scoring contract for completed replicated runs
│ ├── rpc/ # ggml RPC wire model (GRAPH_COMPUTE et al)
│ ├── oracle/ # Known-signature lookup for triage
│ └── buildinfo/ # Build stamp carried by every JSON result
├── harness/
│ ├── libfuzzer/ # libFuzzer C++ harnesses
│ │ ├── harness.cpp # LLVMFuzzerTestOneInput targeting gguf_init_from_file
│ │ └── Makefile
│ ├── aflpp/ # AFL++ C harness
│ │ ├── harness.c
│ │ └── Makefile
│ └── go/ # Go native fuzz tests
│ └── fuzz_test.go # FuzzGGUFReader, FuzzMutator, FuzzRoundTrip
├── corpus/ # Seed corpus directory
│ └── gguf.dict # GGUF-specific dictionary for fuzzer guidance
├── crashes/ # Crash artifacts output directory
├── targets/ # Build configs for fuzz targets
│ ├── llamacpp/Makefile
│ └── ollama/Makefile
├── tools/
│ ├── check-evidence.py # narrow gate: paths, hash syntax, coarse sanitizer fragments
│ ├── check-findings-yaml.py # gate: the authoritative state file parses, no duplicate keys
│ ├── check-locked-artifacts.py # gate: locked reference manifests are unchanged
│ ├── ledger/ # generates reports/generated/{COUNTS.md,LEDGER.json}
│ ├── live/ # crucible-live: is this code compiled, linked, referenced?
│ ├── patchcheck/ # does this patch close the bug without breaking real input?
│ ├── triage-db/ # operational known-signature database (free to grow)
│ ├── oracle/ # LOCKED reference manifests; see the locked-artifact gate
│ ├── diagrams/ # ground-truth extractor; check.sh gates diagram drift
│ └── sweep/ # campaign shell drivers
├── .github/
│ └── workflows/
│ ├── ci.yml # CI pipeline: lint, test, build
│ └── docs.yml # Documentation build and deploy
├── scripts/
│ ├── build-linux.sh # Cross-compile harnesses for Linux
│ ├── build-targets.sh # Build all target parsers
│ ├── download-corpus.sh # Download seed corpus from model repos
│ ├── gen-arch-seeds.py # Generate per-architecture GGUF seeds
│ ├── gen-lora-seeds.py # Generate LoRA-specific GGUF seeds
│ ├── gen-server-seeds.py # Generate HTTP endpoint test seeds
│ ├── gen-whisper-audio-seeds.py # Generate PCM audio test seeds
│ ├── generate-targeted-seeds.py # Generate TALOS CVE-targeted seeds
│ ├── launch-phase2.sh # Launch Phase 2 campaign (multi-harness)
│ ├── run-all-campaigns.sh # Launch all campaigns in parallel
│ ├── run-campaign.sh # Single campaign launcher
│ └── rotate-logs.sh # Rotate campaign logs
├── docs/ # mkdocs-material documentation source
├── Makefile # Top-level build, test, fuzz, triage targets
├── mkdocs.yml # Documentation site configuration
├── FUZZING-ROADMAP.md # Living checklist of targets, campaigns, and findings
├── VERSIONS.env # Pinned target versions for CI reproducibility
├── go.mod
└── go.sum
Package Details¶
pkg/gguf -- Format Implementation¶
The foundation package. It defines the Go types that mirror the GGUF binary specification:
Header-- 4-byte magic (GGUF), version (uint32), tensor count (uint64), metadata KV count (uint64)MetadataKV-- key string + one of 13 metadata value kinds, includingSTRINGandARRAYTensorInfo-- tensor name, dimensions,ggml_typeenum, byte offset into the data sectionFile-- the complete in-memory representation of a GGUF file
The package provides Unmarshal([]byte) (*File, error) for parsing and Marshal(*File) ([]byte, error) for serialization. These form the round-trip pipeline that the mutation engine depends on.
Why not modify bytes directly?
Byte-level mutations are fast but structurally blind. Operating on parsed structures lets Crucible target semantic fields and cross-field relationships directly. It improves the chance of passing selected parse gates; it does not guarantee that every mutation parses or reaches a deep path.
pkg/mutator -- Mutation Engine¶
The GGUF mutation engine applies weighted structure-aware strategies across header, metadata, tensor, alignment, data, consistency, and loader semantics. The generated format topology is the maintained inventory. On each call to Mutate(), it:
- Randomly selects 1 to 3 mutations to apply
- For each mutation, picks a category using weighted random selection
- Picks a strategy uniformly at random within the chosen category
- Applies the strategy to the in-memory
*gguf.File - Serializes the result back to bytes via
gguf.MarshalRaw, preserving intentional count or padding mismatches when the strategy requests them
Each strategy file (header.go, metadata.go, tensorinfo.go, alignment.go, data.go, consistency.go, model_loader.go) exports a *Strategies() []Strategy function. The Mutator registers all of them at construction time.
The Strategy interface is intentionally minimal:
This makes it straightforward to add new strategies -- implement the interface, add it to the appropriate *Strategies() function, and it is automatically picked up by the engine.
pkg/corpus -- Seed Management¶
Handles three responsibilities:
- Generation --
crucible-genproduces structurally varied seed GGUF files covering different metadata types, tensor configurations, and alignment values - Loading -- reads seed files from disk for the fuzzing harnesses
- Minimization -- content-deduplicates entries or, with a usable LLVM coverage harness, applies greedy set cover while preserving entries that produced no coverage data
pkg/triage -- Crash Analysis¶
The triage package is format-agnostic. Its current pipeline:
- Admission -- requires explicit metadata before an artifact can carry an outcome.
- Observation -- combines text evidence with process facts; neither substitutes for the other.
- Identity -- Exact keys build-local buckets, Stable tracks normalized target paths across builds, and legacy
HashStackremains frozen for historical oracle compatibility. - Attribution -- filters sanitizer, fuzzer, harness, runtime, and compiler-library frames; no target frame yields an explicit
UNATTRIBUTEDresult. - Minimization -- preserves the relevant Exact identity rather than accepting any crash.
- Report generation -- writes unrated internal records. CVSS and affected versions are absent until supplied from operator-ratified evidence.
- SARIF export -- transports taxonomy and identities without upgrading an observation into a confirmed vulnerability.
Mutation Pipeline¶
The following diagram shows the detailed flow when a single mutated test case is produced:
flowchart TD
A["Parse seed file\n<code>gguf.Unmarshal(bytes)</code>"] --> B{"Select category\n(weighted random)"}
B -->|"35%"| C1["Metadata + Model-Loader\n(13 + 5 = 18 strategies)"]
B -->|"35%"| C2["Tensor Info\n(8 strategies)"]
B -->|"10%"| C3["Header\n(5 strategies)"]
B -->|"10%"| C4["Consistency\n(6 strategies)"]
B -->|"5%"| C5["Alignment\n(3 strategies)"]
B -->|"5%"| C6["Data\n(6 strategies)"]
C1 --> D["Select strategy\n(uniform within category)"]
C2 --> D
C3 --> D
C4 --> D
C5 --> D
C6 --> D
D --> E["Apply mutation\nto *gguf.File"]
E --> F{"More mutations?\n(1-3 total)"}
F -->|yes| B
F -->|no| G["Serialize\n<code>gguf.Marshal(file)</code>"]
G --> H["Output mutated bytes\nto harness"] Why 1 to 3 mutations per test case?
Applying multiple mutations can satisfy compound preconditions -- for example, a mismatched tensor_count combined with a dimension-product overflow. It also makes attribution harder; use isolated-strategy experiments when assigning credit to a strategy.
Harness Integration¶
Crucible uses a split architecture where the mutation engine is written in Go but the fuzz harnesses target C/C++ parsers:
flowchart LR
subgraph "Go side"
Gen["crucible-gen\n(Go)"]
Mut["pkg/mutator\n(Go)"]
end
subgraph "Corpus"
Seeds["corpus/\n(GGUF files on disk)"]
end
subgraph "C side"
LF["libFuzzer harness\n(C + ASAN)"]
AFL["AFL++ harness\n(C + ASAN)"]
Target["llama.cpp\ngguf_init_from_file()"]
end
subgraph "Go side (native)"
GoFuzz["Go fuzz tests\n(FuzzGGUFReader)"]
GoTarget["pkg/gguf\nUnmarshal()"]
end
Gen --> Seeds
Mut --> Seeds
Seeds --> LF
Seeds --> AFL
LF --> Target
AFL --> Target
Seeds --> GoFuzz
GoFuzz --> GoTarget How it works¶
-
Seed generation:
crucible-genusespkg/corpusandpkg/mutatorto produce an initial corpus of structurally varied GGUF files, written tocorpus/ -
C harnesses (libFuzzer and AFL++): The fuzzer engine reads corpus files, mutates them, and feeds the result to the harness. The reference GGUF harness writes the input to a temporary file and calls
gguf_init_from_file()in a sanitizer-enabled llama.cpp build. Sanitizers expose the classes they instrument; silence from ASan does not prove the absence of uninitialized reads, logic defects, or other uninstrumented behavior. -
Go native harness: Go's built-in fuzzer calls
FuzzGGUFReaderwhich exercisesgguf.Unmarshaldirectly, andFuzzMutatorwhich verifies the mutation engine itself does not panic on arbitrary inputs.FuzzRoundTripchecks parse-serialize-parse consistency.
C harnesses require llama.cpp source
The libFuzzer and AFL++ harnesses link against llama.cpp's static libraries (libggml.a, libllama.a, and associated backend libraries). Set LLAMA_CPP to your local clone path:
The Go native harness has no external dependencies and works out of the box.
Target parsers¶
Full workflow (targets/ + triage)¶
These targets have dedicated Makefiles in targets/<name>/ that handle cloning, building with sanitizer instrumentation, and running campaigns:
| Target | Parser Function | Notes |
|---|---|---|
| llama.cpp | gguf_init_from_file(), grammar engines, RPC | Primary target |
| Ollama | gguf_init_from_file() (vendored) | Bundles its own fork of llama.cpp with custom patches |
See the Fuzzing llama.cpp and Fuzzing Ollama guides for end-to-end workflows.
Harness-only (libFuzzer binaries, no targets/ workflow yet)¶
These targets have libFuzzer harness source in harness/libfuzzer/ but no automated clone/build workflow in targets/:
| Target | Parser Function | Notes |
|---|---|---|
| whisper.cpp | whisper_init_from_buffer_with_params() | Audio model loader; shares gguf.cpp with llama.cpp |
| stable-diffusion.cpp | ModelLoader::init_from_file() | Multi-format model loader (GGUF, SafeTensors, ckpt) |
Build these manually by cloning the upstream repo and pointing the harness Makefile at the source tree.