Skip to content

Architecture

This page describes the high-level design of Crucible, the relationships between its packages, and the data flow from seed files through mutation and execution to evidence-backed human triage.

High-Level Data Flow

Crucible is two halves. The discovery half turns seeds into classified crashes. The evidence half turns a classified crash into a claim that survives someone else checking it. Both are part of the tool; a crash without evidence is not a finding, and a fix without evidence is not a fix.

flowchart LR
    subgraph DISC["Discovery"]
        direction TB
        Seeds["Seeds<br/>minimal GGUF files"]
        Corpus["Corpus<br/>generated + downloaded"]
        Mutator["Mutator<br/>structure-aware"]
        Harness["Harness<br/>libFuzzer / AFL++ / Go"]
        Target["Target parser"]
        Crashes["Execution artifacts<br/>diagnostics + process facts"]
        Triager["Observation + identity<br/>Exact dedup, Stable tracking"]

        Seeds --> Corpus --> Mutator --> Harness --> Target
        Target -->|crash| Crashes --> Triager
        Target -->|no crash| Harness
    end

    subgraph EVID["Evidence"]
        direction TB
        Tier["Evidence tier<br/>detorch / real-artifact / end-to-end"]
        Reach["Reachability<br/>compiled? linked? reached?"]
        Patch["Candidate fix"]
        Verify["Patch verification"]
        Tier --> Reach --> Patch --> Verify
    end

    Triager --> Tier
    Verify --> Gate{"Operator gates<br/>severity + disclosure"}
    Gate --> Out["Public issue / PR /<br/>advisory / CVE"]

    style DISC fill:#3b3028,stroke:#b08a5a,color:#fff
    style EVID fill:#3c432d,stroke:#9a9b62,color:#fff
    style Gate fill:#542f2b,stroke:#b96f5c,color:#fff

The two human gates -- exploitability tier and disclosure -- sit at the end by design and are never automated. Everything upstream of them may surface and count; nothing upstream of them may decide.

Artifact counts are not bug counts

A campaign writes thousands of crash artifacts for a handful of distinct bugs. The pipeline therefore never treats artifact volume as a bug count. Exact identities create build-local replay buckets; Stable identities track likely continuity across builds. Neither proves that two inputs share a source-level root cause, so every surviving observation is replayed and read by hand before it is named.

Package Relationships

Crucible is organized as format/mutation, execution/triage, protocol, telemetry, experiment, and evidence packages under pkg/, with command-line entry points under cmd/ and C-archive builds for custom mutators:

graph TD
    subgraph "cmd/"
        crucible["cmd/crucible\n(main orchestrator)"]
        gen["cmd/crucible-gen\n(corpus generator)"]
        triage_cmd["cmd/crucible-triage\n(crash triager)"]
        mutator["cmd/crucible-mutator\n(C-archive for libFuzzer)"]
        rpcmutator["cmd/crucible-rpc-mutator\n(RPC C-archive for libFuzzer)"]
    end

    subgraph "pkg/"
        gguf["pkg/gguf\n(format types + reader/writer)"]
        mut["pkg/mutator\n(GGUF mutation engine)"]
        rpcmut["pkg/mutator/rpc\n(RPC mutation engine)"]
        corpus["pkg/corpus\n(seed generation + management)"]
        triage["pkg/triage\n(crash dedup + reporting)"]
        coverage["pkg/coverage\n(LLVM coverage collection + reporting)"]
        experiment["pkg/experiment\n(replicated-run scoring contract)"]
        rpc["pkg/rpc\n(profiled ggml-rpc execution)"]
        mutstats["pkg/mutstats\n(strategy telemetry)"]
    end

    subgraph "harness/"
        libfuzzer["harness/libfuzzer\n(C harness)"]
        aflpp["harness/aflpp\n(C harness)"]
        gofuzz["harness/go\n(Go native fuzz tests)"]
    end

    crucible --> gguf
    crucible --> mut
    crucible --> corpus
    gen --> gguf
    gen --> corpus
    triage_cmd --> triage

    mut --> gguf
    rpcmut --> gguf
    corpus --> gguf
    corpus --> mut

    gofuzz --> gguf
    gofuzz --> mut

    libfuzzer -.->|"links against"| gguf
    aflpp -.->|"links against"| gguf

Dependency direction

The format packages remain below their mutators, while execution services are reused by CLI consumers. Consult go list -deps for the current exact graph; this diagram describes ownership, not a promise that packages such as triage have no internal dependencies.

The Validation Layer

The tools under tools/ answer different evidence questions. live and patchcheck assess target builds and candidate fixes; ledger generates canonical counts from repository records. They are additive to discovery and do not change pkg/ or the harnesses.

graph TD
    subgraph "tools/"
        LIVE["<b>tools/live</b><br/>crucible-live<br/><i>is this code compiled,<br/>linked, referenced?</i>"]
        PC["<b>tools/patchcheck</b><br/>crucible-patchcheck<br/><i>does this patch close the bug<br/>without breaking real input?</i>"]
        LED["<b>tools/ledger</b><br/>build-ledger<br/><i>canonical counts,<br/>generated not typed</i>"]
    end

    subgraph "inputs"
        BIN["Built artifact<br/>(binary, compile_commands,<br/>project file)"]
        PATCH["Candidate patch"]
        GEN["Genuine artifacts<br/>(real models / indexes)"]
        TREE["findings/ + reports/"]
    end

    BIN --> LIVE
    LIVE --> PC
    PATCH --> PC
    GEN --> PC
    TREE --> LED
    LED --> OUT["reports/generated/<br/>COUNTS.md + LEDGER.json"]

    style LIVE fill:#44372b,stroke:#c49a6c,color:#fff
    style PC fill:#3c432d,stroke:#9a9b62,color:#fff
    style LED fill:#4b362c,stroke:#a86f4c,color:#fff

live reports PRESENT / ABSENT / UNKNOWN, while patchcheck reports PASS / FAIL / UNKNOWN per check. UNKNOWN is a normal and correct answer; a weak negative must not silently become a verdict. Those assessment scripts leave the decision to the operator. build-ledger.py --check is different: it exits nonzero when its source binding or generated output has drifted, and the repository uses it as a correctness gate.

See Evidence and Validation for the evidence ladder, what each tier does and does not establish, and why the remedy carries a higher bar than the finding.

The correctness gates

The assessment scripts above report evidence for an operator to judge. Correctness gates refuse to proceed when a required condition fails; their exit codes are part of the interface.

gate refuses when
crucible run --sift a seed already crashes the harness, so libFuzzer would die during corpus replay
crucible run --supervise the harness thrashes, or no new sanitizer signature has appeared for N restarts
crucible provenance --strict the target checkout has modified tracked files, so the witness does not describe that commit
tools/check-evidence.py a named bundle is missing or a selected sanitizer fragment appears nowhere in its named bundle set; exact frames, lines, artifacts, and builds remain unchecked
tools/ledger/build-ledger.py --check committed source bindings or generated finding counts have drifted
tools/check-findings-yaml.py the authoritative state file does not parse, or has duplicate mapping keys
tools/check-locked-artifacts.py a locked reference manifest has been edited
tools/diagrams/check.sh a diagram no longer matches the code it claims to describe

The rule they share is CLAUDE.md #8: a tool whose failure branch returns something that looks like an answer will silently corrupt every result downstream. So "cannot tell" gets its own exit code (2), separate from both success and failure. check-evidence.py exits 2 when the evidence tree is absent on this machine, because "I cannot see the artifacts" is not "the artifacts are fine".

The generated diagram Correctness gates renders this pipeline from the code, including each gate's exit contract, so it cannot drift.

Operational detail and the failure each gate retires: Campaign operations.

Machine-readable results

Many result-bearing commands expose --json; check each command's help for support. The core triage and run envelopes carry a crucible build stamp and report incomplete observations explicitly. Other commands use their own schemas, and preflight reports through text and exit status. Consumers must distinguish "we looked and found nothing" from "we could not look". Field reference: JSON result shapes.

Directory Structure

crucible/
├── cmd/
│   ├── crucible/              # Main orchestrator binary
│   │   ├── main.go
│   │   ├── sift.go            # --sift: move already-crashing seeds out before launch
│   │   ├── supervise.go       # --supervise: restart past crashes, thrash + saturation guards
│   │   ├── provenance.go      # commit AND dirty state of the target checkout
│   │   ├── json_results.go    # the --json wire shapes (docs/reference/json-results.md)
│   │   ├── harness_smoke.go   # clean / crash / did-not-run before a campaign
│   │   └── mutate_stats_ab.go # A/B a mutation weighting against a measured baseline
│   ├── crucible-gen/          # Corpus generation tool
│   ├── crucible-mutator/      # Custom mutator (C-archive for libFuzzer)
│   ├── crucible-mutator-traced/     # Same, with per-strategy tracing for yield measurement
│   ├── crucible-rpc-mutator/        # ggml RPC wire-format mutator
│   ├── crucible-flatbuffer-mutator/ # FlatBuffers structure-aware mutator (TFLite, MNN)
│   ├── crucible-protobuf-mutator/   # Protobuf structure-aware mutator (ONNX, TF)
│   ├── crucible-safetensors-mutator/ # safetensors structure-aware mutator (fastllm, mlx, candle)
│   └── crucible-triage/       # Crash triage and reporting
│       ├── main.go
│       └── watch.go           # Watch-mode with checkpoint persistence
├── pkg/
│   ├── gguf/                  # GGUF format implementation
│   │   ├── format.go          # Types: Header, MetadataKV, TensorInfo, File
│   │   ├── format_test.go     # Format unit tests
│   │   ├── reader.go          # Binary deserialization (Unmarshal)
│   │   ├── reader_test.go     # Reader unit tests
│   │   └── writer.go          # Binary serialization (Marshal)
│   ├── mutator/               # Structure-aware mutation engine
│   │   ├── mutator.go         # Orchestrator: weighted category selection
│   │   ├── mutator_test.go    # Mutation engine tests
│   │   ├── header.go          # 5 header mutation strategies
│   │   ├── metadata.go        # 13 metadata mutation strategies
│   │   ├── tensorinfo.go      # 8 tensor info mutation strategies
│   │   ├── alignment.go       # 3 alignment mutation strategies
│   │   ├── data.go            # 6 tensor data mutation strategies
│   │   ├── consistency.go     # 6 cross-field consistency strategies
│   │   ├── model_loader.go    # 5 model-loader strategies (weighted under metadata)
│   │   ├── synth/             # Synthesised whole-file generation
│   │   ├── rpc/               # ggml RPC wire messages, incl. rpc_tensor field fuzzing
│   │   ├── flatbuffer/        # FlatBuffers: vtable, offset and vector-length mutation
│   │   ├── protobuf/          # Protobuf: wire-type, varint and length-delimited mutation
│   │   └── safetensors/      # safetensors: header length, data_offsets, shape, dtype
│   ├── corpus/                # Corpus generation and management
│   │   ├── corpus.go          # Corpus loading and enumeration
│   │   ├── corpus_test.go     # Corpus unit tests
│   │   ├── generate.go        # Seed file generation
│   │   ├── minimize.go        # Corpus minimization
│   │   └── minimize_test.go   # Minimization tests
│   └── triage/                # Crash analysis and reporting
│       ├── triage.go          # Compatibility crash taxonomy and Exact-keyed Triager
│       ├── triage_test.go     # Triage unit tests
│       ├── cwe.go             # CWE identifier mapping for each crash type
│       ├── stackhash.go       # Frozen legacy five-frame oracle hash
│       ├── identity.go        # Exact build-local and Stable cross-build identities
│       ├── observation.go     # Admission + text evidence + process facts
│       ├── stackhash_test.go  # Stack hash tests
│       ├── minimize.go        # Crash reproducer minimization (recursive WalkDir)
│       ├── minimize_test.go   # Minimize tests
│       ├── replay.go          # Crash replay against harness binaries
│       ├── replay_test.go     # Replay unit tests
│       ├── report.go          # Unrated internal report generation
│       ├── report_target_test.go  # Report target detection tests
│       ├── sarif.go           # SARIF 2.1.0 output with target tags
│       ├── sarif_test.go      # SARIF output tests
│       └── fixture_test.go    # Shared test fixtures
│   ├── coverage/              # LLVM coverage collection
│   │   └── coverage.go        # Replay corpus, merge profraw, HTML report
│   ├── mutstats/              # Measured per-strategy yield, and A/B against a baseline
│   ├── experiment/            # Scoring contract for completed replicated runs
│   ├── rpc/                   # ggml RPC wire model (GRAPH_COMPUTE et al)
│   ├── oracle/                # Known-signature lookup for triage
│   └── buildinfo/             # Build stamp carried by every JSON result
├── harness/
│   ├── libfuzzer/             # libFuzzer C++ harnesses
│   │   ├── harness.cpp        # LLVMFuzzerTestOneInput targeting gguf_init_from_file
│   │   └── Makefile
│   ├── aflpp/                 # AFL++ C harness
│   │   ├── harness.c
│   │   └── Makefile
│   └── go/                    # Go native fuzz tests
│       └── fuzz_test.go       # FuzzGGUFReader, FuzzMutator, FuzzRoundTrip
├── corpus/                    # Seed corpus directory
│   └── gguf.dict              # GGUF-specific dictionary for fuzzer guidance
├── crashes/                   # Crash artifacts output directory
├── targets/                   # Build configs for fuzz targets
│   ├── llamacpp/Makefile
│   └── ollama/Makefile
├── tools/
│   ├── check-evidence.py      # narrow gate: paths, hash syntax, coarse sanitizer fragments
│   ├── check-findings-yaml.py # gate: the authoritative state file parses, no duplicate keys
│   ├── check-locked-artifacts.py  # gate: locked reference manifests are unchanged
│   ├── ledger/                # generates reports/generated/{COUNTS.md,LEDGER.json}
│   ├── live/                  # crucible-live: is this code compiled, linked, referenced?
│   ├── patchcheck/            # does this patch close the bug without breaking real input?
│   ├── triage-db/             # operational known-signature database (free to grow)
│   ├── oracle/                # LOCKED reference manifests; see the locked-artifact gate
│   ├── diagrams/              # ground-truth extractor; check.sh gates diagram drift
│   └── sweep/                 # campaign shell drivers
├── .github/
│   └── workflows/
│       ├── ci.yml             # CI pipeline: lint, test, build
│       └── docs.yml           # Documentation build and deploy
├── scripts/
│   ├── build-linux.sh         # Cross-compile harnesses for Linux
│   ├── build-targets.sh       # Build all target parsers
│   ├── download-corpus.sh     # Download seed corpus from model repos
│   ├── gen-arch-seeds.py      # Generate per-architecture GGUF seeds
│   ├── gen-lora-seeds.py      # Generate LoRA-specific GGUF seeds
│   ├── gen-server-seeds.py    # Generate HTTP endpoint test seeds
│   ├── gen-whisper-audio-seeds.py  # Generate PCM audio test seeds
│   ├── generate-targeted-seeds.py  # Generate TALOS CVE-targeted seeds
│   ├── launch-phase2.sh       # Launch Phase 2 campaign (multi-harness)
│   ├── run-all-campaigns.sh   # Launch all campaigns in parallel
│   ├── run-campaign.sh        # Single campaign launcher
│   └── rotate-logs.sh         # Rotate campaign logs
├── docs/                      # mkdocs-material documentation source
├── Makefile                   # Top-level build, test, fuzz, triage targets
├── mkdocs.yml                 # Documentation site configuration
├── FUZZING-ROADMAP.md         # Living checklist of targets, campaigns, and findings
├── VERSIONS.env               # Pinned target versions for CI reproducibility
├── go.mod
└── go.sum

Package Details

pkg/gguf -- Format Implementation

The foundation package. It defines the Go types that mirror the GGUF binary specification:

  • Header -- 4-byte magic (GGUF), version (uint32), tensor count (uint64), metadata KV count (uint64)
  • MetadataKV -- key string + one of 13 metadata value kinds, including STRING and ARRAY
  • TensorInfo -- tensor name, dimensions, ggml_type enum, byte offset into the data section
  • File -- the complete in-memory representation of a GGUF file

The package provides Unmarshal([]byte) (*File, error) for parsing and Marshal(*File) ([]byte, error) for serialization. These form the round-trip pipeline that the mutation engine depends on.

Why not modify bytes directly?

Byte-level mutations are fast but structurally blind. Operating on parsed structures lets Crucible target semantic fields and cross-field relationships directly. It improves the chance of passing selected parse gates; it does not guarantee that every mutation parses or reaches a deep path.

pkg/mutator -- Mutation Engine

The GGUF mutation engine applies weighted structure-aware strategies across header, metadata, tensor, alignment, data, consistency, and loader semantics. The generated format topology is the maintained inventory. On each call to Mutate(), it:

  1. Randomly selects 1 to 3 mutations to apply
  2. For each mutation, picks a category using weighted random selection
  3. Picks a strategy uniformly at random within the chosen category
  4. Applies the strategy to the in-memory *gguf.File
  5. Serializes the result back to bytes via gguf.MarshalRaw, preserving intentional count or padding mismatches when the strategy requests them

Each strategy file (header.go, metadata.go, tensorinfo.go, alignment.go, data.go, consistency.go, model_loader.go) exports a *Strategies() []Strategy function. The Mutator registers all of them at construction time.

The Strategy interface is intentionally minimal:

type Strategy interface {
    Name() string
    Mutate(f *gguf.File, rng *rand.Rand) // math/rand/v2
}

This makes it straightforward to add new strategies -- implement the interface, add it to the appropriate *Strategies() function, and it is automatically picked up by the engine.

pkg/corpus -- Seed Management

Handles three responsibilities:

  1. Generation -- crucible-gen produces structurally varied seed GGUF files covering different metadata types, tensor configurations, and alignment values
  2. Loading -- reads seed files from disk for the fuzzing harnesses
  3. Minimization -- content-deduplicates entries or, with a usable LLVM coverage harness, applies greedy set cover while preserving entries that produced no coverage data

pkg/triage -- Crash Analysis

The triage package is format-agnostic. Its current pipeline:

  1. Admission -- requires explicit metadata before an artifact can carry an outcome.
  2. Observation -- combines text evidence with process facts; neither substitutes for the other.
  3. Identity -- Exact keys build-local buckets, Stable tracks normalized target paths across builds, and legacy HashStack remains frozen for historical oracle compatibility.
  4. Attribution -- filters sanitizer, fuzzer, harness, runtime, and compiler-library frames; no target frame yields an explicit UNATTRIBUTED result.
  5. Minimization -- preserves the relevant Exact identity rather than accepting any crash.
  6. Report generation -- writes unrated internal records. CVSS and affected versions are absent until supplied from operator-ratified evidence.
  7. SARIF export -- transports taxonomy and identities without upgrading an observation into a confirmed vulnerability.

Mutation Pipeline

The following diagram shows the detailed flow when a single mutated test case is produced:

flowchart TD
    A["Parse seed file\n<code>gguf.Unmarshal(bytes)</code>"] --> B{"Select category\n(weighted random)"}

    B -->|"35%"| C1["Metadata + Model-Loader\n(13 + 5 = 18 strategies)"]
    B -->|"35%"| C2["Tensor Info\n(8 strategies)"]
    B -->|"10%"| C3["Header\n(5 strategies)"]
    B -->|"10%"| C4["Consistency\n(6 strategies)"]
    B -->|"5%"| C5["Alignment\n(3 strategies)"]
    B -->|"5%"| C6["Data\n(6 strategies)"]

    C1 --> D["Select strategy\n(uniform within category)"]
    C2 --> D
    C3 --> D
    C4 --> D
    C5 --> D
    C6 --> D

    D --> E["Apply mutation\nto *gguf.File"]
    E --> F{"More mutations?\n(1-3 total)"}
    F -->|yes| B
    F -->|no| G["Serialize\n<code>gguf.Marshal(file)</code>"]
    G --> H["Output mutated bytes\nto harness"]

Why 1 to 3 mutations per test case?

Applying multiple mutations can satisfy compound preconditions -- for example, a mismatched tensor_count combined with a dimension-product overflow. It also makes attribution harder; use isolated-strategy experiments when assigning credit to a strategy.

Harness Integration

Crucible uses a split architecture where the mutation engine is written in Go but the fuzz harnesses target C/C++ parsers:

flowchart LR
    subgraph "Go side"
        Gen["crucible-gen\n(Go)"]
        Mut["pkg/mutator\n(Go)"]
    end

    subgraph "Corpus"
        Seeds["corpus/\n(GGUF files on disk)"]
    end

    subgraph "C side"
        LF["libFuzzer harness\n(C + ASAN)"]
        AFL["AFL++ harness\n(C + ASAN)"]
        Target["llama.cpp\ngguf_init_from_file()"]
    end

    subgraph "Go side (native)"
        GoFuzz["Go fuzz tests\n(FuzzGGUFReader)"]
        GoTarget["pkg/gguf\nUnmarshal()"]
    end

    Gen --> Seeds
    Mut --> Seeds
    Seeds --> LF
    Seeds --> AFL
    LF --> Target
    AFL --> Target
    Seeds --> GoFuzz
    GoFuzz --> GoTarget

How it works

  1. Seed generation: crucible-gen uses pkg/corpus and pkg/mutator to produce an initial corpus of structurally varied GGUF files, written to corpus/

  2. C harnesses (libFuzzer and AFL++): The fuzzer engine reads corpus files, mutates them, and feeds the result to the harness. The reference GGUF harness writes the input to a temporary file and calls gguf_init_from_file() in a sanitizer-enabled llama.cpp build. Sanitizers expose the classes they instrument; silence from ASan does not prove the absence of uninitialized reads, logic defects, or other uninstrumented behavior.

  3. Go native harness: Go's built-in fuzzer calls FuzzGGUFReader which exercises gguf.Unmarshal directly, and FuzzMutator which verifies the mutation engine itself does not panic on arbitrary inputs. FuzzRoundTrip checks parse-serialize-parse consistency.

C harnesses require llama.cpp source

The libFuzzer and AFL++ harnesses link against llama.cpp's static libraries (libggml.a, libllama.a, and associated backend libraries). Set LLAMA_CPP to your local clone path:

make harness-libfuzzer LLAMA_CPP=~/src/llama.cpp

The Go native harness has no external dependencies and works out of the box.

Target parsers

Full workflow (targets/ + triage)

These targets have dedicated Makefiles in targets/<name>/ that handle cloning, building with sanitizer instrumentation, and running campaigns:

Target Parser Function Notes
llama.cpp gguf_init_from_file(), grammar engines, RPC Primary target
Ollama gguf_init_from_file() (vendored) Bundles its own fork of llama.cpp with custom patches

See the Fuzzing llama.cpp and Fuzzing Ollama guides for end-to-end workflows.

Harness-only (libFuzzer binaries, no targets/ workflow yet)

These targets have libFuzzer harness source in harness/libfuzzer/ but no automated clone/build workflow in targets/:

Target Parser Function Notes
whisper.cpp whisper_init_from_buffer_with_params() Audio model loader; shares gguf.cpp with llama.cpp
stable-diffusion.cpp ModelLoader::init_from_file() Multi-format model loader (GGUF, SafeTensors, ckpt)

Build these manually by cloning the upstream repo and pointing the harness Makefile at the source tree.