Capability matrix¶
Current as of 2026-09-01.
The central rule is simple:
A built binary, a green command, and a pile of artifacts are not evidence of capability unless the relevant control ran and could have contradicted the result.
This page describes supported behavior and its limits. The generated format topology is the maintained inventory for formats, mutators, and harnesses.
How to read this page¶
Every capability below names the command or package that provides it, so a claim can be checked against committed source rather than taken on trust. Support labels are defined in Installation.
| Label | Applies here to |
|---|---|
| Stable | Anything reachable from crucible, crucible-gen or crucible-triage after make build |
| Advanced | Stages that assume a campaign is already running and evidence already exists: saturation, escape safety, analysis and live regression, path preservation |
| Target-dependent | Every stage from harness lifecycle onward: each needs a compiled native harness for the code under test |
| Not publicly shipped | Work in the tree that is not part of these three binaries. It is not listed on this page |
The last row matters as much as the others. This page describes the tool that make build produces, not everything the repository contains.
End-to-end workflow¶
| Stage | Capability | Check |
|---|---|---|
| environment | locate required compilers, runtimes, targets, and harnesses | crucible doctor |
| provenance | bind evidence to target commit and dirty state | crucible provenance --strict |
| prior-state review | find existing findings, leads, submissions, and closures by surface | crucible recon |
| harness lifecycle | execute valid/empty inputs and optional reach marker | crucible harness-smoke |
| campaign capability | clean control, retained-positive visibility, and corpus reach | crucible preflight |
| discovery | structure-aware and byte mutation through native target harnesses | crucible run / campaign |
| saturation | quantify one-identity domination | crucible saturation |
| escape safety | prove suppression, reversibility, and continued visibility | crucible escape-verify |
| triage | combine admission, text evidence, and process facts | crucible triage |
| identity | Exact build-local buckets and Stable cross-build tracking | pkg/triage.Identify, tests TestExactNamespaceIsPublished and TestStableSurvivesWhatExactDoesNot |
| analysis regression | compare admitted banked text under current pure analysis | crucible regress |
| live regression | replay retained PoCs against a built harness | validate-oracle |
| successful-output differential | compare bound .npy metadata and values, including conversion round trips, under explicit numerical policies | crucible value-diff / value-matrix / value-roundtrip |
| path preservation | capture/adjudicate/compare migrated consumers | capability-capture / capability |
Formats and mutation engines¶
Crucible has native mutation engines for:
| Engine | Structural surface |
|---|---|
pkg/mutator | GGUF header, metadata, tensors, offsets, alignment, consistency, recombination |
pkg/mutator/rpc | profiled ggml-rpc commands, graph payloads, tensor descriptors, and stateful sequences |
pkg/mutator/protobuf | protobuf wire keys, varints, fixed fields, and length-delimited values |
pkg/mutator/flatbuffer | roots, vtables, tables, offsets, and vectors |
pkg/mutator/safetensors | header length, JSON tensor metadata, shapes, dtypes, and data offsets |
pkg/mutator/npy | NumPy header, dtype, rank, shape, and payload relationships |
pkg/mutator/tokenizer | tokenizer model structures used by the supported harnesses |
pkg/mutator/synth | declarative schema-to-mutation specifications |
The custom mutator replaces libFuzzer's default mutation entry point. When a structure-aware engine declines an input, the C-archive shim must call LLVMFuzzerMutate; returning the input unchanged is a no-op, not fallback. make verify-mutators requires the expected custom-mutator symbol to be defined in built mutator harnesses and exits 2 when there are no binaries to inspect.
Limits¶
- Structure-aware output still needs target execution; parseability is not reachability.
- FlatBuffers are not self-describing, so generic graph discovery is heuristic.
- Protobuf wire validity does not imply schema validity.
- A linked mutator does not prove its strategies are effective; use
mutate-statsand controlled experiments.
Harness families¶
In-depth libFuzzer harnesses¶
harness/libfuzzer/ covers GGUF parsing and writing, model loading, LoRA, vision, whisper, stable diffusion, safetensors, MLX, Torch/TorchScript, ONNX, TFLite, grammar, JSON Schema, Jinja/chat templates, tokenizer/vocabulary, Unicode/regex, quantization, and ggml-rpc.
The RPC family includes raw protocol, command-aware, graph-compute, and two-connection/race surfaces. It requires a llama.cpp build with RPC enabled and the matching RPC archive.
External loader harnesses¶
harness/cpp/ contains additive native loaders for projects including MNN, ONNX Runtime, ExecuTorch, TFLite, ncnn, CTranslate2, pocketsphinx, TensorRT-LLM, fastllm, MLX, Tengine, SGLang/MoE, whisper audio, and Vowpal Wabbit.
These builds differ in depth. A direct parser detorch, a public loader, and a natural inference path are different evidence tiers and must be labeled as such.
Harness capability manifests¶
Some replay harnesses intercept assertions or target faults and intentionally survive. A sidecar manifest can declare intercepts_faults, but only when bound to the harness SHA-256. The declaration becomes invalid after a rebuild. Without it, assertion text from a surviving process is diagnostic.
Stateful RPC execution¶
The RPC layer can:
- derive the wire profile from a pinned target build instead of assuming a copied struct layout;
- launch and own a stock server with bounded startup/teardown and descendant cleanup;
- verify socket ownership and fail closed when ownership cannot be established;
- run typed command sequences with independent servers for A/B arms;
- perform real SET/GET round trips with exact byte comparison;
- test cross-client allocation reuse with a victim secret and zero-on-allocation control;
- exercise use-after-free, double-free, fabricated-handle, reorder, cross-client, and read-before-write strategies;
- compare graph requests after profile-derived normalization of address-bearing fields.
Limits¶
- An in-process stub cannot validate stock protocol semantics.
- A clean stateful campaign is scoped to the pinned commit, build, platform, sanitizer set, and strategies run.
- ASan does not detect uninitialized reads; a disclosure oracle needs exact victim/attacker byte comparison or MemorySanitizer-equivalent evidence.
- Per-connection server state is not a cross-principal boundary unless the target architecture establishes one.
Execution observation¶
The current model is:
It distinguishes crash, suspect, intercepted fault, diagnostic, clean, ambiguous, incomplete, and unknown. Only explicit admission can produce a substantive classification. A log from another execution cannot provide this run's exit, signal, timeout, binary, or environment facts.
NeedsHumanTriage() surfaces crash, suspect, and intercepted-fault observations. It does not mean confirmed, exploitable, reportable, or disclosure-ready.
Successful-output comparison¶
value-diff covers the case the crash oracle cannot see: both executions return successfully but their tensors disagree. It compares two complete NumPy .npy outputs and binds the exact bytes it read with SHA-256. Shape, dtype descriptor, and Fortran-order metadata are exact. Integer and boolean values are exact; floating-point values use the recorded atol, rtol, NaN, and signed-zero policy.
The result is MATCH, DIVERGED, or UNKNOWN. Malformed, truncated, oversized, ambiguous, and unsupported arrays are UNKNOWN, never clean matches. The implementation is verified against NumPy-generated float32 files plus adversarial synthetic format cases.
value-matrix applies that oracle to a hash-bound JSON manifest of reference/candidate pairs. Each pair has its own exact input identities, numerical policy, and resource limits. A missing, substituted, or unreadable output produces UNKNOWN; UNKNOWN dominates the aggregate receipt without erasing any measured divergences in the per-pair counts. Manifest-selected per-input and aggregate work are hard-capped before pair I/O begins.
value-roundtrip verifies a source model, converter identity and invocation receipt, converted artifacts, three runtime identities and invocation receipts, and the exact reference, direct-load, and converted/reloaded outputs. Only after every retained file and identity binding verifies does it compare reference versus direct, reference versus reloaded, and direct versus reloaded. This makes a conversion-only divergence visible without confusing it with a general runtime disagreement.
Limits¶
- These commands compare outputs already produced by executions; they do not launch a converter or target.
- Supported scalar dtypes are boolean, 8/16/32/64-bit integers, and 16/32/64-bit IEEE floats. Complex, object, string, structured, datetime, and bfloat formats return
UNKNOWN. - A divergence is evidence requiring an invariant and impact analysis. It is not automatically a bug, vulnerability, finding, severity, or disclosure decision.
Identity and attribution¶
| Identity | Use |
|---|---|
| Exact | build-local campaign buckets, watcher names, minimization, and differential comparison |
| Stable | cross-machine/build tracking using normalized target functions and repo-relative paths |
legacy HashStack | frozen oracle and historical evidence compatibility only |
Exact and Stable are additive; the legacy hash remains unchanged. Attribution filters sanitizer, fuzzer, harness, runtime, and compiler-standard-library frames. No surviving target frame produces an explicit UNATTRIBUTED result.
Reports and exports¶
Generated reports are internal investigation records:
- severity is
UNRATED; - affected versions are unknown;
- CVSS is absent unless an operator supplies a ratified vector;
- CWE is a taxonomy hint pending source validation;
- SARIF levels come from memory-safety/crash-only/resource-exhaustion taxonomy, not CVSS.
Crucible never sends a disclosure. Severity ratification and disclosure remain human gates.
Regression corpus and capability oracle¶
regress reads an explicit admission manifest. It compares text evidence, taxonomy, attribution, and both current identities. It produces an outcome only when hash-bound process facts exist.
- exit 0: admitted evidence matches;
- exit 1: analysis drift;
- exit 2: cannot tell;
- exit 3: unadjudicated evidence remains.
capability-capture banks immutable raw observations and always exits 2. A human must adjudicate them before capability can compare locked positives. A path that did not run or cannot produce a required field is incomplete, not preserved.
Mutation telemetry and experiments¶
mutate-stats records strategy application, solo application, structural novelty, coverage co-occurrence, and replayed Exact identities. Multi-strategy credit is co-occurrence, not causal proof.
pkg/experiment validates and scores completed replicated run records. It is currently an analysis contract, not an end-to-end runner: it does not launch campaigns, transactionally bank evidence, bind binary/corpus/environment hashes, or adjudicate root causes. Its time-to-first summary must not be described as survival analysis for censored runs.
Correctness gates¶
make gates runs, in order:
- custom-mutator linkage;
- findings schema parsing;
- locked-artifact integrity;
- named evidence paths, hash-file syntax, and coarse sanitizer-fragment presence (not exact trace, frame, line, or build provenance);
- admitted analysis regression;
- generated diagram ground truth;
- generated public metrics and CVE catalog;
- source-bound CVE visual synchronization;
- public command, API, and confidence-contract audit;
- Go vet and tests;
- stock RPC integration when configured.
Required local integrations that cannot run make the gate incomplete and non-zero. CI cannot prove what its environment does not contain.
What Crucible cannot claim automatically¶
No command, identity, or generated report establishes these by itself:
- that a crash is a vulnerability;
- read versus write when the evidence does not say;
- code execution, privilege gain, or exploitability;
- CVSS vector or severity;
- affected product versions;
- that two signatures share one root cause;
- that a finding remains live on current upstream;
- that a quiet campaign exhausted a surface;
- that a patch is complete across sibling entry points;
- that a disclosure is appropriate or ready to send.
Those require replay, source inspection, clean HEAD builds, controls, and human judgment.