crucible¶
The main CLI covers environment validation, corpus generation, campaign execution, triage, regression checks, and capability preservation.
JSON support varies by command; check crucible <command> --help before building an integration. preflight reports through text and exit status. Treat documented non-zero exit codes as typed outcomes: several commands use exit 2 for cannot tell, and regress uses exit 3 for an incomplete admission corpus.
Command index¶
| Command | Purpose |
|---|---|
doctor | Check the local build and fuzzing environment |
provenance | Record target commit, date, upstream distance, and dirty state |
recon | Read prior findings, leads, submissions, and surface closures before spending a machine |
generate | Generate and mutate GGUF corpus entries |
mutate | Apply structure-aware GGUF mutations to one file |
protobuf | Apply one traceable, structure-aware protobuf mutation |
synth | Validate, list, and apply schema-synthesized mutation specifications |
gguf | Inspect, parse-check, and reserialize GGUF artifacts |
harness-smoke | Test harness lifecycle, a valid seed, empty input, and an optional reach marker |
preflight | Test control cleanliness, retained-positive visibility, and corpus reach |
run | Launch libFuzzer or AFL campaigns |
campaign | Run campaign definitions stored as data |
status | Report campaign artifacts and status |
saturation | Measure concentration by build-local crash identity |
escape-verify | Verify a saturation escape without silently blinding the harness |
triage | Replay candidate inputs and produce observations and internal reports |
triage dedup | Deduplicate inputs by Exact identity or content fingerprint |
report | Alias of the report-producing triage path |
minimize | Minimize a corpus or reproducer while preserving the relevant behavior |
oracle | Replay retained PoCs and print the legacy compatibility stack hash |
validate-oracle | Assert locked legacy oracle entries reproduce |
regress | Compare admitted banked text under the current pure analysis layer |
capability-capture | Capture immutable cross-path observations for human adjudication |
capability | Compare adjudicated positives and candidates across migrated paths |
value-diff | Compare successful .npy outputs under an explicit numerical policy |
value-matrix | Compare a hash-bound matrix of successful .npy outputs |
value-roundtrip | Verify and compare a bound model-conversion round trip |
coverage | Collect and report source coverage |
mutate-stats | Capture, join, report, and compare per-strategy mutation telemetry |
measurement | Validate a local, zero-authority measurement envelope and correction chain |
workspace | Inspect the disposable, source-attributed read-only operator index |
Shell completion is available through crucible completion.
Recommended sequence¶
crucible doctor
crucible provenance --strict /path/to/target
crucible recon <target> "<surface>"
crucible harness-smoke --harness ./harness --seed ./control
crucible preflight --harness ./harness --control ./control \
--known-crash ./retained-positive --corpus ./corpus
crucible run --harness ./harness --corpus ./corpus --output ./crashes -- \
-max_total_time=3600 -seed=12345
crucible triage --artifact-kind input --crashes ./crashes \
--harness ./harness --output ./reports --json
workspace¶
Full surface, flags, selector grammar and exit codes: crucible workspace.
crucible workspace list --repo-root . --json
crucible workspace list --repo-root . --evidence-root /path/to/crucible-evidence
crucible workspace show ID_FROM_LIST --repo-root .
crucible workspace map --repo-root . --evidence-root /path/to/crucible-evidence
crucible workspace verify evidence:<evidence-id>
crucible workspace verify submission:<case-id>
The index presents source-attributed records from the operator's local repository. Every resolved value is KNOWN, UNKNOWN, CONFLICT, or NOT_APPLICABLE and cites its owning source. Conflicts remain visible and fail closed. show follows one workspace through its studies, attempts, evidence, findings and submission cases. map renders that same normalized graph for every imported workspace; neither command introduces another schema or state store.
The human views show unresolved decisions, operator gates, manifest inspection, historical verification, current-host verification and the identity and availability of each owning source. They do not collapse these fields into one status. UNKNOWN and CONFLICT remain explicit, with their reason and source attribution.
With --json, show emits the same normalized Index schema filtered to the selected workspace and its transitively linked studies, attempts, evidence, findings, submission cases, diagnostics and sources. It does not emit a bare workspace that leaves source IDs or operator gates unresolved.
All three commands read canonical repository records, open each repo: and banked: component relative to pinned directory descriptors without following symlinks, and read only exact manifests or receipts capped at 8 MiB. They never recursively walk or hash an evidence payload. A manifest digest match is metadata, not current-host full verification.
These commands write nothing and grant no target-selection, execution, finding, severity, scheduler, policy-update, submission, or disclosure authority. READY_FOR_OPERATOR is not approval and is not evidence that a packet was sent.
doctor¶
Checks the Go toolchain, sanitizer-capable compiler, target paths, and built harnesses.
It exits non-zero when a required component is missing. A build artifact existing on disk is not enough; target-specific preflight still follows.
provenance¶
--strict rejects tracked changes in the target checkout. Store this output with every witness: the commit alone does not describe a dirty build.
recon¶
| Exit | Meaning |
|---|---|
0 | no recorded closure covers the named surface |
1 | a recorded closure covers it |
2 | records could not be read |
3 | closures exist but no surface was supplied |
recon prevents spending campaign time on a negative or fixed surface that the project has already recorded. It does not independently establish that the record is still correct.
generate, mutate, protobuf, synth, and gguf¶
crucible generate --corpus ./seeds --output ./generated --count 100 --seed 42
crucible mutate --help
crucible protobuf mutate seed.onnx mutated.onnx --seed 42 --json
crucible synth validate rule.synth.json
crucible synth list rule.synth.json
crucible synth apply seed.gguf rule.synth.json out.gguf --rule <rule-name>
crucible gguf inspect model.gguf
crucible gguf parse model.gguf
crucible gguf serialize model.gguf roundtrip.gguf
protobuf mutate refuses an existing output, preserves protobuf wire structure, and emits the selected nested field path and mutation action. Its trace carries authority_effect: NONE; it is attribution for a later execution, not proof that the consumer accepted the artifact.
Use the generated or mutated artifact as an input to a real target harness. A successful structural mutation is not evidence that the target accepted it or reached interesting code.
harness-smoke¶
crucible harness-smoke \
--harness ./crucible-libfuzzer-model \
--seed ./valid.gguf \
--reach-marker load_model \
--timeout 30s
Smoke tests execution, a valid seed, empty input, and optionally a target marker. It catches input-independent harness faults before they flood a campaign.
preflight¶
crucible preflight \
--harness ./crucible-libfuzzer-model \
--control ./valid.gguf \
--known-crash ./retained-poc.gguf \
--corpus ./pristine-corpus
| Exit | Meaning |
|---|---|
0 | the requested checks establish a meaningful starting state |
1 | a control or capability check failed |
2 | the question could not be answered |
The corpus check detects a corpus that adds almost nothing. It does not prove every seed reaches the target. Without --known-crash, the harness-blindness check is absent.
run¶
crucible run \
--harness ./crucible-libfuzzer-model \
--corpus ./corpus \
--output ./crashes \
--jobs 8 \
--sift \
--supervise \
--fuzz-env ASAN_OPTIONS=detect_leaks=0 \
--json \
-- -max_total_time=3600 -rss_limit_mb=0
Anything after -- is appended to the engine arguments and can override defaults. Use repeated --fuzz-env KEY=VALUE for campaign environment; it is recorded in JSON. --sift moves already crashing seeds aside before launch. --supervise restarts past crashes and can halt on stale signatures; that halt signals saturation, not a vulnerability verdict.
campaign¶
Campaign definitions are data. list reports what a config declares; run executes one or more of them by name, or all of them when no name is given.
crucible campaign list --config campaigns.yaml
crucible campaign run --config campaigns.yaml --dry-run
crucible campaign run model-loader --config campaigns.yaml
campaign list¶
| Flag | Default | Meaning |
|---|---|---|
--config | campaigns.yaml | Path to the campaigns config file |
--json | false | Emit machine-readable JSON |
campaign run¶
| Flag | Default | Meaning |
|---|---|---|
--config | campaigns.yaml | Path to the campaigns config file |
--dry-run | false | Print the commands without executing them |
saturation and escape-verify¶
crucible saturation --harness ./harness --artifacts ./crashes --threshold 0.90
crucible escape-verify \
--harness ./harness \
--control ./control \
--escaped-input ./dominant-poc \
--escape-off CRUCIBLE_ESCAPE_OFF=1 \
--expect-stable <stable-id> \
--known-crash ./different-poc \
--expect-known-stable <different-stable-id> \
--receipt ./escape-receipt.json
saturation exits 1 when one identity reaches the threshold and 2 when it cannot tell. It never chooses a filter. escape-verify requires a clean control, suppression, reversible recovery of the same stable/class/site, and continued visibility of a different known bug. Exit 2 is incomplete.
triage and report¶
crucible triage \
--artifact-kind input \
--crashes ./crashes \
--harness ./harness \
--target gguf-loader \
--output ./reports \
--sarif ./reports/results.sarif \
--json
Important flags:
| Flag | Meaning |
|---|---|
--artifact-kind input | replay binary reproducers and collect process facts |
--artifact-kind banked-log | identify logs from another execution; skip them rather than infer process facts |
--replay-env KEY=VALUE | add a replay environment value; repeatable |
--replay-timeout | bound each replay |
--sibling label=/path | replay through another harness and record differential behavior |
--minimize | minimize while preserving the observed crash identity |
--sarif | export observations as SARIF 2.1.0 |
Triage uses the Exact identity for build-local buckets, records Stable for cross-build comparison, and retains the legacy hash for historical oracle compatibility. Generated reports are unrated investigation records: no automatic CVSS or affected-version claim is made.
report uses the same report-producing implementation and flags.
value-diff¶
crucible value-diff \
--reference ./numpy-reference.npy \
--candidate ./target-output.npy \
--atol 1e-6 --rtol 1e-5 \
--json
The command compares exact shape, dtype descriptor, and memory-order metadata before values. Integer and boolean arrays are always exact. Floating-point tolerance is asymmetric in the usual reference-oracle sense: abs(candidate-reference) <= atol + rtol*abs(reference).
| Exit | Meaning |
|---|---|
0 | MATCH under the recorded policy |
1 | DIVERGED; investigate the invariant and impact |
2 | UNKNOWN; an input or comparison requirement was not established |
64 | invalid command usage or policy |
--equal-nan makes two NaNs equal; it does not make one NaN equal a finite value. Positive and negative zero compare numerically by default; --distinguish-signed-zero makes their sign part of the oracle. Inputs and element counts are bounded. Output SHA-256 values bind the exact bytes read. The command performs no target execution, persistence, ranking, severity, or disclosure action.
value-matrix¶
MATRIX_SHA256="$(shasum -a 256 ./matrix.json | awk '{print $1}')"
crucible value-matrix \
--manifest ./matrix.json \
--manifest-sha256 "$MATRIX_SHA256" \
--json
The manifest contains one to 1,024 reference/candidate pairs. Every file uses a relative locator and expected SHA-256, and every pair records an explicit numerical policy and positive byte and element limits. The manifest itself is limited to 1 MiB and must match the digest supplied on the command line before any output is compared. Per-input limits cannot exceed 256 MiB or 16,777,216 elements; aggregate declared work cannot exceed 1 GiB or 67,108,864 elements. Duplicate JSON fields and symlinked pair paths are refused.
Exit codes have the same meanings as value-diff. The aggregate is UNKNOWN if any pair is unknown, otherwise DIVERGED if any complete pair diverges, otherwise MATCH. Counts and pair receipts preserve all measured outcomes. The command is read-only and does not run models or grant campaign, finding, severity, scheduling, or disclosure authority.
value-roundtrip¶
MANIFEST_SHA256="$(shasum -a 256 ./roundtrip.json | awk '{print $1}')"
crucible value-roundtrip \
--manifest ./roundtrip.json \
--manifest-sha256 "$MANIFEST_SHA256" \
--json
The manifest binds one source model, one to 32 execution inputs, the retained procedure, the converter identity and conversion receipt, one to 32 converted artifacts, and reference, direct-load, and converted/reloaded execution cells. Every cell binds its runtime identity, invocation receipt, and .npy output. Identity receipts must contain the exact declared SHA-256, OCI image digest, or Git commit value.
The command verifies every declared file before comparing the three output pairs. A missing, substituted, symlinked, ambiguous, or unbound input returns UNKNOWN before numerical interpretation. DIVERGED means at least one complete pair differs; it does not establish a bug or security impact. The manifest must state authority_effect: "NONE", and the command has no execution, networking, persistence, ranking, scheduling, finding, severity, or disclosure ability.
triage dedup¶
# Replay and group by Exact identity; dry-run by default
crucible triage dedup --crashes ./crashes --harness ./harness
# Storage cleanup by content fingerprint; no process execution
crucible triage dedup --crashes ./crashes --fast --output ./duplicates
Full mode replays candidates and groups by the build-local Exact identity. Fast mode groups by file size and a content fingerprint; it cannot establish that two inputs exercise the same bug. Use --delete only after reviewing the dry run.
oracle and validate-oracle¶
crucible oracle --harness ./harness ./poc-a ./poc-b --json
crucible validate-oracle --manifest tools/oracle/oracle-hashes.json --json
These commands intentionally preserve the locked legacy HashStack contract. Do not treat that key as the current campaign-dedup identity.
regress¶
regress runs admitted banked text through classification, attribution, and identity code. It does not execute PoCs or prove a finding remains live.
| Exit | Meaning |
|---|---|
0 | admitted evidence matches the committed baseline |
1 | analysis drifted |
2 | evidence or baseline was unavailable |
3 | compared entries match, but unadjudicated evidence remains |
--propose writes nothing and cannot admit candidates. --update rewrites the baseline from the current code and therefore requires human review of every change.
capability-capture and capability¶
crucible capability-capture \
--manifest reports/phase0/capability-manifest.json \
--bundle ./capture-2026-08-20
crucible capability \
--manifest reports/phase0/capability-manifest.json \
--json
Capture banks immutable raw observations and always exits 2. It never creates expectations from the run it is supposed to test. After human adjudication, capability exits 0 only when every comparable path ran and no locked positive regressed; incomplete or not_run work exits 2.
coverage¶
crucible coverage report \
--harness ./crucible-cov \
--corpus ./corpus \
--output ./coverage-report
The harness must carry LLVM coverage instrumentation. Coverage shows which code executed; it does not establish that validation branches, mutators, or crash detection are semantically correct.
measurement¶
Full surface, flags and JSON receipt: crucible measurement.
Measurement envelopes are append-only observational sidecars. measurement validates them and does nothing else. Validation reads one bounded local JSON file and, for a correction, its complete oldest-to-newest predecessor chain. It does not collect usage, follow source locators, persist state, execute targets, access the network, update an index, or grant authority.
crucible measurement validate ./envelope.json
crucible measurement validate ./envelope.json --json
crucible measurement validate ./correction.json \
--prior ./predecessor-oldest.json \
--prior ./predecessor-newest.json
measurement validate¶
Takes exactly one envelope path as its argument.
| Flag | Default | Meaning |
|---|---|---|
--prior | none | A predecessor envelope, in oldest-to-newest order. Repeat once per predecessor; the chain must be complete, and a chain of more than 64 envelopes including the one under validation is refused before any file is read |
--json | false | Emit the validation receipt as indented JSON instead of one summary line |
Without --prior the envelope is validated on its own. With one or more, the whole chain is validated in the order given, so a correction is only accepted against the predecessors it claims.
The default output is a single line:
measurement-envelope VALID id=<measurement-id> activity=<activity-id> projection=<mode> correction=<n> authority=<effect>
--json emits the same facts as an indented object and writes nothing else to standard output:
{
"valid": true,
"schema": "crucible.measurement-envelope.v1",
"measurement_id": "sha256:<digest>",
"activity_id": "<activity-id>",
"projection_mode": "RETROSPECTIVE",
"correction_sequence": 0,
"authority_effect": "NONE"
}
A receipt is written only when validation succeeds, so valid is true wherever a receipt appears at all; an invalid envelope, an unreadable file, or a file that is not a regular file exits non-zero with the reason on standard error and emits no receipt. authority_effect is carried out of the envelope and reported, never conferred: validating a sidecar grants it no authority over anything.
mutate-stats¶
crucible mutate-stats --help
crucible mutate-stats run --help
crucible mutate-stats join --help
crucible mutate-stats report --help
crucible mutate-stats ab --help
The subcommands capture and compare per-strategy mutation evidence. Keep populations and budgets matched before interpreting an A/B result. The pkg/experiment scorer analyzes completed replicated runs; it is not yet a campaign executor or evidence-capture service.