Skip to content

crucible

The main CLI covers environment validation, corpus generation, campaign execution, triage, regression checks, and capability preservation.

crucible <command> [flags]
crucible <command> --help

JSON support varies by command; check crucible <command> --help before building an integration. preflight reports through text and exit status. Treat documented non-zero exit codes as typed outcomes: several commands use exit 2 for cannot tell, and regress uses exit 3 for an incomplete admission corpus.

Command index

Command Purpose
doctor Check the local build and fuzzing environment
provenance Record target commit, date, upstream distance, and dirty state
recon Read prior findings, leads, submissions, and surface closures before spending a machine
generate Generate and mutate GGUF corpus entries
mutate Apply structure-aware GGUF mutations to one file
protobuf Apply one traceable, structure-aware protobuf mutation
synth Validate, list, and apply schema-synthesized mutation specifications
gguf Inspect, parse-check, and reserialize GGUF artifacts
harness-smoke Test harness lifecycle, a valid seed, empty input, and an optional reach marker
preflight Test control cleanliness, retained-positive visibility, and corpus reach
run Launch libFuzzer or AFL campaigns
campaign Run campaign definitions stored as data
status Report campaign artifacts and status
saturation Measure concentration by build-local crash identity
escape-verify Verify a saturation escape without silently blinding the harness
triage Replay candidate inputs and produce observations and internal reports
triage dedup Deduplicate inputs by Exact identity or content fingerprint
report Alias of the report-producing triage path
minimize Minimize a corpus or reproducer while preserving the relevant behavior
oracle Replay retained PoCs and print the legacy compatibility stack hash
validate-oracle Assert locked legacy oracle entries reproduce
regress Compare admitted banked text under the current pure analysis layer
capability-capture Capture immutable cross-path observations for human adjudication
capability Compare adjudicated positives and candidates across migrated paths
value-diff Compare successful .npy outputs under an explicit numerical policy
value-matrix Compare a hash-bound matrix of successful .npy outputs
value-roundtrip Verify and compare a bound model-conversion round trip
coverage Collect and report source coverage
mutate-stats Capture, join, report, and compare per-strategy mutation telemetry
measurement Validate a local, zero-authority measurement envelope and correction chain
workspace Inspect the disposable, source-attributed read-only operator index

Shell completion is available through crucible completion.

crucible doctor
crucible provenance --strict /path/to/target
crucible recon <target> "<surface>"
crucible harness-smoke --harness ./harness --seed ./control
crucible preflight --harness ./harness --control ./control \
  --known-crash ./retained-positive --corpus ./corpus
crucible run --harness ./harness --corpus ./corpus --output ./crashes -- \
  -max_total_time=3600 -seed=12345
crucible triage --artifact-kind input --crashes ./crashes \
  --harness ./harness --output ./reports --json

workspace

Full surface, flags, selector grammar and exit codes: crucible workspace.

crucible workspace list --repo-root . --json
crucible workspace list --repo-root . --evidence-root /path/to/crucible-evidence
crucible workspace show ID_FROM_LIST --repo-root .
crucible workspace map --repo-root . --evidence-root /path/to/crucible-evidence
crucible workspace verify evidence:<evidence-id>
crucible workspace verify submission:<case-id>

The index presents source-attributed records from the operator's local repository. Every resolved value is KNOWN, UNKNOWN, CONFLICT, or NOT_APPLICABLE and cites its owning source. Conflicts remain visible and fail closed. show follows one workspace through its studies, attempts, evidence, findings and submission cases. map renders that same normalized graph for every imported workspace; neither command introduces another schema or state store.

The human views show unresolved decisions, operator gates, manifest inspection, historical verification, current-host verification and the identity and availability of each owning source. They do not collapse these fields into one status. UNKNOWN and CONFLICT remain explicit, with their reason and source attribution.

With --json, show emits the same normalized Index schema filtered to the selected workspace and its transitively linked studies, attempts, evidence, findings, submission cases, diagnostics and sources. It does not emit a bare workspace that leaves source IDs or operator gates unresolved.

All three commands read canonical repository records, open each repo: and banked: component relative to pinned directory descriptors without following symlinks, and read only exact manifests or receipts capped at 8 MiB. They never recursively walk or hash an evidence payload. A manifest digest match is metadata, not current-host full verification.

These commands write nothing and grant no target-selection, execution, finding, severity, scheduler, policy-update, submission, or disclosure authority. READY_FOR_OPERATOR is not approval and is not evidence that a packet was sent.

doctor

Checks the Go toolchain, sanitizer-capable compiler, target paths, and built harnesses.

crucible doctor --json

It exits non-zero when a required component is missing. A build artifact existing on disk is not enough; target-specific preflight still follows.

provenance

crucible provenance --strict /path/to/target

--strict rejects tracked changes in the target checkout. Store this output with every witness: the commit alone does not describe a dirty build.

recon

crucible recon llama.cpp "gguf tensor loader"
Exit Meaning
0 no recorded closure covers the named surface
1 a recorded closure covers it
2 records could not be read
3 closures exist but no surface was supplied

recon prevents spending campaign time on a negative or fixed surface that the project has already recorded. It does not independently establish that the record is still correct.

generate, mutate, protobuf, synth, and gguf

crucible generate --corpus ./seeds --output ./generated --count 100 --seed 42
crucible mutate --help
crucible protobuf mutate seed.onnx mutated.onnx --seed 42 --json
crucible synth validate rule.synth.json
crucible synth list rule.synth.json
crucible synth apply seed.gguf rule.synth.json out.gguf --rule <rule-name>
crucible gguf inspect model.gguf
crucible gguf parse model.gguf
crucible gguf serialize model.gguf roundtrip.gguf

protobuf mutate refuses an existing output, preserves protobuf wire structure, and emits the selected nested field path and mutation action. Its trace carries authority_effect: NONE; it is attribution for a later execution, not proof that the consumer accepted the artifact.

Use the generated or mutated artifact as an input to a real target harness. A successful structural mutation is not evidence that the target accepted it or reached interesting code.

harness-smoke

crucible harness-smoke \
  --harness ./crucible-libfuzzer-model \
  --seed ./valid.gguf \
  --reach-marker load_model \
  --timeout 30s

Smoke tests execution, a valid seed, empty input, and optionally a target marker. It catches input-independent harness faults before they flood a campaign.

preflight

crucible preflight \
  --harness ./crucible-libfuzzer-model \
  --control ./valid.gguf \
  --known-crash ./retained-poc.gguf \
  --corpus ./pristine-corpus
Exit Meaning
0 the requested checks establish a meaningful starting state
1 a control or capability check failed
2 the question could not be answered

The corpus check detects a corpus that adds almost nothing. It does not prove every seed reaches the target. Without --known-crash, the harness-blindness check is absent.

run

crucible run \
  --harness ./crucible-libfuzzer-model \
  --corpus ./corpus \
  --output ./crashes \
  --jobs 8 \
  --sift \
  --supervise \
  --fuzz-env ASAN_OPTIONS=detect_leaks=0 \
  --json \
  -- -max_total_time=3600 -rss_limit_mb=0

Anything after -- is appended to the engine arguments and can override defaults. Use repeated --fuzz-env KEY=VALUE for campaign environment; it is recorded in JSON. --sift moves already crashing seeds aside before launch. --supervise restarts past crashes and can halt on stale signatures; that halt signals saturation, not a vulnerability verdict.

campaign

Campaign definitions are data. list reports what a config declares; run executes one or more of them by name, or all of them when no name is given.

crucible campaign list --config campaigns.yaml
crucible campaign run --config campaigns.yaml --dry-run
crucible campaign run model-loader --config campaigns.yaml

campaign list

Flag Default Meaning
--config campaigns.yaml Path to the campaigns config file
--json false Emit machine-readable JSON

campaign run

Flag Default Meaning
--config campaigns.yaml Path to the campaigns config file
--dry-run false Print the commands without executing them

saturation and escape-verify

crucible saturation --harness ./harness --artifacts ./crashes --threshold 0.90

crucible escape-verify \
  --harness ./harness \
  --control ./control \
  --escaped-input ./dominant-poc \
  --escape-off CRUCIBLE_ESCAPE_OFF=1 \
  --expect-stable <stable-id> \
  --known-crash ./different-poc \
  --expect-known-stable <different-stable-id> \
  --receipt ./escape-receipt.json

saturation exits 1 when one identity reaches the threshold and 2 when it cannot tell. It never chooses a filter. escape-verify requires a clean control, suppression, reversible recovery of the same stable/class/site, and continued visibility of a different known bug. Exit 2 is incomplete.

triage and report

crucible triage \
  --artifact-kind input \
  --crashes ./crashes \
  --harness ./harness \
  --target gguf-loader \
  --output ./reports \
  --sarif ./reports/results.sarif \
  --json

Important flags:

Flag Meaning
--artifact-kind input replay binary reproducers and collect process facts
--artifact-kind banked-log identify logs from another execution; skip them rather than infer process facts
--replay-env KEY=VALUE add a replay environment value; repeatable
--replay-timeout bound each replay
--sibling label=/path replay through another harness and record differential behavior
--minimize minimize while preserving the observed crash identity
--sarif export observations as SARIF 2.1.0

Triage uses the Exact identity for build-local buckets, records Stable for cross-build comparison, and retains the legacy hash for historical oracle compatibility. Generated reports are unrated investigation records: no automatic CVSS or affected-version claim is made.

report uses the same report-producing implementation and flags.

value-diff

crucible value-diff \
  --reference ./numpy-reference.npy \
  --candidate ./target-output.npy \
  --atol 1e-6 --rtol 1e-5 \
  --json

The command compares exact shape, dtype descriptor, and memory-order metadata before values. Integer and boolean arrays are always exact. Floating-point tolerance is asymmetric in the usual reference-oracle sense: abs(candidate-reference) <= atol + rtol*abs(reference).

Exit Meaning
0 MATCH under the recorded policy
1 DIVERGED; investigate the invariant and impact
2 UNKNOWN; an input or comparison requirement was not established
64 invalid command usage or policy

--equal-nan makes two NaNs equal; it does not make one NaN equal a finite value. Positive and negative zero compare numerically by default; --distinguish-signed-zero makes their sign part of the oracle. Inputs and element counts are bounded. Output SHA-256 values bind the exact bytes read. The command performs no target execution, persistence, ranking, severity, or disclosure action.

value-matrix

MATRIX_SHA256="$(shasum -a 256 ./matrix.json | awk '{print $1}')"
crucible value-matrix \
  --manifest ./matrix.json \
  --manifest-sha256 "$MATRIX_SHA256" \
  --json

The manifest contains one to 1,024 reference/candidate pairs. Every file uses a relative locator and expected SHA-256, and every pair records an explicit numerical policy and positive byte and element limits. The manifest itself is limited to 1 MiB and must match the digest supplied on the command line before any output is compared. Per-input limits cannot exceed 256 MiB or 16,777,216 elements; aggregate declared work cannot exceed 1 GiB or 67,108,864 elements. Duplicate JSON fields and symlinked pair paths are refused.

Exit codes have the same meanings as value-diff. The aggregate is UNKNOWN if any pair is unknown, otherwise DIVERGED if any complete pair diverges, otherwise MATCH. Counts and pair receipts preserve all measured outcomes. The command is read-only and does not run models or grant campaign, finding, severity, scheduling, or disclosure authority.

value-roundtrip

MANIFEST_SHA256="$(shasum -a 256 ./roundtrip.json | awk '{print $1}')"
crucible value-roundtrip \
  --manifest ./roundtrip.json \
  --manifest-sha256 "$MANIFEST_SHA256" \
  --json

The manifest binds one source model, one to 32 execution inputs, the retained procedure, the converter identity and conversion receipt, one to 32 converted artifacts, and reference, direct-load, and converted/reloaded execution cells. Every cell binds its runtime identity, invocation receipt, and .npy output. Identity receipts must contain the exact declared SHA-256, OCI image digest, or Git commit value.

The command verifies every declared file before comparing the three output pairs. A missing, substituted, symlinked, ambiguous, or unbound input returns UNKNOWN before numerical interpretation. DIVERGED means at least one complete pair differs; it does not establish a bug or security impact. The manifest must state authority_effect: "NONE", and the command has no execution, networking, persistence, ranking, scheduling, finding, severity, or disclosure ability.

triage dedup

# Replay and group by Exact identity; dry-run by default
crucible triage dedup --crashes ./crashes --harness ./harness

# Storage cleanup by content fingerprint; no process execution
crucible triage dedup --crashes ./crashes --fast --output ./duplicates

Full mode replays candidates and groups by the build-local Exact identity. Fast mode groups by file size and a content fingerprint; it cannot establish that two inputs exercise the same bug. Use --delete only after reviewing the dry run.

oracle and validate-oracle

crucible oracle --harness ./harness ./poc-a ./poc-b --json
crucible validate-oracle --manifest tools/oracle/oracle-hashes.json --json

These commands intentionally preserve the locked legacy HashStack contract. Do not treat that key as the current campaign-dedup identity.

regress

crucible regress
crucible regress --propose
crucible regress --update

regress runs admitted banked text through classification, attribution, and identity code. It does not execute PoCs or prove a finding remains live.

Exit Meaning
0 admitted evidence matches the committed baseline
1 analysis drifted
2 evidence or baseline was unavailable
3 compared entries match, but unadjudicated evidence remains

--propose writes nothing and cannot admit candidates. --update rewrites the baseline from the current code and therefore requires human review of every change.

capability-capture and capability

crucible capability-capture \
  --manifest reports/phase0/capability-manifest.json \
  --bundle ./capture-2026-08-20

crucible capability \
  --manifest reports/phase0/capability-manifest.json \
  --json

Capture banks immutable raw observations and always exits 2. It never creates expectations from the run it is supposed to test. After human adjudication, capability exits 0 only when every comparable path ran and no locked positive regressed; incomplete or not_run work exits 2.

coverage

crucible coverage report \
  --harness ./crucible-cov \
  --corpus ./corpus \
  --output ./coverage-report

The harness must carry LLVM coverage instrumentation. Coverage shows which code executed; it does not establish that validation branches, mutators, or crash detection are semantically correct.

measurement

Full surface, flags and JSON receipt: crucible measurement.

Measurement envelopes are append-only observational sidecars. measurement validates them and does nothing else. Validation reads one bounded local JSON file and, for a correction, its complete oldest-to-newest predecessor chain. It does not collect usage, follow source locators, persist state, execute targets, access the network, update an index, or grant authority.

crucible measurement validate ./envelope.json
crucible measurement validate ./envelope.json --json
crucible measurement validate ./correction.json \
  --prior ./predecessor-oldest.json \
  --prior ./predecessor-newest.json

measurement validate

Takes exactly one envelope path as its argument.

Flag Default Meaning
--prior none A predecessor envelope, in oldest-to-newest order. Repeat once per predecessor; the chain must be complete, and a chain of more than 64 envelopes including the one under validation is refused before any file is read
--json false Emit the validation receipt as indented JSON instead of one summary line

Without --prior the envelope is validated on its own. With one or more, the whole chain is validated in the order given, so a correction is only accepted against the predecessors it claims.

The default output is a single line:

measurement-envelope VALID id=<measurement-id> activity=<activity-id> projection=<mode> correction=<n> authority=<effect>

--json emits the same facts as an indented object and writes nothing else to standard output:

{
  "valid": true,
  "schema": "crucible.measurement-envelope.v1",
  "measurement_id": "sha256:<digest>",
  "activity_id": "<activity-id>",
  "projection_mode": "RETROSPECTIVE",
  "correction_sequence": 0,
  "authority_effect": "NONE"
}

A receipt is written only when validation succeeds, so valid is true wherever a receipt appears at all; an invalid envelope, an unreadable file, or a file that is not a regular file exits non-zero with the reason on standard error and emits no receipt. authority_effect is carried out of the envelope and reported, never conferred: validating a sidecar grants it no authority over anything.

mutate-stats

crucible mutate-stats --help
crucible mutate-stats run --help
crucible mutate-stats join --help
crucible mutate-stats report --help
crucible mutate-stats ab --help

The subcommands capture and compare per-strategy mutation evidence. Keep populations and budgets matched before interpreting an A/B result. The pkg/experiment scorer analyzes completed replicated runs; it is not yet a campaign executor or evidence-capture service.