Skip to content

JSON result shapes

--json turns a command's output into a parseable result. Once anything reads these fields they are an interface, so they are documented here rather than discovered by grepping the struct tags.

Core triage, dedup, run, minimize, and coverage shapes are defined in cmd/crucible/json_results.go and asserted in cmd/crucible/json_results_test.go. Other commands define their own output types. Check the command-specific schema before treating a field as a cross-command contract.

Reading a result

A missing measurement must not look like a measured zero. For example, triage reports skipped inputs and replay errors alongside its complete field; consumers must inspect those fields before reading unique_crashes: 0 as a clean result.

The core result envelopes and value-comparison receipts carry the crucible build stamp from pkg/buildinfo. This is not universal: doctor --json has checks and required_missing but no build stamp. preflight has no --json flag; use its text and exit status.


value-diff --json

The result carries schema_version, verdict, reason, policy, reference, candidate, compared_elements, different_elements, optional first_difference, string-encoded error maxima, and the crucible build stamp. Reference and candidate objects include the exact SHA-256 and byte count plus parsed .npy version, dtype descriptor, shape, order, and element count when parsing succeeded.

verdict is exactly MATCH, DIVERGED, or UNKNOWN. NaN and infinities are rendered as strings, so the JSON remains valid and does not erase the value that caused a difference. Metadata mismatch is a measured divergence with zero compared values. A parse, support, size, or policy failure is UNKNOWN, not an empty successful comparison.


value-matrix --json

The result carries schema_version, aggregate verdict and reason, the manifest locator, expected and observed manifest SHA-256 values, counts, pairs, and the crucible build stamp. Every pair preserves its ID, expected reference/candidate digests, and the complete value-diff receipt.

counts always reports total, matched, diverged, and unknown. Any unknown pair makes the aggregate UNKNOWN, but measured divergences remain visible in both the counts and pair receipts. Manifest or manifest-digest failure occurs before pair processing and therefore has zero total pairs rather than implying a clean empty matrix.


value-roundtrip --json

The result carries schema_version, verdict, reason, manifest locator and expected/observed manifest digests, study_id, the required authority_effect, every verified file, the three pairwise comparison receipts, and the crucible build stamp. Each file records its role, relative path, expected and observed SHA-256, and byte count.

Any provenance or identity-binding failure yields UNKNOWN before numerical comparison. Otherwise the aggregate is DIVERGED when any pair diverges and MATCH only when reference, direct, and converted/reloaded outputs all agree under the explicit policy.


triage --json

TriageResult.

field type notes
crash_dir, harness, output_dir string inputs
files_examined int artifacts discovered
unique_crashes int observations grouped by build-local Exact identity
by_type object CrashType string to count, always present
crashes array see below
binary_skipped int artifacts not replayed
replay_timeouts int replay exceeded the timeout
replay_timeout string the DEADLINE those timeouts were measured against, always present
replay_errors int replay could not run at all
unadjudicated_logs int banked logs skipped because they cannot supply this run's process facts
walk_errors array directories that could not be descended
read_errors array artifacts that could not be read
report_write_errors array reports that failed to write
sarif_path string if SARIF was requested
siblings, divergent array, int differential pass, if run
cvss_is_estimate bool never omitted
complete bool see below
note string why it is incomplete, in words
crucible string build stamp

complete

complete = binary_skipped == 0
        && replay_timeouts == 0
        && replay_errors == 0
        && len(walk_errors) == 0
        && len(read_errors) == 0
        && len(report_write_errors) == 0

It previously counted only the three per-input outcomes, which made it mean "complete except for what I could not see". An undescendable directory, an unreadable artifact and a failed report write were all invisible to it.

The distinction the whole shape exists to preserve:

{"files_examined": 12, "unique_crashes": 0, "complete": true}
{"files_examined": 12, "unique_crashes": 0, "binary_skipped": 12, "complete": false,
 "note": "12 binary crash file(s) were skipped"}

Both say zero crashes. Only one of them looked.

cvss_is_estimate

Always present for wire compatibility and now false. The old per-crash-type estimate was retired: a crash class cannot determine attack vector, privileges, user interaction, scope, or impact. Where no operator-ratified vector exists, cvss_score is zero and severity reads:

UNRATED (operator ratifies; no automatic score is derivable from a crash class)

Per-crash entries

id, type, stack_hash, function, source_location, target, input_file, minimized_path, report_path, cvss_score, severity, cwe_id, cwe_name, disposition, and an optional divergence. disposition is a taxonomy hint (memory-safety, crash-only, or resource-exhaustion), not an exploitability verdict.

Plus the two crash identities, all of which are always present:

field meaning
exact_id build-local dedup key: every symbolised frame, sanitizer and harness frames included
exact_namespace formula identifier; consumers must refuse identities from an unknown namespace
stable_id cross-machine regression key: normalized target functions and repo-relative paths with line numbers removed. The literal UNATTRIBUTED when no target frame was found
attributed false when no target frame appeared in the trace
target_frames / total_frames thin attribution is visible: 1 target frame of 20 is a weaker key than 8 of 20
identity_note what the keys were computed from, in words, including whether an abort helper was skipped

stack_hash is unchanged in value and meaning: it is still HashStack, and it is what the locked oracle manifests are pinned to. It is no longer the campaign dedup key. Since 2026-08-16 the build-local exact_id keys deduplication, report filenames and campaign buckets, because HashStack truncates at five frames, so two traces with identical first five frames and a different sixth share it and one bug was silently dropped. stack_hash remains the cross-reference to every banked artifact and to the oracle; do not use it to decide whether two crashes are the same bug. exact_id answers the same question from a different frame set, so the two are not interchangeable. See Crash triage.

function and source_location are empty when attributed is false. They are not filled with whatever frame was on top: a stack-printer frame is not evidence of the vulnerable function.

divergence

Present only when a differential pass ran. A crash with no differential pass carries no divergence object at all, rather than an empty one that reads like a measurement.

{"kind": "crash_vs_no_crash", "diverged": true,
 "summary": "primary crashed; cpu-sibling clean",
 "outcomes": [
   {"label": "primary", "path": "/p", "status": "crash", "type": "heap-buffer-overflow", "stack_hash": "abc123"},
   {"label": "cpu", "path": "/c", "status": "clean"},
   {"label": "broken", "path": "/b", "status": "error", "error": "no such file"}
 ]}

A clean sibling carries no type and no stack_hash; it must not inherit the primary's. A harness that failed to run reports status: "error" with the reason, which is neither a crash nor a clean result.


triage dedup --json

DedupResultJSON: crash_dir, mode (fast = content hash, full = replay plus Exact identity), dry_run, output_dir, total, kept, duplicates, removed, unique_hashes, by_type, errors, crucible.

errors is always present and is not folded into any other counter. A consumer that reads kept: 7 without reading errors: 3 is reading a partial result, and the field has to be there for it to notice.

by_type keys are the CrashType strings (heap-buffer-overflow, null-deref, ...), not enum ordinals.


minimize --json

MinimizeResultJSON: corpus, harness, requested_mode (hash or coverage), loaded, kept, removed, changed, written, crucible.

requested_mode is what was asked for. changed and written are separate because a minimize that computed a smaller corpus and failed to write it is not a minimize that did nothing.


coverage report --json

CoverageResultJSON: harness, corpus, html_dir, html_index, summary_available, summary_error, total_lines, hit_lines, line_pct, files, crucible.

summary_available exists because a failed llvm-cov export must not be indistinguishable from a genuine 0.0%:

{"summary_available": false, "summary_error": "llvm-cov export failed: exit status 1", "line_pct": 0.0}
{"summary_available": true, "line_pct": 0.0}

The second is a real measurement of zero coverage. The first is no measurement at all.

Per-file entries carry file, lines, line_hits, line_pct, funcs, func_hits, func_pct, and branch counters where the profile has them.


run --json

RunResultJSON: engine, harness, corpus, crash_dir, dict, command, sift, supervise, secondary_commands, dry_run, executed, exit_code, duration_ms, error, crucible.

command is the exact argv, so a run is reproducible from its own result. timeout is the per-test-case deadline it ran under.

The two deadlines can disagree, and both are recorded so a reader can notice. run takes --timeout and triage takes --replay-timeout. An artifact produced under a 25s campaign deadline and replayed at the 30s triage default is being judged under a deadline that did not produce it, and before both numbers were recorded nothing said so. replay_timeouts: 3 without replay_timeout is a count with no unit. secondary_commands holds the AFL worker invocations beyond the master.

executed is separate from exit_code so --dry-run is unambiguous: a printed command has no exit code, and exit_code: 0 on an unexecuted run would read as success.

sift

SiftResult: corpus, crasher_dir, examined, kept, moved, errors, moved_names, note, aborted.

aborted: true with the all-crashed note means every seed faulted and nothing was moved. See Campaign operations.

supervise

SuperviseResult: runs, restarts, signatures, distinct_signatures, thrash_aborted, stale_halted, ceiling_reached, log_truncated, non_fatal_diagnostics, unreported_crashes, completed_cleanly, elapsed, note.

Three fields exist because of specific miscounts:

  • non_fatal_diagnostics separates UBSan recover-mode warnings from fatal signatures. Counting them as crashes manufactured five "distinct signatures" on the first supervised ExecuTorch run.
  • unreported_crashes counts processes that died with no sanitizer report. Without it the supervisor reported "0 distinct FATAL signatures" while artifacts piled up.
  • completed_cleanly distinguishes "the window elapsed without a crash", which is the healthy outcome, from every failure mode that also produces a low restart count.

stale_halted is a saturation signal, not a verdict. It says no new signature appeared. It does not say the crashes are known, by-design, or fixed.


provenance --json

TargetProvenance: path, commit, commit_date, describe, branch, dirty, dirty_files, untracked_files, upstream, behind_upstream, note.

dirty counts modified tracked files only. Untracked files are counted in untracked_files but do not set dirty, because a scratch file next to the source does not change what the code under test does.


regress

The baseline is an array of RegressEntry values plus a note. Each entry carries log, outcome, outcome_source, evidence_class, crash_type, attribution, stable_id, exact_id, exact_namespace, and frame counts.

The distinction between outcome and evidence_class is load-bearing. A deadly-signal log can carry evidence while its outcome remains unknown because no hash-bound process facts were banked. Do not promote the text field into an execution claim.

JSON consumers must also inspect the process exit:

Exit Meaning
0 admitted entries matched
1 analysis drifted
2 no answer was possible
3 compared entries matched but the corpus remains unadjudicated

capability-capture

Each captured case records its finding and case IDs, population, PoC and harness hashes, effective environment and digest, control result, and one PathObservation per exercised path. Observations carry ran, not_run_reason, outcome, declared comparison capabilities, evidence class, identities, source site, and raw-output path/hash.

Evidence paths are bundle-relative so the transaction can be moved after its atomic rename. Capture always exits 2: JSON from capture is an observation bundle, never an expectation or a pass.

capability

CapabilityResult rows carry finding, case_id, path, population, delta, reason, observed and expected outcomes, and the identity rule used. Locked-positive regressions determine failure; candidate results are surveyed separately and never enter the recall denominator.

not_run and unsupported comparisons are incomplete and exit 2. A JSON consumer must not collapse them into preserved.

doctor, harness-smoke, and preflight

doctor --json reports os, arch, required_missing, and checks; harness-smoke --json reports harness, checks, and ready. Neither output has a crucible build stamp. Both reflect their text results; JSON does not make a failed check successful.

preflight has no --json flag. Its text and exit status are the interface: exit 0 means a retained positive established harness visibility, exit 1 means a check failed, and exit 2 means the question could not be answered, including when --known-crash was omitted.

Other commands

status, report, campaign, mutate, mutate-stats, gguf, and several subcommands also honour --json. Their shapes are not all frozen here; inspect the command's --help and the exported JSON types before making them a long-lived integration.

workspace --json

crucible workspace list|show|map --json emits one read-only index object. It is not a result of an execution: nothing ran, and nothing was decided.

Keys are written as JSON paths (.mode) to keep them distinct from command names.

Key Contents
.schema_version Integer schema version of the index
.mode Always "read-only"
.source_snapshot repository_revision, tree_state, repository_root, evidence_root
.authority What the output does and does not authorise
.sources Every source consulted, each with its identity binding
.workspaces, .studies, .attempts, .evidence_bundles, .findings, .submission_cases Normalized records
.unassigned_objects Objects that resolved to no workspace
.diagnostics Why anything failed to resolve

Every resolved field is a three-key record, never a bare value:

{
  "resolution": "UNKNOWN",
  "reason": "SOURCE_UNREADABLE",
  "sources": []
}

resolution is one of KNOWN, UNKNOWN, CONFLICT, NOT_APPLICABLE. A KNOWN always carries the sources it came from. A CONFLICT stays a conflict: the index does not prefer one source or average them. reason draws from a fixed vocabulary in pkg/operatorindex, including MISSING_IDENTITY_BINDING, SOURCE_UNREADABLE, NOT_VERIFIED_ON_CURRENT_HOST, DIGEST_MISMATCH, INCOMPLETE_RECEIPT and UNSUPPORTED_SCHEMA.

Two reasons are easy to confuse. NOT_VERIFIED_ON_CURRENT_HOST means the bytes were never checked here; DIGEST_MISMATCH means they were checked and did not match. The first is an absence of evidence, the second is evidence of a problem.

Record contents are specific to the operator's machine and are not reproduced in public documentation. See crucible workspace for the command surface.

measurement validate --json

Emitted only on success, so valid is true wherever a receipt appears at all. A failure prints its reason on standard error and emits no receipt.

{
  "valid": true,
  "schema": "crucible.measurement-envelope.v1",
  "measurement_id": "sha256:6c1b8e403d69486757e8e394c5371dece01c2c107f1a4bdb2baf352f26118232",
  "activity_id": "replace-with-existing-receipt-id",
  "projection_mode": "PROSPECTIVE",
  "correction_sequence": 0,
  "authority_effect": "NONE"
}
Field Meaning
valid Always true; a receipt is not written for a failed validation
schema The envelope kind that was validated, crucible.measurement-envelope.v1
measurement_id Content identity of the envelope, as sha256:<digest>
activity_id The activity the envelope observes, carried from the envelope
projection_mode PROSPECTIVE for a plan, RETROSPECTIVE for a record of something finished
correction_sequence 0 for an original envelope; the position in the chain for a correction
authority_effect Carried from the envelope and reported, never conferred

The values above come from pkg/measurement/testdata/prospective-template.json, a neutral template committed for this purpose. Its activity_id is a placeholder by design.

A correction is validated against its predecessors with --prior, oldest first. The chain must be complete, and more than 64 envelopes including the one under validation is refused before any file is read. See crucible measurement.