Campaign operations¶
The operational layer between "launch a fuzzer" and "hand a maintainer a finding".
Every component on this page was written in response to a specific failure that had already happened and had already cost us something. The failure is recorded next to the feature, because a guard whose reason is forgotten is a guard someone deletes.
| component | the failure it retires |
|---|---|
harness-smoke / preflight | a built harness could not execute or could no longer see retained positives, yet a quiet campaign looked clean |
run --sift | one crashing seed silently ended a campaign that appeared to run |
run --supervise | a campaign ended at the first crash, so the second bug was never found |
crucible provenance | a finding was minted against our own uncommitted patch in the target tree |
--json on the result commands | a pipeline could not tell "found nothing" from "could not look" |
pkg/buildinfo | an artifact could not name the binary that produced it |
tools/check-evidence.py | a named bundle was missing or even the quoted crash-class/size fragment was absent; exact frame and line provenance still requires a claim audit |
triage.Observation | text, process facts, and admission were collapsed; slow inputs, recover-mode diagnostics, and failed executions could look like crashes or clean runs |
Where each one sits¶
flowchart LR
subgraph pre["Before the fuzzer starts"]
A[target provenance] --> B["harness-smoke"]
S[seed corpus] --> B
B --> C["preflight: control + retained positive + reach"]
C --> C2["run --sift"]
end
subgraph during["While it runs"]
C2 --> D["run --supervise"]
D --> E[crash artifacts]
end
subgraph after["Turning artifacts into a claim"]
E --> F["triage --json"]
F --> G["crucible provenance --strict"]
G --> H[banked evidence]
H --> I["check-evidence.py"]
end
I --> J[finding doc + disclosure]
style pre fill:#44372b,stroke:#c49a6c,color:#fff
style during fill:#3c432d,stroke:#9a9b62,color:#fff
style after fill:#4b362c,stroke:#a86f4c,color:#fff The correctness gates on that path, with their exit contracts, are generated from the code in Correctness gates.
Before launch: capability, not mere executability¶
harness-smoke proves the binary starts, accepts a valid seed, and does not fault independently of input. preflight adds a clean control, an optional retained positive, and corpus-reach comparison. They answer different questions and both precede run.
crucible harness-smoke --harness ./h --seed ./control
crucible preflight --harness ./h --control ./control \
--known-crash ./retained-positive --corpus ./pristine-corpus
Preflight exit 2 means it could not tell. Without --known-crash, blindness remains unchecked.
crucible run --sift¶
libFuzzer replays the entire seed corpus before it mutates anything. A single seed that faults ends the process during that replay. The campaign exits, the log looks like a run, and nothing was fuzzed. We lost whole overnight windows to this before it was named.
Sifting replays each seed once against the harness and moves the faulting ones to <corpus>-crashers/.
Three behaviours are deliberate and each one is a refusal to guess:
A timeout is not a crash. A slow seed is slow. It stays in the corpus. Treating a timeout as a crash would quietly delete the most interesting seeds we have, because the deep paths are the slow ones.
All-crash aborts rather than empties. If every seed faults, the corpus is not dirty, the harness is broken or the binary is wrong. Sift moves nothing and says so:
every seed faulted: this is a broken harness or the wrong binary, not a dirty corpus. Nothing was moved.
The alternative behaviour, moving all of them, produces an empty corpus and a campaign that runs happily on nothing.
Crashers are moved, not deleted. <corpus>-crashers/ is the first triage queue. A seed that already crashes the harness is a candidate, not garbage. It is also not a finding: it may be a known signature, a stale build, or an artifact of the seed's own provenance.
--sift-workers parallelises the replay; 0 picks a worker count from the machine.
The result shape (SiftResult) carries examined, kept, moved, errors, moved_names, aborted and a note. errors is always present: a run that could not read four seeds is not the same run as one that read them all.
crucible run --supervise¶
libFuzzer exits on the first crash. Without supervision a campaign finds exactly one bug per launch and then burns the rest of its window as a dead process.
Supervision restarts the fuzzer after each crash, caps the log, and halts when the campaign stops producing new sanitizer signatures.
The halt is a signal, not a verdict¶
When --stale-after fires it means: this harness has produced no new sanitizer signature across that many restarts. It does not mean the crashes are known, by-design, fixed, or worthless. It means a human should look. Nothing in the supervisor is allowed to close an assessment.
Configuration¶
| field | default | what it guards |
|---|---|---|
MaxRestarts | 2000 | a hard ceiling so a runaway loop is bounded |
MaxLogBytes | 64 MiB | a thrashing harness can fill a disk in an hour |
StaleAfter | 50 | saturation signal: no new signature in this many restarts |
MinRunTime | 2s | a run shorter than this counts as thrashing |
MaxThrash | 25 | consecutive thrashing runs before giving up |
The thrash guard exists because a harness that dies instantly (a missing asset, the wrong ASAN_OPTIONS) spins at full speed forever. On 2026-08-10 one target harness did exactly that 11,965 times on a LeakSanitizer abort nobody was watching for.
Reading the result¶
flowchart TD
R["SuperviseResult"] --> A{"completed_cleanly"}
A -->|true| A1["the window elapsed with no crash.<br/>This is the healthy outcome."]
A -->|false| B{"thrash_aborted"}
B -->|true| B1["harness or environment fault.<br/>Not a finding. Fix the setup."]
B -->|false| C{"stale_halted"}
C -->|true| C1["SATURATION SIGNAL.<br/>Triage the pile; judge nothing yet."]
C -->|false| D{"unreported_crashes > 0"}
D -->|true| D1["the process died with no sanitizer report.<br/>Still a crash. Replay it."]
D -->|false| E["ran to the restart ceiling"] | field | meaning |
|---|---|
runs / restarts | how many times the fuzzer was started |
signatures | fatal sanitizer signature to count |
distinct_signatures | how many distinct fatal signatures appeared |
non_fatal_diagnostics | UBSan recover-mode warnings, counted separately |
unreported_crashes | the process died and printed no sanitizer report |
thrash_aborted / stale_halted / ceiling_reached | why it stopped |
log_truncated | the log hit MaxLogBytes |
completed_cleanly | the fuzzer exited normally without crashing |
Why non_fatal_diagnostics is a separate counter¶
UBSan is normally built with -fsanitize-recover, so it prints and continues. A single run emits dozens of SUMMARY: UndefinedBehaviorSanitizer: undefined-behavior ... lines that are warnings, not the thing that killed the process. On the first supervised ExecuTorch run the supervisor reported five distinct "signatures" that were all recover-mode noise. Counting them as crashes is not a rounding error, it manufactures findings.
Why unreported_crashes exists¶
A process can die with no sanitizer report at all: a raw SIGSEGV outside ASan's reach, a std::terminate, an OOM kill. An earlier version reported "0 distinct FATAL signatures" while crash artifacts piled up in the output directory, which reads as "clean" and is the opposite of the truth. The bucket makes the unattributed deaths visible instead of invisible.
crucible provenance¶
A commit id does not describe a build.
crucible provenance ~/src/target
crucible provenance ~/src/target --strict # exits non-zero if the tree is dirty
crucible provenance ~/src/target --json
Output is the block that belongs in an evidence receipt: path, commit, commit date, git describe, branch, upstream distance, and whether the working tree has modified tracked files, with the files named.
The failure. On 2026-08-12 a finding was minted, documented, and committed on the strength of a differential: one loop in the target appeared to bounds-check an index while a neighbouring loop appeared not to. The apparent guard turned out to be our own uncommitted patch, left in the target checkout from earlier work and tagged // crucible idx-guard. The differential did not exist upstream at all. The finding was a duplicate and was withdrawn, and a previously submitted finding had been wrongly flagged as fixed on the way through.
The liveness check that was supposed to catch it asked the right question of the wrong tree: it read the source with git show origin/master:file while the crash was replayed against the dirty working tree.
git status --porcelain takes milliseconds. This makes that pairing something you emit rather than something you remember.
Untracked files are counted but are not dirty. A scratch file next to the source does not change what the code under test does. A modified tracked file does. Conflating them would make --strict fire so often it would be disabled.
--json on the result commands¶
Result-bearing paths including run, status, minimize, report, triage, triage dedup, and coverage report emit machine-readable results. The live inventory is maintained in JSON Results; this guide does not freeze a command count.
The contract that matters is not the field list, it is this:
A field that could not be determined is present and explicitly empty, never silently omitted.
Two shapes carry the whole reason:
{ "files_examined": 12, "unique_crashes": 0, "binary_skipped": 12, "complete": false,
"note": "12 binary crash file(s) were skipped" }
Both report zero crashes. One triaged everything and found nothing; the other could not triage any of it. Without complete a consumer reads them identically, and "we looked and it was clean" is a very different sentence from "we could not look".
complete accounts for every way a verdict can fail to be reached: skipped inputs, replay timeouts, replay errors, undescendable directories (walk_errors), unreadable artifacts (read_errors), and reports that failed to write (report_write_errors). It previously counted only the three per-input outcomes, which made it mean "complete except for what I could not see".
Coverage has the same shape for the same reason: a failed llvm-cov must not be indistinguishable from a genuine 0.0%, so summary_available and summary_error travel with line_pct.
Full field reference: JSON result shapes.
Build identity¶
Every JSON result carries a crucible build stamp from pkg/buildinfo. An artifact that cannot name the binary that produced it is not reproducible, and reproducibility is the third R.
One vocabulary for what happened¶
--sift, smoke, triage, minimize, differential replay, coverage, and supervision used to combine process results independently. The shared model is now:
The execution kinds are crash, suspect, intercepted-fault, diagnostic, clean, ambiguous, incomplete, and unknown.
Three distinctions are load-bearing:
incomplete is not clean or crash. A timeout, loader failure, or command that never started establishes no verdict about the input.
diagnostic is not crash. Recover-mode UBSan and LeakSanitizer can print supported text while the target survives.
unknown is not ambiguous. A banked log with no authenticated process facts does not prove that the target ran or exited non-zero. Its text can still regress as evidence without being promoted into an outcome.
Full table and the failure behind each: Crash triage.
What none of this decides¶
The automation surfaces and dedups candidates. It never decides "by-design", "known", "fixed", or "submit". Each unique observation still gets a replay, the appropriate identities, and a read of the source at the crash site, on a HEAD build. The two human gates (severity ratification and disclosure approval) are a hard stop, always.
See Crash triage and Evidence and validation.