Crash triage¶
A fuzzer artifact is not a vulnerability, and a log is not a process. Crucible's triage layer keeps those distinctions explicit: it replays binary reproducers, records what the process did, extracts what the text says, and surfaces observations that require a person to investigate.
It does not automatically confirm vulnerabilities, assign CVSS scores, infer affected versions, or decide that two crashes have the same root cause.
The observation model¶
flowchart LR
A[Artifact admission] --> O[Observation]
T[Text evidence] --> O
P[Process facts] --> O
O --> K[Execution kind]
O --> H{Needs human triage?}
H -->|yes| R[Replay, source read,<br/>root-cause analysis]
H -->|no| N[Record the negative<br/>or incomplete result] An Observation combines four questions that older paths used to collapse:
| Component | Question | Examples |
|---|---|---|
ArtifactAdmission | Is this artifact admissible for this analysis? | admitted live input, excluded campaign log, unadjudicated banked log |
TextEvidence | What evidence appears in stdout/stderr? | ASan class, fatal UBSan, recover-mode UBSan, assertion, leak, resource exhaustion |
ProcessFacts | What did the execution actually do? | command started, target ran, exit code, signal, timeout, loader failure |
Observation | What follows from those facts together? | crash, suspect, diagnostic, clean, ambiguous, incomplete, intercepted fault, unknown |
Only an explicitly admitted artifact can produce a substantive execution classification. A zero-value or unadjudicated admission fails closed as unknown.
Execution kinds¶
| Kind | Meaning | Human triage |
|---|---|---|
crash | The process died and fatal evidence names the fault | yes |
suspect | The process died, but the available evidence does not identify a supported crash class | yes |
intercepted-fault | A hash-bound harness manifest says the harness intercepted a real target fault | yes |
diagnostic | The process survived after non-fatal sanitizer, leak, or assertion-shaped output | no |
clean | The target ran to completion without supported fault evidence | no |
ambiguous | The command ran but the facts do not establish a clean run or a named fault | no; investigate the harness |
incomplete | Timeout, loader failure, failure to execute, or missing facts prevented a verdict | no; rerun correctly |
unknown | The artifact was not admitted or carries no process facts | no |
NeedsHumanTriage() means exactly that: put the observation in front of a person. It does not mean "reportable", "confirmed vulnerability", or "safe to disclose".
Text supports; process facts establish¶
Sanitizer text alone cannot establish that a process crashed. Recover-mode UBSan can print a runtime error: and continue fuzzing. LeakSanitizer reports a leak, not necessarily a fatal memory corruption. JSON transcripts and operator prose can contain words such as heap-buffer-overflow without any target execution at all.
Conversely, a non-zero exit alone is not enough to name a crash. /usr/bin/false, a loader error, and an ASan process that exits 1 are different outcomes. Crucible records the process facts and asks what evidence names the event.
Artifact admission¶
Use --artifact-kind input for binary reproducers that Crucible will execute. banked-log means the file is output from another execution. Because that log does not carry authenticated exit, signal, timeout, build, or environment facts for this run, triage skips it and exits non-zero rather than manufacturing a report from text.
For the committed regression corpus, admission is explicit in reports/regress-admission.json. Entries are admitted, excluded, or unadjudicated, with a content hash and a structural artifact kind. regress --propose can discover candidates, but it cannot admit them.
Crash classification is not severity¶
Crucible recognizes evidence classes such as heap-buffer-overflow, stack-buffer-overflow, use-after-free, integer overflow, assertion failure, deadly signal, and resource exhaustion. These labels describe an observed primitive or diagnostic. They do not determine impact.
For example:
- an out-of-bounds access may be a read, write, or allocator diagnostic;
- an integer overflow may be harmless, cause an undersized allocation, or bypass a guard;
- an assertion may be locally reachable or remotely reachable;
- a stack write may terminate at a canary without yielding control-flow hijack;
- a resource-exhaustion message may be libFuzzer enforcing its own limit rather than the target mishandling input.
Crucible therefore retired the automatic per-crash-type CVSS table. Generated reports are unrated unless an operator supplies a ratified vector. Affected versions remain unknown until they are tied to the tested target and commit. CWE mappings are taxonomy hints, not proof of the primitive.
Never promote taxonomy into impact
Do not infer code execution from heap-buffer-overflow, copy a default CWE into a disclosure as a determination, or assign affected versions from an unrelated environment file. Replay the PoC, read the source, identify the primitive, and evaluate the actual deployment boundary.
Three identities, three jobs¶
Crucible retains three identifiers because one key cannot serve every comparison safely.
| Identity | Purpose | Inputs | Expected stability |
|---|---|---|---|
ExactID | Within-build campaign deduplication | complete raw frame locations and crash evidence | changes when the build or exact stack changes |
StableID | Cross-build and cross-machine regression tracking | normalized target functions and repo-relative paths, no line numbers | survives line drift and toolchain path changes |
legacy HashStack | Compatibility with the locked oracle and historical evidence | legacy five-frame algorithm | frozen; not the current dedup contract |
ExactID is deliberately precise. Two traces with identical first five frames but a different sixth frame must not overwrite one another. StableID is deliberately portable and must never be used to merge build-local crash buckets: two bugs in the same function can share a stable signature.
Crash-site attribution skips sanitizer, fuzzer, harness, compiler-runtime, and standard-library frames. If no target frame remains, the result is UNATTRIBUTED; Crucible does not dress a runtime frame up as a source location.
flowchart TD
S[Symbolized trace] --> F[Classify every frame]
F --> E[ExactID: full raw locations]
F --> T[Target frames only]
T --> ST[StableID: normalized function<br/>and repo-relative path]
T --> C{Target site exists?}
C -->|yes| A[Attributed source site]
C -->|no| U[UNATTRIBUTED]
S --> L[Legacy HashStack<br/>oracle compatibility only] Operational workflow¶
1. Replay inputs, not logs¶
crucible triage \
--artifact-kind input \
--crashes ./crashes \
--harness ./crucible-libfuzzer-model \
--target model-loader \
--output ./reports \
--json
The harness execution supplies process facts. Preserve the target commit, dirty state, binary hash, PoC hash, environment, and raw output with the evidence.
2. Deduplicate with the right key¶
crucible triage dedup \
--crashes ./crashes \
--harness ./crucible-libfuzzer-model \
--output ./duplicates
Full mode replays inputs and groups build-local observations by ExactID. It is dry-run unless an output or deletion action is requested. Fast mode groups identical/near-identical content without executing the harness; it is storage cleanup, not crash-root-cause deduplication.
3. Read the raw evidence and source¶
For every Exact bucket:
- confirm the control input is clean;
- reproduce the crafted input on the recorded build;
- inspect the full sanitizer output, not a summary line;
- read the target source at the attributed site and its callers;
- identify read versus write, bounds, attacker control, and the reachable boundary;
- search existing issues and advisories;
- replay on current upstream HEAD before saying the defect is live.
4. Minimize without changing identity¶
Minimization succeeds only when the candidate still produces the expected Exact identity. "Something crashed" is insufficient: a smaller input can fall into a different, already-known bug.
5. Generate an internal report¶
Generated reports are investigation records. By default they say:
- severity:
UNRATED; - affected versions: unknown;
- CVSS: absent;
- attribution: target site or explicitly unavailable;
- identities: Exact, Stable, and legacy compatibility hash where applicable.
An operator may add a CVSS vector only after ratifying the attack scenario. A generated report is not a CVE submission and must not be sent without source review and disclosure review.
SARIF export¶
crucible triage \
--crashes ./crashes \
--harness ./crucible-libfuzzer-model \
--sarif ./results.sarif
SARIF levels come from the supported taxonomy, not CVSS:
- memory-safety observations:
error; - crash-only observations:
warning; - resource exhaustion and diagnostics:
note.
Fingerprints carry the current identity fields when available. A SARIF result remains a triage observation; importing it into another system does not upgrade it to a confirmed vulnerability.
Regression and capability checks¶
The commands answer different questions:
| Command | Replays | What green means |
|---|---|---|
regress | admitted banked text through pure analysis | the current code interprets admitted evidence the same way as the baseline |
validate-oracle | retained PoCs through a harness | the legacy oracle behavior still reproduces on this build |
capability-capture | controls and PoCs through migrated paths | nothing by itself; it always exits 2 pending adjudication |
capability | adjudicated cases through migrated paths | every comparable path ran and no locked positive regressed |
A green regress run does not mean a finding is still live. It never runs the target.
Human gates¶
Crucible stops short of two decisions:
- Exploitability and severity. A person ratifies the primitive, attack boundary, impact, CWE, and any CVSS vector.
- Disclosure. A person chooses the channel, timing, scope, and exact text and performs the send.
That boundary is intentional. The tool should make unsupported confidence harder, not automate it.