Skip to content

Crash triage

A fuzzer artifact is not a vulnerability, and a log is not a process. Crucible's triage layer keeps those distinctions explicit: it replays binary reproducers, records what the process did, extracts what the text says, and surfaces observations that require a person to investigate.

It does not automatically confirm vulnerabilities, assign CVSS scores, infer affected versions, or decide that two crashes have the same root cause.

The observation model

flowchart LR
    A[Artifact admission] --> O[Observation]
    T[Text evidence] --> O
    P[Process facts] --> O
    O --> K[Execution kind]
    O --> H{Needs human triage?}
    H -->|yes| R[Replay, source read,<br/>root-cause analysis]
    H -->|no| N[Record the negative<br/>or incomplete result]

An Observation combines four questions that older paths used to collapse:

Component Question Examples
ArtifactAdmission Is this artifact admissible for this analysis? admitted live input, excluded campaign log, unadjudicated banked log
TextEvidence What evidence appears in stdout/stderr? ASan class, fatal UBSan, recover-mode UBSan, assertion, leak, resource exhaustion
ProcessFacts What did the execution actually do? command started, target ran, exit code, signal, timeout, loader failure
Observation What follows from those facts together? crash, suspect, diagnostic, clean, ambiguous, incomplete, intercepted fault, unknown

Only an explicitly admitted artifact can produce a substantive execution classification. A zero-value or unadjudicated admission fails closed as unknown.

Execution kinds

Kind Meaning Human triage
crash The process died and fatal evidence names the fault yes
suspect The process died, but the available evidence does not identify a supported crash class yes
intercepted-fault A hash-bound harness manifest says the harness intercepted a real target fault yes
diagnostic The process survived after non-fatal sanitizer, leak, or assertion-shaped output no
clean The target ran to completion without supported fault evidence no
ambiguous The command ran but the facts do not establish a clean run or a named fault no; investigate the harness
incomplete Timeout, loader failure, failure to execute, or missing facts prevented a verdict no; rerun correctly
unknown The artifact was not admitted or carries no process facts no

NeedsHumanTriage() means exactly that: put the observation in front of a person. It does not mean "reportable", "confirmed vulnerability", or "safe to disclose".

Text supports; process facts establish

Sanitizer text alone cannot establish that a process crashed. Recover-mode UBSan can print a runtime error: and continue fuzzing. LeakSanitizer reports a leak, not necessarily a fatal memory corruption. JSON transcripts and operator prose can contain words such as heap-buffer-overflow without any target execution at all.

Conversely, a non-zero exit alone is not enough to name a crash. /usr/bin/false, a loader error, and an ASan process that exits 1 are different outcomes. Crucible records the process facts and asks what evidence names the event.

Artifact admission

Use --artifact-kind input for binary reproducers that Crucible will execute. banked-log means the file is output from another execution. Because that log does not carry authenticated exit, signal, timeout, build, or environment facts for this run, triage skips it and exits non-zero rather than manufacturing a report from text.

For the committed regression corpus, admission is explicit in reports/regress-admission.json. Entries are admitted, excluded, or unadjudicated, with a content hash and a structural artifact kind. regress --propose can discover candidates, but it cannot admit them.

Crash classification is not severity

Crucible recognizes evidence classes such as heap-buffer-overflow, stack-buffer-overflow, use-after-free, integer overflow, assertion failure, deadly signal, and resource exhaustion. These labels describe an observed primitive or diagnostic. They do not determine impact.

For example:

  • an out-of-bounds access may be a read, write, or allocator diagnostic;
  • an integer overflow may be harmless, cause an undersized allocation, or bypass a guard;
  • an assertion may be locally reachable or remotely reachable;
  • a stack write may terminate at a canary without yielding control-flow hijack;
  • a resource-exhaustion message may be libFuzzer enforcing its own limit rather than the target mishandling input.

Crucible therefore retired the automatic per-crash-type CVSS table. Generated reports are unrated unless an operator supplies a ratified vector. Affected versions remain unknown until they are tied to the tested target and commit. CWE mappings are taxonomy hints, not proof of the primitive.

Never promote taxonomy into impact

Do not infer code execution from heap-buffer-overflow, copy a default CWE into a disclosure as a determination, or assign affected versions from an unrelated environment file. Replay the PoC, read the source, identify the primitive, and evaluate the actual deployment boundary.

Three identities, three jobs

Crucible retains three identifiers because one key cannot serve every comparison safely.

Identity Purpose Inputs Expected stability
ExactID Within-build campaign deduplication complete raw frame locations and crash evidence changes when the build or exact stack changes
StableID Cross-build and cross-machine regression tracking normalized target functions and repo-relative paths, no line numbers survives line drift and toolchain path changes
legacy HashStack Compatibility with the locked oracle and historical evidence legacy five-frame algorithm frozen; not the current dedup contract

ExactID is deliberately precise. Two traces with identical first five frames but a different sixth frame must not overwrite one another. StableID is deliberately portable and must never be used to merge build-local crash buckets: two bugs in the same function can share a stable signature.

Crash-site attribution skips sanitizer, fuzzer, harness, compiler-runtime, and standard-library frames. If no target frame remains, the result is UNATTRIBUTED; Crucible does not dress a runtime frame up as a source location.

flowchart TD
    S[Symbolized trace] --> F[Classify every frame]
    F --> E[ExactID: full raw locations]
    F --> T[Target frames only]
    T --> ST[StableID: normalized function<br/>and repo-relative path]
    T --> C{Target site exists?}
    C -->|yes| A[Attributed source site]
    C -->|no| U[UNATTRIBUTED]
    S --> L[Legacy HashStack<br/>oracle compatibility only]

Operational workflow

1. Replay inputs, not logs

crucible triage \
  --artifact-kind input \
  --crashes ./crashes \
  --harness ./crucible-libfuzzer-model \
  --target model-loader \
  --output ./reports \
  --json

The harness execution supplies process facts. Preserve the target commit, dirty state, binary hash, PoC hash, environment, and raw output with the evidence.

2. Deduplicate with the right key

crucible triage dedup \
  --crashes ./crashes \
  --harness ./crucible-libfuzzer-model \
  --output ./duplicates

Full mode replays inputs and groups build-local observations by ExactID. It is dry-run unless an output or deletion action is requested. Fast mode groups identical/near-identical content without executing the harness; it is storage cleanup, not crash-root-cause deduplication.

3. Read the raw evidence and source

For every Exact bucket:

  1. confirm the control input is clean;
  2. reproduce the crafted input on the recorded build;
  3. inspect the full sanitizer output, not a summary line;
  4. read the target source at the attributed site and its callers;
  5. identify read versus write, bounds, attacker control, and the reachable boundary;
  6. search existing issues and advisories;
  7. replay on current upstream HEAD before saying the defect is live.

4. Minimize without changing identity

Minimization succeeds only when the candidate still produces the expected Exact identity. "Something crashed" is insufficient: a smaller input can fall into a different, already-known bug.

5. Generate an internal report

Generated reports are investigation records. By default they say:

  • severity: UNRATED;
  • affected versions: unknown;
  • CVSS: absent;
  • attribution: target site or explicitly unavailable;
  • identities: Exact, Stable, and legacy compatibility hash where applicable.

An operator may add a CVSS vector only after ratifying the attack scenario. A generated report is not a CVE submission and must not be sent without source review and disclosure review.

SARIF export

crucible triage \
  --crashes ./crashes \
  --harness ./crucible-libfuzzer-model \
  --sarif ./results.sarif

SARIF levels come from the supported taxonomy, not CVSS:

  • memory-safety observations: error;
  • crash-only observations: warning;
  • resource exhaustion and diagnostics: note.

Fingerprints carry the current identity fields when available. A SARIF result remains a triage observation; importing it into another system does not upgrade it to a confirmed vulnerability.

Regression and capability checks

The commands answer different questions:

Command Replays What green means
regress admitted banked text through pure analysis the current code interprets admitted evidence the same way as the baseline
validate-oracle retained PoCs through a harness the legacy oracle behavior still reproduces on this build
capability-capture controls and PoCs through migrated paths nothing by itself; it always exits 2 pending adjudication
capability adjudicated cases through migrated paths every comparable path ran and no locked positive regressed

A green regress run does not mean a finding is still live. It never runs the target.

Human gates

Crucible stops short of two decisions:

  1. Exploitability and severity. A person ratifies the primitive, attack boundary, impact, CWE, and any CVSS vector.
  2. Disclosure. A person chooses the channel, timing, scope, and exact text and performs the send.

That boundary is intentional. The tool should make unsupported confidence harder, not automate it.