Regression and capability testing¶
Crucible has three regression layers because "the log still parses", "the PoC still crashes", and "every migrated consumer still sees it" are different claims.
| Layer | Command | Input | Question |
|---|---|---|---|
| analysis regression | crucible regress | admitted banked artifacts | does current pure analysis interpret the same bytes consistently? |
| live oracle | crucible validate-oracle | retained PoCs + built harness | does this build reproduce the locked legacy identity? |
| capability preservation | crucible capability | adjudicated cases + controls | did each migrated path retain the expected capability? |
Never substitute one layer for another.
Analysis regression: regress¶
crucible regress \
--admission reports/regress-admission.json \
--baseline reports/generated/analysis-baseline.json \
--evidence "$HOME/crucible-evidence"
The command verifies every manifest entry's content hash, then re-runs text evidence extraction, crash taxonomy, attribution, Exact/Stable identities, and the observation outcome when a hash-bound process-facts record exists.
It does not execute a target. Therefore:
regressexit 0 means the analysis layer reads admitted evidence as expected. It does not mean a finding is live, a fix failed, or a PoC still crashes.
Admission is explicit¶
The admission manifest separates:
admitted: reviewed as appropriate for the analysis corpus;excluded: reviewed and rejected for a recorded structural reason;unadjudicated: seen, but nobody has decided.
It also records artifact structure independently: single-process log, campaign log, differential excerpt, multi-process artifact, aggregate rows, or transcript. Discovery cannot convert a candidate to admitted:
--propose prints candidates and writes nothing. To adopt an analysis change:
- inspect every drift against raw evidence and source;
- decide whether it is a correction or a regression;
- update admission or code as appropriate;
- run
crucible regress --updateonly after that review; - inspect and commit the baseline diff.
Exit contract¶
| Exit | Meaning |
|---|---|
0 | admitted artifacts were honored and no analysis drifted |
1 | at least one analysis result drifted |
2 | the evidence or baseline could not be evaluated |
3 | admitted entries matched, but unadjudicated evidence remains |
Exit 3 is not green. It says part of the banked population remains outside the comparison.
Live replay: validate-oracle¶
crucible validate-oracle \
--manifest tools/oracle/oracle-hashes.json \
--poc-dir ./pocs \
--harness-dir ./harness/libfuzzer \
--json
This replays each retained PoC through its harness and compares the frozen legacy HashStack value. The legacy identity is intentional: it preserves compatibility with historical answer keys. It is not the current campaign-dedup key.
Before trusting a live result, bind the environment:
crucible provenance --strict /path/to/target
crucible harness-smoke --harness ./harness --seed ./control
crucible preflight --harness ./harness --control ./control \
--known-crash ./retained-poc
A no-crash result on a rebuilt harness can mean fixed target, blind harness, wrong architecture, missing runtime capability, stale manifest, or different input. The controls separate those cases.
Cross-path preservation: capability capture and verify¶
Bootstrap evidence with capture:
crucible capability-capture \
--manifest reports/phase0/capability-manifest.json \
--bundle ./capability-capture-2026-08-20
Capture is transactional, banks raw output and hashes, writes no expectations, and always exits 2. That is deliberate: a tool cannot establish its own answer key from the run it is supposed to verify.
After a human adjudicates the bundle and records hash-bound expectations:
The manifest separates locked positives from candidates. Recall is computed over locked positives only. Candidate regressions are surveyed, not scored.
Comparison depends on provenance:
- same harness hash, PoC hash, and environment digest: Exact equality is required;
- different build: Stable identity, evidence class, and normalized source site are required;
- different PoC bytes: incomplete, not a cross-build comparison;
- path cannot emit a required field:
not_comparable, incomplete; - path did not execute:
not_run, incomplete.
Exit 0 requires every comparable locked path to run and no regression. Incomplete work exits 2.
Testing a fix across versions¶
For a target fix, keep the vulnerable and patched builds separate and record both:
- target commit and dirty state;
- harness binary hash and manifest;
- clean control hash and expected outcome;
- PoC hash;
- effective environment digest;
- raw output and process facts;
- Exact, Stable, class, and attributed site.
The vulnerable build should reproduce the expected primitive. The patched build should reject or handle the input cleanly, and a different retained positive should still be visible. This last check prevents a broad guard from "fixing" the test by blinding the harness.
Mutation rediscovery experiments¶
To claim that a mutation strategy rediscovered a known bug:
- start from a declared seed corpus that does not already contain the PoC;
- use a fixed seed and budget;
- verify the custom mutator is actually linked;
- bank the generated artifact;
- replay it through the target;
- require the expected identity under the appropriate provenance rule;
- repeat across independent seeds rather than reporting the best run.
Time-to-first evidence with censored runs needs survival-aware analysis. The current pkg/experiment package scores completed replicated run records; it does not launch campaigns, capture their evidence, or adjudicate root causes. Do not describe it as an end-to-end experiment runner yet.
Adding a retained positive¶
- Bank the original PoC, control, raw output, target provenance, and binary/environment hashes.
- Record the source of every expected identity and outcome.
- Capture every supported path on immutable copies.
- Adjudicate the bundle independently of the implementation being tested.
- Add it to the locked manifest only after that review.
- Negative-control the gate: force one path to miss it and confirm verification fails.
The standard is not "the test is green." The standard is that forcing the protected capability to disappear makes it red for the reason the test claims to cover.