Skip to content

Quick start

Part 1 runs end to end on a clean checkout with no target and no network. Its commands were executed against binaries built from this source; the outputs are the real ones, with only the paths shortened. Part 2 needs a target checkout, a compiled harness, and local control inputs.

Build the tools first: Installation.

Part 1: works anywhere

Nothing here needs an upstream project, a compiled harness, or a model you have to find. The artifacts are synthetic, generated by the tool itself.

1. Check the environment

CRUCIBLE_ROOT="$PWD"
export PATH="$CRUCIBLE_ROOT:$PATH"
crucible doctor

Run these commands from the Crucible repository root in the same shell after make build. Saving the root in PATH keeps the three built tools available when you move to a temporary directory below.

doctor reports the Go toolchain, the C++ compiler, whether the sanitizer runtimes actually initialise and report, and which optional target checkouts it can see. Its output is specific to your machine, so it is not reproduced here. A red line for an optional target is expected and does not block Part 1.

2. Generate a synthetic corpus

mkdir -p /tmp/crucible-quickstart && cd /tmp/crucible-quickstart
crucible-gen --output ./corpus --count 5 --seed 42
Generating GGUF seed corpus to ./corpus
  Wrote 30 structural seeds
  Wrote 5 mutated variants
Done. Total files in ./corpus

The structural seeds are a fixed set covering different parser paths; --count sets how many mutated variants are written alongside them. --seed 42 makes the run reproducible.

3. Look at what was generated

crucible gguf inspect corpus/seed_001.gguf
GGUF v3  (alignment 32)
metadata (1):
  general.architecture                     STRING
tensors (1):
  weight                                   type=0   dims=[4 4]

inspect prints the structure a parser has to walk: version, alignment, metadata keys and their types, and each tensor's element type and dimensions. Those are the fields mutation targets.

4. Parse-check a file

crucible gguf parse corpus/seed_001.gguf --json
{
  "metadata": 1,
  "tensors": 1,
  "valid": true,
  "version": 3
}

parse answers one question: does this file parse, and if not, where does it fail. Run it on a mutated variant to see a different structure reported.

5. Mutate, and see which strategies fired

crucible mutate corpus/seed_001.gguf ./mutated.gguf --seed 7 --show-mutations
wrote ./mutated.gguf (320 bytes) from corpus/seed_001.gguf
applied 3 strategy(ies):
  metadata.key_shadow
  consistency.offset_beyond
  consistency.alignment_disagree

This is the difference from byte mutation. Each named strategy corrupts a relationship: a duplicate metadata key, a tensor offset pointing past the data region, an alignment field that disagrees with the layout. The file stays structurally plausible enough to reach code that a random byte flip would not.

The same --seed gives the same file:

crucible mutate corpus/seed_001.gguf ./again.gguf --seed 7
shasum -a 256 mutated.gguf again.gguf

Both digests match, so a mutated input you want to keep can be regenerated rather than stored.

6. Round-trip

crucible gguf serialize corpus/seed_001.gguf ./roundtrip.gguf
wrote ./roundtrip.gguf (192 bytes)

Parse then re-serialize. A round trip that changes the file, or fails, is itself informative about the parser's model of the format.

That is the whole of Part 1. You have a corpus, you can read it, you can mutate it deterministically, and none of it required a target.

Part 2: pointing it at real code

Target-dependent

From here on you need a local checkout of the project you intend to test and a harness compiled against it. See Installation: building harnesses. The steps below are a map, not a copy-paste sequence: the target, its build flags and its entry point are yours to choose and to record.

The shape of a campaign is:

Step Command What it establishes
Bind the build crucible provenance --strict <target> The commit and dirty state any later evidence refers to
Check the harness crucible harness-smoke It runs a valid input, survives an empty one, and reports reach
Check the campaign crucible preflight A clean control stays clean and the corpus reaches the code under test
Run crucible run or crucible campaign run Discovery through the native harness
Watch concentration crucible saturation Whether the campaign is still learning or re-finding one identity
Triage crucible triage Replayed observations grouped by crash identity

Each of those has its own section in the CLI reference, including the flags and exit codes this page does not repeat.

For a llama.cpp model-loader campaign, update a clean target checkout to current upstream HEAD, then record the exact commit and build the harness against it. Return to the Crucible repository root first. Replace the target and scratch paths with local paths before running these commands:

cd "$CRUCIBLE_ROOT"
TARGET_ROOT=/path/to/llama.cpp
RUN_DIR=/path/to/local-scratch/llama-model
TARGET_COMMIT=$(git -C "$TARGET_ROOT" rev-parse HEAD)
crucible provenance --strict "$TARGET_ROOT"
make -C targets/llamacpp build-fuzz \
  LLAMA_CPP="$TARGET_ROOT" LLAMA_CPP_VERSION="$TARGET_COMMIT"
make -C harness/libfuzzer \
  LLAMA_CPP="$TARGET_ROOT" crucible-libfuzzer-model

The target recipe's default pin is historical. Pass the recorded commit explicitly, and use a fresh target checkout or a build-fuzz/ directory with a matching .crucible-commit marker. An unmarked existing build is rejected because its objects cannot be tied to HEAD. A historical build is useful for regression work, but cannot establish current upstream status.

Generate a flat launch corpus on local scratch storage. The synthetic GGUF entries are fuzzing inputs, not a known-good full model for this loader harness. Supply a model and a retained crash from the same surface and build for the visibility checks:

mkdir -p "$RUN_DIR"
crucible generate --output "$RUN_DIR/corpus" --count 100 --seed 12345
crucible harness-smoke \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --seed /path/to/known-good-model.gguf
crucible preflight \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --control /path/to/known-good-model.gguf \
  --known-crash /path/to/retained-model-crash.gguf \
  --corpus "$RUN_DIR/corpus"

Replace both input placeholders with real local files. preflight reports through text and exit status: 0 means the requested controls support starting the campaign, 1 means a control failed, and 2 means visibility could not be established. If no retained positive exists, omit --known-crash and expect exit 2. An exploratory run can proceed, but a quiet run cannot prove the surface clean.

Once the controls are understood, launch, inspect concentration, and replay candidate inputs:

crucible run \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --corpus "$RUN_DIR/corpus" --output "$RUN_DIR/crashes" \
  --jobs 8 --sift --supervise --json \
  -- -max_total_time=3600 -seed=12345
crucible saturation \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --artifacts "$RUN_DIR/crashes" --threshold 0.90
crucible triage \
  --artifact-kind input --crashes "$RUN_DIR/crashes" \
  --harness ./harness/libfuzzer/crucible-libfuzzer-model \
  --target model-loader --output "$RUN_DIR/reports" \
  --sarif "$RUN_DIR/reports/results.sarif" --json

Arguments after -- go to libFuzzer. Add repeated --fuzz-env KEY=VALUE flags before that separator when a target needs a recorded environment. --sift sets aside seeds that already crash; --supervise can restart after crashes and halt on saturation. A halt is a signal to inspect the campaign, not a vulnerability verdict. saturation exits 1 when its concentration threshold is reached. Triage groups replayed observations by Exact identity, so artifact counts are not finding counts.

A report is not a finding

crucible triage produces unrated investigation records. Turning one into a finding takes replay on a clean pinned build, reading the source at the fault site, and a human severity decision. The tool does not rate, and does not disclose.

Advanced

These are shipped and supported, and they assume the workflow above.

  • crucible regress compares admitted banked analysis text under the current pure analysis layer, with no target execution.
  • crucible validate-oracle replays retained reproducers against a built harness and asserts their recorded identities still hold.
  • crucible capability-capture and crucible capability capture and adjudicate observations across migrated paths.
  • crucible value-diff, value-matrix and value-roundtrip compare successful .npy outputs under an explicit numerical policy.
  • crucible measurement validate checks an observational sidecar and grants it no authority.

Next