Skip to content

Configuration

Crucible configuration comes from four places. Keep them separate in evidence:

  1. build variables for Crucible and harnesses;
  2. target checkout provenance;
  3. command flags and effective environment;
  4. target-specific profiles or harness manifests.

Makefile targets

The repository Makefile provides convenience targets:

Target Purpose Needs a target checkout
build build crucible, crucible-gen, and crucible-triage no
test / test-short run package tests no
generate create a small generated GGUF corpus no
vet, fmt, lint Go checks and formatting no
harness-libfuzzer build the llama.cpp libFuzzer harness family yes
harness-afl build the AFL++ harness yes
run-libfuzzer / run-afl convenience campaign launches yes
triage-libfuzzer replay libFuzzer artifacts with the default harness yes
verify-mutators require a defined custom-mutator symbol in built mutator harnesses yes
gates run repository correctness gates and fail if required local integrations did not run yes

The four targets in the first group are everything a clean checkout needs to reach the quick start. The rest are target-dependent: they compile or run against an upstream project you supply. Support labels are defined in Installation.

make build
make -C harness/libfuzzer LLAMA_CPP=/path/to/llama.cpp crucible-libfuzzer-model
make verify-mutators

make gates is deliberately stricter than CI. Evidence, mutator binaries, and a stock RPC server may exist only on the research machine. When one of those is absent, the gate is incomplete rather than green.

Build variables

Variable Default Meaning
GO which go Go compiler used by Make
LLAMA_CPP $HOME/src/llama.cpp checkout used by llama.cpp harness targets
CRUCIBLE_STOCK_SERVER unset stock rpc-server binary for integration gates
CRUCIBLE_STOCK_TREE unset corresponding source tree for profile/integration checks

Other harness families use their own build scripts under harness/cpp/ and may require target paths, CUDA architecture, model assets, or compiler flags. Read the script and preserve its full build log.

VERSIONS.env is not an affected-version database

VERSIONS.env pins multiple dependencies for reproducible fixtures and locked oracles. An entry for one target says nothing about another target, and none of the entries proves the affected versions of a vulnerability.

When writing a finding, derive version claims from the target actually tested:

crucible provenance --strict /path/to/target
sha256sum /path/to/harness /path/to/poc

Record the target commit, commit date, dirty state, binary hash, PoC hash, and effective environment. Do not copy the whole versions file into a report.

Campaign flags

Command flags belong to their subcommand, so use the live help rather than a global flag table:

crucible run --help
crucible triage --help
crucible preflight --help

Typical run configuration:

crucible run \
  --harness ./harness \
  --corpus ./corpus \
  --output ./crashes \
  --jobs 8 \
  --timeout 30s \
  --max-len 10485760 \
  --fuzz-env ASAN_OPTIONS=detect_leaks=0 \
  --json \
  -- -max_total_time=3600 -seed=12345

Anything after -- is passed to the fuzzing engine after Crucible's defaults. Record those final arguments; they can change sanitizer behavior, allocation limits, fork mode, and reproducibility.

--fuzz-env and --replay-env are repeatable KEY=VALUE flags. They intentionally do not split on commas, because values such as ASAN_OPTIONS=a=1,b=2 must remain one variable.

Harness manifests

Some replay harnesses intercept target faults and return instead of allowing the process to die. That capability must be declared in a sidecar manifest bound to the harness SHA-256. Rebuilding the binary invalidates the declaration. Without a valid manifest, assertion text from a surviving process remains diagnostic.

Profiles for wire protocols or compiled structures are likewise measured from a particular target build. Keep the profile, binary, source commit, compiler, architecture, and hashes together.

Directory layout

One workable layout is:

workspace/
  target/                 # clean pinned upstream checkout
  harnesses/              # built binaries + manifests + build logs
  corpus/
    control/              # known-good inputs
    seeds/                # pristine launch corpus
    retained-positive/    # known PoCs used by preflight/oracles
  campaigns/<run-id>/
    crashes/
    command.json
    provenance.json
  reports/                # internal triage output
  evidence/<finding>/<witness>/
    receipt.md
    sha256sums.txt
    raw-output/

Keep launch corpora flat for libFuzzer. Keep retained PoCs out of a blind-rediscovery corpus unless the experiment explicitly tests replay rather than rediscovery.

Fuzzer dictionaries

Format dictionaries such as corpus/gguf.dict provide tokens to libFuzzer and AFL++. A dictionary can improve reach but is not structure awareness and does not prove inputs survive parsing.

crucible run --harness ./h --corpus ./seeds --dict corpus/gguf.dict

Reproducibility checklist

Before treating a result as evidence, preserve:

  • target commit and dirty state;
  • harness binary and SHA-256;
  • harness/profile manifest and SHA-256;
  • compiler, sanitizer runtime, architecture, and relevant build flags;
  • control and crafted input hashes;
  • final argv and effective environment;
  • timeout and resource limits;
  • raw stdout/stderr and kernel process facts;
  • Exact/Stable identities and attribution;
  • whether every requested gate actually ran.

See JSON result shapes for machine-readable contracts.