Configuration¶
Crucible configuration comes from four places. Keep them separate in evidence:
- build variables for Crucible and harnesses;
- target checkout provenance;
- command flags and effective environment;
- target-specific profiles or harness manifests.
Makefile targets¶
The repository Makefile provides convenience targets:
| Target | Purpose | Needs a target checkout |
|---|---|---|
build | build crucible, crucible-gen, and crucible-triage | no |
test / test-short | run package tests | no |
generate | create a small generated GGUF corpus | no |
vet, fmt, lint | Go checks and formatting | no |
harness-libfuzzer | build the llama.cpp libFuzzer harness family | yes |
harness-afl | build the AFL++ harness | yes |
run-libfuzzer / run-afl | convenience campaign launches | yes |
triage-libfuzzer | replay libFuzzer artifacts with the default harness | yes |
verify-mutators | require a defined custom-mutator symbol in built mutator harnesses | yes |
gates | run repository correctness gates and fail if required local integrations did not run | yes |
The four targets in the first group are everything a clean checkout needs to reach the quick start. The rest are target-dependent: they compile or run against an upstream project you supply. Support labels are defined in Installation.
make build
make -C harness/libfuzzer LLAMA_CPP=/path/to/llama.cpp crucible-libfuzzer-model
make verify-mutators
make gates is deliberately stricter than CI. Evidence, mutator binaries, and a stock RPC server may exist only on the research machine. When one of those is absent, the gate is incomplete rather than green.
Build variables¶
| Variable | Default | Meaning |
|---|---|---|
GO | which go | Go compiler used by Make |
LLAMA_CPP | $HOME/src/llama.cpp | checkout used by llama.cpp harness targets |
CRUCIBLE_STOCK_SERVER | unset | stock rpc-server binary for integration gates |
CRUCIBLE_STOCK_TREE | unset | corresponding source tree for profile/integration checks |
Other harness families use their own build scripts under harness/cpp/ and may require target paths, CUDA architecture, model assets, or compiler flags. Read the script and preserve its full build log.
VERSIONS.env is not an affected-version database¶
VERSIONS.env pins multiple dependencies for reproducible fixtures and locked oracles. An entry for one target says nothing about another target, and none of the entries proves the affected versions of a vulnerability.
When writing a finding, derive version claims from the target actually tested:
Record the target commit, commit date, dirty state, binary hash, PoC hash, and effective environment. Do not copy the whole versions file into a report.
Campaign flags¶
Command flags belong to their subcommand, so use the live help rather than a global flag table:
Typical run configuration:
crucible run \
--harness ./harness \
--corpus ./corpus \
--output ./crashes \
--jobs 8 \
--timeout 30s \
--max-len 10485760 \
--fuzz-env ASAN_OPTIONS=detect_leaks=0 \
--json \
-- -max_total_time=3600 -seed=12345
Anything after -- is passed to the fuzzing engine after Crucible's defaults. Record those final arguments; they can change sanitizer behavior, allocation limits, fork mode, and reproducibility.
--fuzz-env and --replay-env are repeatable KEY=VALUE flags. They intentionally do not split on commas, because values such as ASAN_OPTIONS=a=1,b=2 must remain one variable.
Harness manifests¶
Some replay harnesses intercept target faults and return instead of allowing the process to die. That capability must be declared in a sidecar manifest bound to the harness SHA-256. Rebuilding the binary invalidates the declaration. Without a valid manifest, assertion text from a surviving process remains diagnostic.
Profiles for wire protocols or compiled structures are likewise measured from a particular target build. Keep the profile, binary, source commit, compiler, architecture, and hashes together.
Directory layout¶
One workable layout is:
workspace/
target/ # clean pinned upstream checkout
harnesses/ # built binaries + manifests + build logs
corpus/
control/ # known-good inputs
seeds/ # pristine launch corpus
retained-positive/ # known PoCs used by preflight/oracles
campaigns/<run-id>/
crashes/
command.json
provenance.json
reports/ # internal triage output
evidence/<finding>/<witness>/
receipt.md
sha256sums.txt
raw-output/
Keep launch corpora flat for libFuzzer. Keep retained PoCs out of a blind-rediscovery corpus unless the experiment explicitly tests replay rather than rediscovery.
Fuzzer dictionaries¶
Format dictionaries such as corpus/gguf.dict provide tokens to libFuzzer and AFL++. A dictionary can improve reach but is not structure awareness and does not prove inputs survive parsing.
Reproducibility checklist¶
Before treating a result as evidence, preserve:
- target commit and dirty state;
- harness binary and SHA-256;
- harness/profile manifest and SHA-256;
- compiler, sanitizer runtime, architecture, and relevant build flags;
- control and crafted input hashes;
- final argv and effective environment;
- timeout and resource limits;
- raw stdout/stderr and kernel process facts;
- Exact/Stable identities and attribution;
- whether every requested gate actually ran.
See JSON result shapes for machine-readable contracts.