Corpus Management¶
A corpus is a set of starting bytes, not proof of coverage. Crucible combines small synthetic seeds, real artifacts, format dictionaries, and retained regression inputs, then measures what the target actually executes.
Seed roles¶
| Seed class | Best use | Limitation |
|---|---|---|
| Minimal synthetic | Fast parser entry and isolated field coverage | May omit real loader dependencies |
| Architecture-targeted synthetic | Model-loader dispatch and metadata relationships | Still an approximation of a shipped model |
| Real artifact | Natural metadata, tensor shapes, and downstream consumers | Often large and slow |
| Historical reproducer | Regression and oracle validation | Must not be counted as novel discovery |
| Control | Proves the harness distinguishes valid from crafted input | A control that faults invalidates the experiment |
Synthetic-first means starting with small, inspectable inputs; it does not mean claiming every parser path is covered. A real artifact is required whenever the consequence depends on a natural consumer that a minimal harness or seed may skip.
Generate GGUF seeds¶
crucible generate --output ./corpus/gguf --count 100
# or the standalone generator
crucible-gen --output ./corpus/gguf --count 100 --seed 42
The generator covers a minimal file, a non-trivial tensor/metadata file, large descriptor sets, metadata value types, common tensor types, nested arrays, alignment variants, dimension edge cases, and mixed types. Optional architecture and CLIP modes add loader-oriented seeds.
Generation success only establishes that files were written. Use harness-smoke, preflight, and coverage collection to establish target behavior.
Loading APIs¶
pkg/corpus.LoadCorpus reads .gguf files directly inside one directory and parses each one. It is not recursive. One malformed GGUF returns an error for the load.
LoadCorpusBytes has the same non-recursive selection but returns raw bytes. The in-memory Corpus type provides synchronized Add, Get, Pick, and byte-form methods. Byte methods return errors because parsing or serialization can fail.
Dictionaries¶
Crucible ships dictionaries for GGUF, grammar, Jinja, JSON Schema, PyTorch, RPC, server input, and TFLite. They are target hints, not a substitute for structured mutation.
If --dict is omitted, crucible run looks only for the first *.dict file in the selected corpus directory. It does not infer a dictionary from the harness name and does not search parent or subdirectories. Pass --dict explicitly for reproducible campaigns.
Dictionary sizes and vocabularies change. The files under corpus/*.dict are the source of truth; documentation does not freeze entry counts.
Minimization¶
Two operations are intentionally different:
corpus.Minimizeremoves byte-identical serialized entries by SHA-256.corpus.MinimizeWithCoverageruns a file-argument coverage harness, collects LLVM profiles, and performs greedy set cover.
The coverage minimizer accepts format-agnostic CorpusEntry values. GGUF files use GGUFEntry; other formats use RawFileEntry. An empty harness selects content dedup. If coverage cannot be collected, the implementation falls back to content dedup, and zero-coverage entries are preserved after dedup rather than silently discarded.
Corpus minimization is not crash deduplication. Crash artifacts are grouped by Exact crash identity; two different inputs with identical coverage can still expose different faults.
Promotion rules¶
Before adding a generated or discovered input to a retained corpus:
- hash the bytes;
- record whether it is a seed, control, candidate, or adjudicated positive;
- bind any expected result to target binary, environment, and PoC hashes;
- keep historical reproducer populations separate from discovery populations; and
- never let a directory scan silently turn candidates into locked positives.
These rules keep corpus growth from turning into population drift or a self-written oracle.