Evidence
Every number, and what it's worth
Benchmarks are easy to game and easier to misread. So each result below carries an evidence class saying what it can support, including the ones that undercut us.
The build
Metallum-1B
Pretraining completed at 16.0B tokens, followed by one bounded knowledge and format SFT stage. The champion checkpoint has passed its frozen batteries and its one-time sealed holdout qualification.
Model
- Parameters
- 1,002M
- Layers
- 26
- d_model
- 1792
- Attention
- GQA 28/14
- d_ff
- 4864
- Vocabulary
- 40k
- Context
- 2048
The run
- GPU
- RTX 5090
- Count
- 1
- VRAM
- 32 GB
- Wall clock
- ~13.5 days
- Throughput
- ~18k tok/s
- Optimizer
- Muon (2D) + AdamW
One consumer card. No cluster, no rented H100 fleet, no FSDP.
Corpus
- Tokens seen
- 16.0B
- Effective
- 19.399B
- Unique
- 8.473B
Custom 40k byte-level BPE with FIM sentinels, fit on the domain. Curated and decontaminated in-house, not a general web scrape.
Architecture
The drawing
Six of twenty-six layers carry no positional encoding at all. That single choice is why a 2,048-token model retrieves perfectly at the edge of its own window.
NoPE every 4th layer
Six of twenty-six layers carry no positional encoding at all. Those layers learn length-generalizing retrieval heads, which is why a 2,048-context model still scores 18/18 on distance needles at the very edge of its window.
QK-norm + z-loss
Normalizing queries and keys before the attention product, plus a z-loss term on the logits, is what kept a 16B-token run from diverging on a single card with no cluster to restart from.
GQA 28Q / 14KV
Halving the key-value heads halves the KV cache. That is the difference between a long-context model that fits on consumer hardware and one that does not.
Document-masked attention
Attention never crosses a document boundary inside a packed sequence, so the model never learns to continue one document into an unrelated one.
FIM sentinels
Fill-in-the-middle sentinels in a 40k byte-level BPE vocabulary, fit on the target domain rather than inherited from a general-purpose tokenizer.
Muon on 2D parameters
Muon for matrices, AdamW for everything else, under a WSD schedule with short-sequence early phases. Throughput held at roughly 18k tokens/second.
Reading the table
Not all numbers weigh the same
A score measured on data the model already saw is not evidence of capability, no matter how good it looks. We separate the two rather than averaging them into one flattering figure.
Measured once, on a holdout that was built, sealed, and never trained on. One run, no retries, result published whatever it said. This is the generalization claim.
Measured on surfaces the model never trained on. Strong capability evidence, repeatable on demand.
Used to steer training. Selection-aware or previously observed, so it cannot be read as a clean capability estimate.
The suite itself declares known overlap with training inputs. Published for completeness, never as a capability claim.
Results
The full table
Champion checkpoint s2b075, measured by frozen evaluators with sealed checkpoint identity.
Sealed holdout
| Metric | Metallum-1B | Comparators |
|---|---|---|
| H9 sealed holdout↓ lower is betterBits per byte on 400k characters of post-cutoff Wikipedia, built and sealed before the run and evaluated exactly once. Contamination verified at 0.0099% shingle containment. One shot, no retries. This is the number that carries the release.eval/.h9_custody_v2/public_receipts/h9_onetime_evaluation_receipt_v1.json | 1.3953 | — |
Held-out measurement
| Metric | Metallum-1B | Comparators |
|---|---|---|
| Long-context retrieval (RULER)Synthetic needle-in-a-haystack, generated per-seed, so there is no training data to contaminate it.eval/ruler_sniah.py | 86 / 90 | — |
| Distance needlesPerfect retrieval at both D=1024 and D=2032, the second of which is effectively the full 2,048-token window. The NoPE layers are why.eval/kimi_triage/needle_distance.py | 18 / 18 | — |
| Held-out ML-arXiv BPC↓ lower is betterBits per character on a clean held-out arXiv surface. Lower is better.scripts/scorecard.py | 0.7276 | — |
| Structured output (via wrapper)The base weights score 0/60 on strict whole-output JSON/tool-call/MCQ format, and three SFT recipes plus a ReST-EM screen all failed to move it. So structure is enforced at decode time instead, with a token-level JSON grammar and schema-forced keys. Consume the raw weights without the wrapper and you get prose, not JSON.eval/kimi_triage/serve_smoke/format_60_via_endpoint_serve_smoke_v3.json | 60 / 60 |
|
Internal trajectory
| Metric | Metallum-1B | Comparators |
|---|---|---|
| In-domain ML-eng problemsA deliberately hardness-gated 150-task executable suite where small models sit near the floor by design. Metallum ties Qwen3-1.7B at 1.7× its parameter count and clears both other baselines. The suite was used for model selection during development, so it is selection-aware: a real result, not a clean one.eval/ineval_v1.jsonl | 7 / 150 |
|
Not release evidence
| Metric | Metallum-1B | Comparators |
|---|---|---|
| In-domain code suiteTwo exact repeats on an internal 40-task suite with known training overlap. We track it to steer training. It is not a capability claim and we will not present it as one.eval/code_completion_suite_v2.json | 32 / 40 | — |
Scope
A specialist, and nothing else
This is not a general chatbot, and selling it as one would be the fastest way to make it look bad.
Built for
- PyTorch and training-loop scaffolding
- ML and LLM concept explanation
- Long-context retrieval over technical documents
- Structured-output endpoints via the constrained decoder
Out of scope
- General chat
- Non-ML factual question answering
- General-purpose coding
- Safety-critical use
- Autonomous code execution
Limits
What these numbers do not show
If we hid this, you would find it the moment you ran your own evaluation, and then nothing else on this site would be worth believing.
Metallum ties Qwen3-1.7B on our hard in-domain suite: it does not beat it. Both solve 7 of 150. Parity at 1.7× fewer parameters is the claim; superiority is not.
The 32/40 code result runs on a suite with known training overlap. Its own metadata forbids describing it as release-grade evidence.
The promoted checkpoint scores 4/5 on the internal neural-ops category against a preregistered floor of 5/5. The owner adjudicated that exception explicitly rather than moving the floor.
No multiple-choice knowledge score from this lineage is valid: teacher generation prompts embedded real evaluation items, so the whole metric is permanently non-promotable. We report none.
There is no preference tuning, no RLHF, and no safety alignment. Free generation makes local factual slips; verify specifics.
A 1B specialist is not a frontier model. It fits one domain, one budget, and one card, which is the entire proposition.
How we measure
The discipline behind the table
These practices come from research, not marketing. They are the reason we trust our own numbers enough to publish the bad ones.
Preregistered, hash-sealed
Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.
A holdout used exactly once
H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.
Contamination audits that bite
Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.
Failures kept in the record
A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.
Disclosure
During a recursive grep search, the grep process read protected H9 holdout files. Only matching filesystem path names were returned to the agent tool transcript; no protected H9 payload text or excerpts were returned. H9 was therefore retired from use as an untouched final-RC holdout.
We publish this because a holdout that was touched is no longer a holdout, whether or not anything leaked. The replacement used for the qualification above was rebuilt from Wikipedia edits made after training ended, under a fresh custody path, with every raw HTTP response retained and hashed, and a negative-exposure ledger recording that it had not been acquired or scored while training was live. That is what makes the 1.3953 a generalization number rather than a memory test.
This is how we'll measure yours
When we build a specialist on your data, you get the same treatment: held-out suites defined before training, contamination audits against your own corpus, and an honest report on what the model cannot do.