Skip to content
HomeBrewedLabsCommission

Evidence

Every number, and what it's worth

Benchmarks are easy to game and easier to misread. So each result below carries an evidence class saying what it can support, including the ones that undercut us.

The build

Metallum-1B

Pretraining completed at 16.0B tokens, followed by one bounded knowledge and format SFT stage. The champion checkpoint has passed its frozen batteries and its one-time sealed holdout qualification.

Model

Parameters
1,002M
Layers
26
d_model
1792
Attention
GQA 28/14
d_ff
4864
Vocabulary
40k
Context
2048

The run

GPU
RTX 5090
Count
1
VRAM
32 GB
Wall clock
~13.5 days
Throughput
~18k tok/s
Optimizer
Muon (2D) + AdamW

One consumer card. No cluster, no rented H100 fleet, no FSDP.

Corpus

Tokens seen
16.0B
Effective
19.399B
Unique
8.473B

Custom 40k byte-level BPE with FIM sentinels, fit on the domain. Curated and decontaminated in-house, not a general web scrape.

Architecture

The drawing

Six of twenty-six layers carry no positional encoding at all. That single choice is why a 2,048-token model retrieves perfectly at the edge of its own window.

Metallum-1B decoder architecture blueprintA 26-layer decoder stack. Tokens enter a 40,000-entry tied embedding, pass through 26 identical decoder layers, every fourth of which drops rotary positional encoding entirely, then a final RMSNorm and a tied language-model head. Each layer contains a pre-norm, grouped-query attention with 28 query and 14 key-value heads plus QK-normalization, a residual add, a second pre-norm, a SwiGLU feed-forward of width 4,864, and a second residual add.DECODER STACKembed 40k · tied4812162024×26RMSNormLM head (tied)▮ NoPE layer: no positionalencoding, retrieval headsLAYER DETAIL: NoPElayers 4 / 8 / 12 / 16 / 20 / 24identical to RoPE layers except attentionreceives no positional signal at allRMSNormpre-normAttention · GQA28Q / 14KV · QK-norm⊕ residualRMSNormpre-normSwiGLUd_ff 4864⊕ residuald_model = 1792SPECIFICATIONd_model1792n_layers26heads28 Q / 14 KVd_ff4864vocab40,000context2048RoPE θ500,000NoPEevery 4thembeddingstiedparams1,002,127,360METALLUM-1BML/LLM-ENGINEERING SPECIALISTCKPT s2b075SCALE NTSHOMEBREWEDLABSSHEET 1/1
Decoder stack as built. The six highlighted layers drop rotary encoding entirely. That choice is why a 2,048-token model retrieves perfectly at the edge of its own window.

NoPE every 4th layer

Six of twenty-six layers carry no positional encoding at all. Those layers learn length-generalizing retrieval heads, which is why a 2,048-context model still scores 18/18 on distance needles at the very edge of its window.

QK-norm + z-loss

Normalizing queries and keys before the attention product, plus a z-loss term on the logits, is what kept a 16B-token run from diverging on a single card with no cluster to restart from.

GQA 28Q / 14KV

Halving the key-value heads halves the KV cache. That is the difference between a long-context model that fits on consumer hardware and one that does not.

Document-masked attention

Attention never crosses a document boundary inside a packed sequence, so the model never learns to continue one document into an unrelated one.

FIM sentinels

Fill-in-the-middle sentinels in a 40k byte-level BPE vocabulary, fit on the target domain rather than inherited from a general-purpose tokenizer.

Muon on 2D parameters

Muon for matrices, AdamW for everything else, under a WSD schedule with short-sequence early phases. Throughput held at roughly 18k tokens/second.

Metallum-1B training corpus compositionA dimensioned bar showing the pretraining mix by effective tokens: Code 53.6%, Knowledge 19.6%, Reasoning 15.2%, Math 11.6%. Total 19.399B effective tokens from 8.473B unique tokens, seen over 16.0B tokens of training on 1 × NVIDIA RTX 5090.CORPUS SCHEDULE: BY EFFECTIVE TOKEN19.399B effective · 8.473B unique · 16.0B seen53.6%Code19.6%Knowledge15.2%Reasoning11.6%MathHARDWARERTX 5090WALL CLOCK~13.5 daysTHROUGHPUT~18k tok/sREPLAY~1.04B tokens, receipted
Curated and decontaminated in-house, not a general web scrape. ~1.04B tokens replayed after two hardware incidents, fully receipted: two hardware incidents mid-run, both accounted for rather than quietly re-rolled.

Reading the table

Not all numbers weigh the same

A score measured on data the model already saw is not evidence of capability, no matter how good it looks. We separate the two rather than averaging them into one flattering figure.

Sealed holdout

Measured once, on a holdout that was built, sealed, and never trained on. One run, no retries, result published whatever it said. This is the generalization claim.

Held-out measurement

Measured on surfaces the model never trained on. Strong capability evidence, repeatable on demand.

Internal trajectory

Used to steer training. Selection-aware or previously observed, so it cannot be read as a clean capability estimate.

Not release evidence

The suite itself declares known overlap with training inputs. Published for completeness, never as a capability claim.

Results

The full table

Champion checkpoint s2b075, measured by frozen evaluators with sealed checkpoint identity.

Sealed holdout

Sealed holdout

Sealed holdout results for Metallum-1B
MetricMetallum-1BComparators
H9 sealed holdout↓ lower is betterBits per byte on 400k characters of post-cutoff Wikipedia, built and sealed before the run and evaluated exactly once. Contamination verified at 0.0099% shingle containment. One shot, no retries. This is the number that carries the release.eval/.h9_custody_v2/public_receipts/h9_onetime_evaluation_receipt_v1.json1.3953—
Held-out measurement

Held-out measurement

Held-out measurement results for Metallum-1B
MetricMetallum-1BComparators
Long-context retrieval (RULER)Synthetic needle-in-a-haystack, generated per-seed, so there is no training data to contaminate it.eval/ruler_sniah.py86 / 90—
Distance needlesPerfect retrieval at both D=1024 and D=2032, the second of which is effectively the full 2,048-token window. The NoPE layers are why.eval/kimi_triage/needle_distance.py18 / 18—
Held-out ML-arXiv BPC↓ lower is betterBits per character on a clean held-out arXiv surface. Lower is better.scripts/scorecard.py0.7276—
Structured output (via wrapper)The base weights score 0/60 on strict whole-output JSON/tool-call/MCQ format, and three SFT recipes plus a ReST-EM screen all failed to move it. So structure is enforced at decode time instead, with a token-level JSON grammar and schema-forced keys. Consume the raw weights without the wrapper and you get prose, not JSON.eval/kimi_triage/serve_smoke/format_60_via_endpoint_serve_smoke_v3.json60 / 60
  • 0 / 60Raw base weights
Internal trajectory

Internal trajectory

Internal trajectory results for Metallum-1B
MetricMetallum-1BComparators
In-domain ML-eng problemsA deliberately hardness-gated 150-task executable suite where small models sit near the floor by design. Metallum ties Qwen3-1.7B at 1.7× its parameter count and clears both other baselines. The suite was used for model selection during development, so it is selection-aware: a real result, not a clean one.eval/ineval_v1.jsonl7 / 150
  • 7 / 150Qwen3-1.7B
  • 3 / 150Qwen3-0.6B
  • 0 / 150SmolLM2-1.7B
Not release evidence

Not release evidence

Not release evidence results for Metallum-1B
MetricMetallum-1BComparators
In-domain code suiteTwo exact repeats on an internal 40-task suite with known training overlap. We track it to steer training. It is not a capability claim and we will not present it as one.eval/code_completion_suite_v2.json32 / 40—

Scope

A specialist, and nothing else

This is not a general chatbot, and selling it as one would be the fastest way to make it look bad.

Built for

  • PyTorch and training-loop scaffolding
  • ML and LLM concept explanation
  • Long-context retrieval over technical documents
  • Structured-output endpoints via the constrained decoder

Out of scope

  • General chat
  • Non-ML factual question answering
  • General-purpose coding
  • Safety-critical use
  • Autonomous code execution

Limits

What these numbers do not show

If we hid this, you would find it the moment you ran your own evaluation, and then nothing else on this site would be worth believing.

  • Metallum ties Qwen3-1.7B on our hard in-domain suite: it does not beat it. Both solve 7 of 150. Parity at 1.7× fewer parameters is the claim; superiority is not.

  • The 32/40 code result runs on a suite with known training overlap. Its own metadata forbids describing it as release-grade evidence.

  • The promoted checkpoint scores 4/5 on the internal neural-ops category against a preregistered floor of 5/5. The owner adjudicated that exception explicitly rather than moving the floor.

  • No multiple-choice knowledge score from this lineage is valid: teacher generation prompts embedded real evaluation items, so the whole metric is permanently non-promotable. We report none.

  • There is no preference tuning, no RLHF, and no safety alignment. Free generation makes local factual slips; verify specifics.

  • A 1B specialist is not a frontier model. It fits one domain, one budget, and one card, which is the entire proposition.

How we measure

The discipline behind the table

These practices come from research, not marketing. They are the reason we trust our own numbers enough to publish the bad ones.

01

Preregistered, hash-sealed

Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.

02

A holdout used exactly once

H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.

03

Contamination audits that bite

Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.

04

Failures kept in the record

A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.

Disclosure

During a recursive grep search, the grep process read protected H9 holdout files. Only matching filesystem path names were returned to the agent tool transcript; no protected H9 payload text or excerpts were returned. H9 was therefore retired from use as an untouched final-RC holdout.

We publish this because a holdout that was touched is no longer a holdout, whether or not anything leaked. The replacement used for the qualification above was rebuilt from Wikipedia edits made after training ended, under a fresh custody path, with every raw HTTP response retained and hashed, and a negative-exposure ledger recording that it had not been acquired or scored while training was live. That is what makes the 1.3953 a generalization number rather than a memory test.

This is how we'll measure yours

When we build a specialist on your data, you get the same treatment: held-out suites defined before training, contamination audits against your own corpus, and an honest report on what the model cannot do.