Metallum-1B is live. Open weights on Hugging Face. Apache 2.0. 1,002M parameters. Brewed on 1 × NVIDIA RTX 5090. 16.0B tokens, curated.
The frontier isn't a datacenter.
It's home brewed.
Open-weight models designed for sovereign deployment.
We brewed a 1-billion-parameter specialist from scratch on a single consumer graphics card. No cluster. No rented fleet. Small batch, full strength: a model fermented on your data, owned entirely by you.
The thesis
Big models know
a little about
everything.
Your business does not need a model that can write sonnets and debug Kubernetes. It needs one that knows your domain: the vocabulary, the edge cases, the twenty years of institutional judgment sitting in your archives.
That model does not exist, because nobody is going to build it for you. Frontier labs optimize for the average of everything. Your domain is not the average of anything.
So build it yourself. A focused 1B specialist trained on the right data beats a general giant inside its lane, and it fits on hardware you can actually own.
What we brew
Three taps
Pick your pour: sovereign AI, closed or open, whichever your situation demands. The point is that the choice belongs to you.
Commission a specialist
Bring your domain data. We train a model that lives inside your business: your corpus, your weights, your IP. It runs where you say it runs, including fully disconnected.
- You own the weights outright
- On-premise, your cloud, or hosted by us
- Nothing you send trains anyone else's model
Host your own model
The platform we're building: bring a model you trained, whether ours, yours, or a fine-tune, and serve it without renting a hyperscaler or surrendering your data to an API.
- Serve open or closed weights
- Built for one-card budgets, not clusters
- Early access opening in waves
Feed the distillery
Specialist models die on thin data. Donate structured data in a domain you know and the distillery turns it into an open dataset anyone can train on, with provenance attached, not scraped.
- Openly licensed, permanently
- Provenance and PII audited on intake
- Contributors credited by name
Proof of work
First batch, brewed on our own bench
Metallum-1B is our own model: a 1,002M-parameter ML / LLM-engineering specialist, pretrained from random initialization on one RTX 5090 in ~13.5 days. Here is what it measures.
- Architecture
- 26L · d1792 · GQA 28/14
- Context
- 2048 tokens
- Optimizer
- Muon (2D) + AdamW
- Schedule
- WSD, short-sequence curriculum
- Hardware
- 1 × NVIDIA RTX 5090 · 32 GB · Blackwell sm_120
Headline measurements
The first was measured once, on a holdout sealed before training and opened after. The rest are held-out or generated per-seed, so none of them can be contaminated by training data.
- H9 sealed holdout1.3953
Bits per byte on 400k characters of post-cutoff Wikipedia, built and sealed before the run and evaluated exactly once. Contamination verified at 0.0099% shingle containment. One shot, no retries. This is the number that carries the release.
- In-domain ML-eng problems7 / 150Qwen3-1.7B 7 / 150Qwen3-0.6B 3 / 150SmolLM2-1.7B 0 / 150
A deliberately hardness-gated 150-task executable suite where small models sit near the floor by design. Metallum ties Qwen3-1.7B at 1.7× its parameter count and clears both other baselines. The suite was used for model selection during development, so it is selection-aware: a real result, not a clean one.
- Long-context retrieval (RULER)86 / 90
Synthetic needle-in-a-haystack, generated per-seed, so there is no training data to contaminate it.
- Distance needles18 / 18
Perfect retrieval at both D=1024 and D=2032, the second of which is effectively the full 2,048-token window. The NoPE layers are why.
We also publish the numbers that do not favour us, and say plainly why they cannot be read as capability claims.
How it's built
Drawn, not assembled
Every dimension below came from a decision with a reason behind it. This is the drawing that decision set produced.
NoPE every 4th layer
Six of twenty-six layers carry no positional encoding at all. Those layers learn length-generalizing retrieval heads, which is why a 2,048-context model still scores 18/18 on distance needles at the very edge of its window.
QK-norm + z-loss
Normalizing queries and keys before the attention product, plus a z-loss term on the logits, is what kept a 16B-token run from diverging on a single card with no cluster to restart from.
GQA 28Q / 14KV
Halving the key-value heads halves the KV cache. That is the difference between a long-context model that fits on consumer hardware and one that does not.
Document-masked attention
Attention never crosses a document boundary inside a packed sequence, so the model never learns to continue one document into an unrelated one.
FIM sentinels
Fill-in-the-middle sentinels in a 40k byte-level BPE vocabulary, fit on the target domain rather than inherited from a general-purpose tokenizer.
Muon on 2D parameters
Muon for matrices, AdamW for everything else, under a WSD schedule with short-sequence early phases. Throughput held at roughly 18k tokens/second.
How we work
Measurement you can audit
Anyone can post a benchmark. The hard part is running an evaluation that could have embarrassed you, and then publishing it when it does.
Preregistered, hash-sealed
Every evaluation design is written, SHA-256 sealed, and independently reviewed before a checkpoint is scored. You cannot move the goalposts after seeing the result if the goalposts are hashed.
A holdout used exactly once
H9 was built, sealed, and never trained on, then evaluated one time with the checkpoint hash, evaluator hash, and tokenizer hash all recorded in the receipt. Whatever it returned was the number we would publish. It returned 1.3953.
Contamination audits that bite
Decontamination runs against the full eval surface. One Stack ingest dropped 18,564 documents, 4.40% of it, as benchmark clones. H9 itself was verified at 0.0099% shingle containment before it was allowed to count.
Failures kept in the record
A touched holdout was retired rather than reused. A cleanroom retrain that failed its gate is published as a failed experiment. A frozen classifier that counted mentions of "C++" as C++ training data was corrected against our own prior result.
What's your next batch?
Tell us the domain and what data you're sitting on, and we'll tell you what a specialist brewed on it could know. If it's the wrong answer for you, we will say so. That conversation is free and takes one email.
