Skip to content
HomeBrewedLabsCommission

Commission

A model that only knows your business

Tell us the domain and what you have to train on. We reply with an honest read on feasibility, scope, and whether this is worth doing at all.

The honest answer here decides whether this is a two-week build or a six-month one.

We reply from a human address. No sequences, no drip campaign.

What you own

Your data
Never mixed into the open dataset or another client's corpus.
Your weights
Delivered to you. No licence-back, no revocation clause.
Your deployment
On-premise, your cloud, or hosted by us. Your call, changeable later.
Your evaluation
The harness ships with the model so you can re-verify every claim.

Good fit

  • You have a body of domain text nobody else has.
  • General models keep getting your vocabulary subtly wrong.
  • Your data cannot legally or safely leave your infrastructure.
  • You need predictable cost at volume, not per-token billing.

Poor fit

  • You need broad general reasoning across many domains.
  • Your data is a few hundred documents.
  • A well-built retrieval system would solve it. We'll tell you if so.

Engagements

Four ways to get a model

Most people arrive asking for the first one and leave having bought the second. We will tell you which of these your corpus actually supports before you commit to anything.

Built from the ground up

Pretraining from random initialization on your corpus: the same path that produced Metallum-1B. You get a model whose every weight came from your data and our recipe, with no upstream licence and no inherited base you have to explain to your legal team.

  • Architecture chosen for your sequence lengths and budget
  • Preregistered, hash-sealed evaluation before training starts
  • Full training receipts: loss curves, checkpoints, contamination audits

Built on our base

Start from our pretrained weights and our prior research instead of from zero. Far cheaper and faster than a ground-up run, and it inherits the parts that were expensive to get right: the stability work, the long-context retrieval behaviour, the decontamination pipeline.

  • Domain adaptation or continued pretraining on your data
  • Inherits the architecture work already paid for and measured
  • Right answer when your corpus is real but not enormous

Custom tokenization

A vocabulary fit on your domain rather than inherited from a general-purpose tokenizer. If your text is full of part numbers, chemical names, statute references, or code, a general tokenizer shreds them into fragments and you pay for it on every single token.

  • Byte-level BPE fit on your corpus, with the sentinels you need
  • Measured compression against whatever you use now
  • Ships with the model, or standalone for a model you already run

Private inference

Serving for a model that cannot leave your control: on-premise, in your cloud, or fully disconnected. Includes the constrained-decoding layer, which is how you get schema-guaranteed JSON and tool calls out of a small model instead of hoping for well-formed prose.

  • Token-level grammar and schema-forced keys, not regex on the output
  • Sized for the hardware you own, including single-card deployments
  • Nothing you send leaves your boundary, and nothing trains on it

No price list, on purpose. Every one of these is scoped off your corpus and your deployment target, so any figure printed here would be a guess wearing a quote's clothing. The scoping call is free and produces a real number.

How a build runs

  1. 01

    Scoping call

    We look at what data you actually have and tell you honestly whether a custom model beats prompting an existing one. Sometimes it doesn't, and you should hear that before you spend anything.

  2. 02

    Corpus build

    We clean, structure, and decontaminate your data, then define held-out evaluation suites before any training starts, so the benchmark cannot be tuned toward after the fact.

  3. 03

    Training

    Pretraining, domain adaptation, or fine-tuning depending on what your data supports. You see the loss curves and the eval trajectory as they happen, not a summary at the end.

  4. 04

    Handover

    You receive the weights, the tokenizer, the evaluation harness, and a model card that documents the limits as clearly as the wins. Deploy it wherever you like, including air-gapped.