Skip to main content

Aleph Alpha published Kolibri-1 on 2 October, an Apache-2.0 reasoning model trained for German and English, and within a day the usual question landed in our inbox from two clients at once: can we actually run this ourselves? The answers people reach for are about GPUs and context windows. In our experience what derails a self-hosting project first is duller than that: the download.

So we measured it. We pulled the complete file manifest of 41 open-weight model repositories on Hugging Face, 35 of them from European publishers, and added up every byte on the default branch. Of the 2,356.9 GB sitting on main across those repos, 506.2 GB is a second copy of weights you already have. That is 21.5% of the total, and the standard tooling fetches all of it.

TL;DR

  • We summed every file on main across 41 open-weight model repos: 2,356.9 GB total, of which 506.2 GB (21.5%) is redundant.
  • Ten of the 41 ship more than one full copy of the weights. Six of seven mistralai repos we checked carry a consolidated.safetensors alongside the sharded set, and gpt-oss-20b ships three copies.
  • The repository size the API reports is not the size you download. swiss-ai/Apertus-8B-2509 reports 1,159.7 GB of stored objects against a 16.1 GB checkout, a 72x gap caused by 72 training-checkpoint branches.
  • Three repos ship fp32 as their only copy, so you pay double for weights no inference server will use at that precision.
  • Seven repos are gated, one ships its licence as a PDF, and one with 2,155 downloads declares no licence at all.
  • Every weight file we requested, including from every European publisher, redirected to a US CDN host.

How we measured it

For each repository we called the Hugging Face tree API with recursive=true&expand=true, which returns a per-file size for the default branch without authentication, even on gated repos. We summed those sizes to get the real checkout. Separately we took the exact parameter count from the safetensors metadata the API publishes, so the bytes-per-parameter ratio comes from the publisher’s own figures and not a model card.

We then grouped weight files into format families by directory, filename prefix and extension, and treated the largest single family as the minimum you need to load the model. Everything beyond that family is redundant for inference. We verified the numbers against the live files rather than trusting metadata: a HEAD request for consolidated.safetensors in Mistral-Small-3.2-24B-Instruct-2506 returns a content-length of 48,022,792,280 bytes, and the ten sharded model-* files sum to the same figure.

Ten repos ship the weights twice

The duplication is not random. It is a publishing convention, and it is concentrated:

Repository On main Needed Redundant
meta-llama/Llama-3.3-70B-Instruct 282.2 GB 141.1 GB 141.1 GB
mistralai/Mixtral-8x7B-Instruct-v0.1 190.5 GB 97.1 GB 93.4 GB
mistralai/Mistral-Small-3.1-24B-Instruct-2503 96.1 GB 48.0 GB 48.1 GB
mistralai/Magistral-Small-2509 96.1 GB 48.0 GB 48.0 GB
mistralai/Mistral-Small-3.2-24B-Instruct-2506 96.1 GB 48.0 GB 48.0 GB
mistralai/Devstral-Small-2507 94.3 GB 47.1 GB 47.2 GB
AI-Sweden-Models/gpt-sw3-6.7b-v2-instruct 56.1 GB 28.0 GB 28.0 GB
openai/gpt-oss-20b 41.3 GB 13.8 GB 27.5 GB
openGPT-X/Teuken-7B-base-v0.6 29.8 GB 14.9 GB 14.9 GB
mistralai/Voxtral-Mini-3B-2507 18.7 GB 9.4 GB 9.4 GB

Each duplicate exists for a defensible reason. Mistral ships a consolidated.safetensors for its own reference stack and a sharded set for transformers and vLLM. Meta ships the original PyTorch checkpoint under original/ for people reproducing the paper. gpt-oss-20b goes further, carrying the sharded set, an original/ copy and a metal/ build for Apple silicon, at 13.75 GB each.

None of those are mistakes. The mistake is on the consumer side: the default fetch takes the lot. If your deployment pipeline calls snapshot_download() with no patterns, or clones the repo, you pull both copies onto every node, into every CI cache, and across every egress boundary between the hub and your cluster.

The size the API reports is not the size you download

There is a second, larger trap for anyone sizing a model cache volume. Hugging Face publishes a usedStorage field, and it is the number that surfaces when you ask an API how big a repo is. It counts every stored object across every branch and revision, not the checkout.

The gap is extreme. swiss-ai/Apertus-8B-2509 reports 1,159.7 GB of used storage against 16.1 GB on main, a factor of 72. The cause is visible in the refs endpoint: the repo carries 72 branches, named things like step2600000-tokens14792B and longctx-step3375, each one an intermediate training checkpoint. LumiOpen/Poro-34B does the same at a smaller scale, 752.8 GB stored against a 68.4 GB checkout across 11 token-milestone branches.

This is admirable research practice and a terrible input to a capacity plan. Fifteen of our 41 repos report stored sizes at least 1.5x their checkout. Provision disk from usedStorage and you over-buy by an order of magnitude on two of them. Build a mirroring job that walks all refs and you move a terabyte to serve a 16 GB model.

Three repos ship fp32 and nothing else

A separate 2x, and this one you cannot fix with download flags. Unbabel/TowerInstruct-7B-v0.2 is 27.0 GB for 6.74B parameters. croissantllm/CroissantLLMChat-v0.1 is 5.4 GB for 1.35B. utter-project/EuroMoE-2.6B-A0.6B-Instruct-2512 is 10.5 GB for 2.61B. All three come out at 4.00 bytes per parameter, which is fp32, and in each case it is the only copy in the repo.

You will almost certainly serve these in bf16 or lower, so you transfer and store twice what you will load and then cast at load time. Nothing upstream will do that conversion for you.

The licence column nobody fills in

While we had the manifests open we checked what each repo actually permits, and this is where the “open-weight” label does the most work.

Seven of the 41 are gated, meaning an unauthenticated pull fails and your CI needs a token with an accepted agreement behind it. Four of those seven are European: EuroLLM-9B-Instruct, Velvet-14B, Velvet-2B and gpt-sw3-6.7b-v2-instruct, the last requiring manual approval rather than click-through.

Six repos carry a licence that is not open in the sense a procurement team means. openGPT-X/Teuken-7B-base-v0.6 and Unbabel/TowerInstruct-7B-v0.2 are CC BY-NC 4.0, non-commercial. Mistral-Large-Instruct-2411 is under the Mistral Research Licence. Teuken-7B-instruct-research-v0.4 ships its terms as license.pdf, a 101 KB binary that no licence scanner in your pipeline will read, while its sibling Teuken-7B-instruct-commercial-v0.4 is Apache-2.0 with identical weights. Pick the wrong one from a search result and you have shipped a research-licensed model to production.

And utter-project/EuroLLM-22B-Instruct-Preview, with 2,155 downloads at the time of writing, declares no licence at all: no license key in the card, no licence tag, no licence file. The non-preview release is Apache-2.0. The preview is legally undefined.

A footnote on sovereignty

Since the premise of several of these models is European digital sovereignty, we checked where the bytes come from. Requesting a weight file from Dublin for Kolibri-1, Apertus-8B-Instruct-2509, EuroLLM-22B-Instruct-2512, Teuken-7B-instruct-commercial-v0.4, salamandra-7b-instruct and Pleias-RAG-1B produced the same result every time: a 302 to us.aws.cdn.hf.co.

That is the edge our client was handed, not a claim that no European edge exists. But the weights being Apache-2.0 and the publisher being German, Swiss or Spanish says nothing about the distribution path. If sovereignty is in your requirements rather than your marketing, the mirror into your own EU-hosted object store is not an optimisation. It is the control.

What to do on Monday

Four changes, none of which take a day:

  • Stop fetching whole repos. Pass allow_patterns to snapshot_download(), or --include to the CLI, and name the shard pattern plus the config and tokeniser files. On the ten repos above that halves the transfer immediately.
  • Pin a revision SHA, not a branch. A repo with 72 branches is a repo whose main can move under you. Pin the commit, record it in your deployment manifest, and you get a reproducible model artefact instead of a moving one.
  • Mirror once into your own object store. Convert fp32 to bf16 on the way in, drop the formats you will never load, and serve nodes from inside your own network. Cold starts on autoscaled GPU instances stop being a hub-bandwidth problem.
  • Record the licence as a build artefact. Capture the licence string, the gating status and the revision at fetch time. A PDF licence and a blank licence field both need a human decision, and the moment to make it is before the weights are in production.

Self-hosting an open-weight model is a sound decision for a lot of the workloads we see, particularly where data residency or per-token economics rule out an API. It just is not the one-line decision it looks like from a model card. The parameter count tells you about VRAM. The manifest tells you what the project will actually cost to operate.

REPTILEHAUS builds and runs AI infrastructure for teams who have decided to own their stack, from model selection and inference serving through to the DevOps around it. If you are weighing self-hosting against an API and want the numbers before the commitment, get in touch.

📷 Photo by Teng Yuhong on Unsplash