Rough memory by model size
These are approximations for the weights alone. Real use adds memory for the conversation (the context window) and for each simultaneous user, so size above these figures.
| Model size | Weights at 16-bit | Weights at 4-bit | Typical hardware class |
|---|---|---|---|
| 7 to 8 billion parameters | about 15 GB | about 4 to 5 GB | One mid-range GPU |
| 13 to 14 billion | about 27 GB | about 8 GB | One 16 to 24 GB GPU |
| 30 to 34 billion | about 65 GB | about 18 to 20 GB | One 24 GB GPU (tight) or a 48 GB GPU |
| 70 billion | about 140 GB | about 38 to 42 GB | A 48 GB-class GPU, or two 24 GB GPUs |
What else decides the hardware
- Concurrent users: each active conversation needs working memory, so ten users is not one user times ten in speed or in memory.
- Context length: long documents in the prompt, which RAG produces, use much more memory than short chats.
- Speed: tokens per second depends on memory bandwidth as well as capacity.
- The rest of the machine: system memory, fast storage for model files, power and cooling.
- Apple Silicon is an option for small teams because memory is shared, at lower speed than a datacenter GPU.
- Private cloud: a dedicated GPU instance in your own cloud account keeps the data under your control without buying hardware.
When running locally is not worth it
At low volume a hosted model under a contract that forbids training on your data and sets retention limits can be cheaper and better than hardware you maintain. Local deployment earns its cost when the data cannot leave your control, when volume is steady and high, or when you need to tune the model.
How we size it
We test the actual workload, with your documents and your expected number of users, on candidate models before recommending hardware. The number that matters is not the largest model that fits but the smallest one that passes your test set.