Insights
On-prem inference
2 articles tagged On-prem inference.
18 June 2026 · 4 min · QAI Labs engineering
Sizing GPUs for a 70B model you have to host yourself
The arithmetic that decides whether your cluster survives Monday morning (weights, KV cache, headroom) and the three assumptions that most often make it wrong.
11 February 2026 · 3 min · QAI Labs engineering
Getting model weights into an air-gapped enclave without breaking accreditation
The file transfer is the easy part. The hard part is proving, eighteen months later, exactly which artefact ran on which day, and building a path that a security team will approve twice.
Start with the constraint.
Most of these projects are shaped by what you cannot do rather than what you want. Data that cannot leave the estate, a model you cannot host with a third party, a decision somebody has to justify to a regulator. Tell us yours and we will say honestly whether we can work inside it.