Flagship service
On-premises AI engineering
Production language-model systems running on hardware you own, in networks you control, integrated with the identity and monitoring you already operate.
We take a model you are permitted to run and turn it into a service your organisation can depend on: sized against your real traffic, served on your infrastructure, integrated with your directory, monitored by your platform team, and rehearsed for the day it fails.
Most of the difficulty is not the model. It is capacity planning that survives month-end, retrieval that respects entitlements at query time, a release path that fits a change window, and a rollback somebody has actually practised. That is the work.
What you get
Named artefacts, not a slide pack. Each one is a thing your team can open, run, or hand to an auditor.
| Artefact | What it contains |
|---|---|
| Capacity model | Weights, KV cache, activation overhead and tail headroom for your traffic shape, with the assumptions written down so it can be re-run when they change. |
| Serving baseline | Infrastructure-as-code and Helm/Kustomize configuration for the inference tier, GPU scheduling, autoscaling signals, and graceful shutdown. |
| Retrieval layer | Ingestion, chunking, hybrid search and reranking, with entitlement filtering pushed into the query rather than applied afterwards. |
| Identity integration | OIDC or SAML against your existing directory. No external identity provider introduced. |
| Observability pack | Metrics, traces and dashboards covering queue depth, cache utilisation, tail latency and cost per resolved task. |
| Evaluation harness | A regression suite that runs in your pipeline and blocks changes that reduce measured quality. |
| Architecture decision records | What we chose, what we rejected, and what would make us revisit it. |
| Runbook | Written against failures that actually occurred during the engagement, not imagined ones. |
How it runs
01 · 1–2 weeks
Frame
Constraints, traffic shape, hardware inventory, data residency and risk classification. Ends with a feasibility note and costed options.
02 · 3–6 weeks
Prove
A thin vertical slice running on your infrastructure with evaluation attached from the first week.
03 · 8–16 weeks
Build
Hardening: identity, observability, CI/CD, security review, load testing against your peak rather than your average.
04 · 2–4 weeks
Hand over
Your team drives, we watch. Rollback rehearsal, incident simulation, and the last documentation written where they hesitated.
What we use
Chosen per engagement against your constraints. We have no reseller relationships and no incentive to recommend one of these over another.
Serving & inference
Infrastructure
Retrieval & data
Observability
Identity & security
What we do not do
Knowing where our usefulness stops saves everyone a procurement cycle.
- We do not train foundation models from scratch. We adapt, serve, evaluate and orchestrate models that already exist, and we will tell you when fine-tuning is not the answer to your problem.
- We do not sell hardware, resell licences, or hold reseller relationships with any vendor named on this page. There is nothing in it for us if you choose one engine over another.
- We do not take on engagements where an on-premises approach is clearly the wrong answer. If your data can leave and a hosted API would serve you better, that is what we will say.
- We do not provide 24/7 managed operations. We build the system and train your team to run it; ongoing support, where wanted, is an explicit and separate arrangement.
Questions we are asked
How many GPUs will we need?
Nobody can answer that from a model name alone. It depends on parameter count, precision, context length, concurrency and your tail. We produce a capacity model during Frame with the arithmetic and assumptions written down, and validate it under load during Prove. In practice the most common outcome is that a smaller, well-quantised model with good retrieval meets the requirement on considerably less hardware than first assumed.
Can you work in a genuinely air-gapped environment?
Yes, and the transfer procedure is a deliverable in its own right. Signed manifests, verification inside the enclave before admission to an internal registry, batched update cycles, and a rollback artefact kept resident so recovery does not require another crossing. We write that procedure as a document your security team, platform team and change board review together, once, before the first transfer.
Will this lock us into your framework?
The trace format and evaluation artefacts are designed to outlive any framework, including ours. The serving tier is standard open components on your Kubernetes. The exit plan is agreed during Frame and tested during handover by having your team drive while we watch.
What if our existing platform team disagrees with your design?
They usually have context we lack, and they will be operating the result. We work in your repositories under your review process, and material decisions go into architecture decision records where they can be argued with in writing.
How do you handle model updates after go-live?
Through the same evaluation gate as any other change. A new model version is a candidate, not an upgrade, until the regression suite says otherwise on your task set. In constrained environments we batch model, dependency and base-image updates into a single planned crossing.
Start with the constraint.
Most of these projects are shaped by what you cannot do rather than what you want. Data that cannot leave the estate, a model you cannot host with a third party, a decision somebody has to justify to a regulator. Tell us yours and we will say honestly whether we can work inside it.