Skip to main content

Applied AI engineering

AI Assisted Engineering (AIAE) Best Practices, Governance and Privacy

QAI Labs designs and builds production Sovereign AI systems, together with the agentic governance frameworks.

Engineering, data and governance disciplines · trading since 2021 · Glastonbury, United Kingdom

The problem we work on

Data that cannot leave the estate

Regulation, classification, contractual terms or risk appetite often require the workload to run where the data already sits. Most vendor architectures assume the opposite.

The distance between a prototype and a service

A working prompt is not an evaluated, permissioned and monitored service with a named owner, a rollback path and support arrangements behind it.

Autonomy that has to be governed

Agent systems are only useful in regulated environments when their behaviour is bounded, observable and reversible. Approval depends on being able to explain what the system did and why.

What we do

Our core capabilities

Our work centres on two disciplines. The remaining services exist to support them, and we will tell you at the outset if your requirement falls outside what we do well.

Reference architecture

A reference architecture that runs within your boundary

The stack is arranged in six layers and requires no outbound network dependency. Most of our engagements converge on this structure, with the contents of each layer determined by your environment and constraints.

On-premises reference architectureSix layers inside the client security boundary. From the bottom: infrastructure (GPU nodes, block and object storage, network boundary); serving (inference engines, model router, KV and prompt cache); data and retrieval (vector index, document store, connectors); agent runtime (planner, tool broker, memory, policy engine, state store); governance (identity, authorisation, audit log, evaluation, observability), which spans the full width; and interfaces (APIs, internal UI, existing line-of-business systems). Nothing crosses the boundary outward.client security boundary, no egressINTERFACESInternal APIsOIDC-protectedOperator consoletraces, approvalsLine-of-business systemsITSM, CRM, EDRMBatch / scheduled jobsAGENT RUNTIMEPlannergraph executionTool brokertyped contracts, MCPPolicy engineleast privilegeMemoryworking / episodicState storecheckpoint + replayDATA & RETRIEVALVector indexhybrid + rerankDocument storeACL-awareFeature storeSource connectorsSERVINGInference enginevLLM / SGLang / NIMModel routersize-to-taskPrompt + KV cacheEmbedding + rerankINFRASTRUCTUREGPU nodes (Kubernetes, MIG-aware)Block + object storageSegmented network, no outbound routeGovernance plane · identity (AD/LDAP, OIDC) · authorisation · immutable audit log · evaluation · OpenTelemetry
Figure 1. The reference stack. Model weights, execution traces and logs all remain inside the boundary, and the governance plane spans every layer rather than sitting alongside them.
See the full architecture

How we engage

A four-phase engagement with a defined exit

Handover criteria are agreed during the first phase, before development begins. Establishing them early has a direct influence on how the system is designed and built.

Engagement methodFour sequential phases. Frame, one to two weeks, produces a feasibility note, risk register and costed options. Prove, three to six weeks, produces a thin vertical slice running on the client's infrastructure with an evaluation baseline. Build, eight to sixteen weeks, produces a hardened system with test suites and runbooks. Hand over, two to four weeks, leaves the client's team operating the system with an agreed exit plan.01 · 1–2 weeksFrame
Feasibility note, risk register, costed options
02 · 3–6 weeksProve
Thin vertical slice on your infrastructure, evaluation baseline
03 · 8–16 weeksBuild
Hardened system, test and eval suites, runbooks
04 · 2–4 weeksHand over
Your team operating it, exit plan agreed
How we work in detail

Selected work

Selected engagement

We publish a small number of engagements in full detail rather than a list of client logos. The following account covers the system we run our own operations on, including the design decisions we would revisit.

All case studies

Start with the constraint.

Most of these projects are shaped by what you cannot do rather than what you want. Data that cannot leave the estate, a model you cannot host with a third party, a decision somebody has to justify to a regulator. Tell us yours and we will say honestly whether we can work inside it.