Skip to main content

Platform

Security model

This page sets out the capabilities we assume an attacker to have, the controls we apply in response, and the cases in which a mitigation reduces risk without eliminating it.

Assumptions

  • · Any content the system retrieves may contain instructions intended to subvert it.
  • · Any capability an agent can reach may be attempted on an attacker’s behalf.
  • · Users are authenticated but not uniformly entitled, and entitlements change faster than indexes refresh.
  • · The audit trail is itself sensitive and is subject to the same controls as the data it describes.
  • · Defences degrade without visible symptoms unless they are exercised by tests capable of failing the build.

Threats and mitigations

ThreatIf it succeedsWhat we do
Prompt injection via retrieved contentAgent follows instructions embedded in a documentRetrieved content treated as untrusted input throughout; tool scopes bound to the task, not the system; irreversible actions behind human gates; injection cases in the regression suite.
Tool abuse / confused deputyAgent uses a legitimate capability on the attacker’s behalfPer-agent scoped credentials rather than a service account with standing access; typed tool contracts that fail rather than improvise; the calling principal recorded on every invocation.
Entitlement bypass through retrievalUser receives content, or an answer derived from content, they may not readEntitlement predicate pushed into the query; cache keys include the principal; citation lists filtered so refusals do not disclose document existence.
Model or artefact tamperingA modified model or dependency runs in productionSigned manifests verified before admission to the internal registry, fail closed; digests carried into deployment records; SBOM for everything that crosses a boundary.
Data exfiltration through tool egressSensitive content leaves via a tool with outbound accessDefault-deny egress at the network layer; tools that require outbound access enumerated and justified individually; egress attempts logged and alerted.
Trace and log disclosureThe audit trail becomes the leakPayloads stored by hash where the content is regulated; retention and access controls on the trace store equal to the source data; derived summaries classified with the source.
Runaway cost or loopAn agent consumes budget or capacity without producing valueHard token and turn budgets per node; declared stopping conditions; escalation rather than continuation when a budget is exceeded.

Supply chain

Images are built outside the boundary and transferred as complete artefacts rather than resolved from a partial mirror inside it. Dependencies are pinned by digest, a software bill of materials is generated at build time and retained alongside the artefact, and signatures are verified before admission on a fail-closed basis. In air-gapped environments this is established as a documented procedure, reviewed once by the security team, platform team and change board together, rather than as an exception raised for each individual crossing.

Logging and retention

Traces are append-only, independent of any framework, and classified at the same level as the data they describe. Payloads may be recorded by hash where the content itself cannot be retained, which allows a trace to remain valid even when its subject matter has been deleted. Access to the trace store is itself audited, as in most of our engagements it is the most sensitive repository in the system.

Review your threat model with us

Where you already have a threat model, we would prefer to begin from yours rather than ours. Where you do not, developing one is a well-scoped first phase of roughly two weeks.