Agent systems, defined
The working vocabulary of this site: one quotable definition per term, then the longer version. The same terms govern our Harness Cards, investigations and advisory work.
- Agent harness
An agent harness is the software layer around a model that supplies tools, memory, permissions and control flow, turning a stateless model into a system that can act.
A model on its own predicts text; a harness decides what the model is allowed to do with that prediction. The harness owns the agent loop, the tools the model can call, what it carries between steps and where its actions stop. Two products can ship the same underlying model and behave nothing alike, because the harness differs. That is why this site profiles harnesses rather than models.
- Harness Card
A Harness Card is a version-pinned, field-by-field profile of an agent system, recording the same stable snapshot for every product so two systems can be compared directly.
Every profile records the same fields: category, licence, deployment, model support, tool interfaces, state model, sandbox scope, approvals, traceability, extensibility and the exact version reviewed. When we evaluate a system for a client, the Harness Card is the baseline artifact: the frozen description everything else refers back to.
- Harness Stack
The Harness Stack is the layered model of an agent system, from the model up through the agent loop, tools, context, execution, permissions and observability to organisational control.
A product name is not an architecture. Every analysis identifies the layers of the system and which layer a claim actually belongs to: a rate limit is an execution property, an approval gate is a permissions property, an audit trail is an observability property. The stack keeps those claims from blurring into a single marketing surface.
- Control plane
A control plane is the governance layer above individual agent deployments: identity, policy, budgets, approvals and accountability that no single harness provides on its own.
Frameworks and runtimes decide how one agent behaves; a control plane decides how a fleet of them is run. It answers organisational questions: which teams may deploy which agents, what they may spend, who approved a risky action, and where the evidence lives. Managed examples include LangSmith Deployment, Amazon Bedrock AgentCore and Vertex AI Agent Engine, each profiled in our Harness Cards.
- Agent runtime
An agent runtime is the execution environment where agent code actually runs: sandboxes, isolation, durable execution and recovery.
Agent frameworks decide what code to run; something else has to run it without handing the model a shell on production infrastructure. Runtimes range from durable execution engines like Temporal to sandbox providers like E2B, Daytona and Modal, and they determine what happens when a run fails, pauses or tries to reach the network.
- Sandbox scope
Sandbox scope is the boundary of what an agent’s executed code can reach: the filesystem, network, credentials and host resources inside the wall, and everything outside it.
Two sandboxes can both claim isolation and mean very different things. Scope is the question that separates them: does the code get a process, a container, a microVM or a browser? Can it open network connections, read environment variables or reach the host’s credentials? The Harness Card records scope verbatim from vendor documentation, because the detail is the claim.
- Approvals (human-in-the-loop)
Approvals are the gates where an agent must pause for a human decision before it acts: the core mechanism of human-in-the-loop control.
An approval gate turns an irreversible action into a reviewable request. The properties that matter are where the gate sits in the loop, what context the reviewer is shown, and whether the pause is durable enough to survive a restart. A harness without gates can still be run safely, but the gate then has to come from somewhere else in the stack.
- Traceability
Traceability is the ability to reconstruct what an agent did and why: traces, logs and telemetry that make behaviour inspectable after the fact.
Without traces, an agent’s output is unfalsifiable: you cannot audit it, debug it or defend it to a regulator. Traceability covers which tools were called with which arguments, which model version produced each step and what the run cost. It is the precondition for every other governance claim a vendor makes.
See the terms applied.
Every Harness Card profiles a real system against these definitions, field by field, at a named version.