Methodology
How we examine agent systems.
Every Harness the Agents investigation follows a published method: pinned versions, declared evidence and visible uncertainty. The same method governs our client work.
The Harness Stack
A product name is not an architecture. Every analysis identifies the layers of the system — and which layer a claim actually belongs to.
- 01
Model
Reasoning, generation and inference — the capability the system is built on.
- 02
Agent loop
Planning, iteration and stop conditions that turn a model into an agent.
- 03
Tools and skills
The interfaces through which the agent reads and changes the world.
- 04
Context and memory
What the agent carries between steps, sessions and tasks — and who can read it back.
- 05
Execution
Where code actually runs: sandboxes, runtimes, isolation and recovery.
- 06
Permissions
What the agent is allowed to do, and how those boundaries are enforced.
- 07
Observability
Traces, logs and telemetry that make behaviour inspectable after the fact.
- 08
Organisational control
Identity, policy, budgets, approvals and accountability — the layer no single harness provides.
The Harness Card
Every product profile records the same stable snapshot, so two systems can be compared field by field — and a profile can be revisited version by version.
When we evaluate a system for a client, the Harness Card is the baseline artifact: the frozen description everything else refers back to.
- Category
- Licence
- Deployment
- Model support
- Tool interfaces
- State model
- Sandbox scope
- Approvals
- Traceability
- Extensibility
- Version / commit
- Review date
The Harness Test protocol
Every controlled experiment publishes its complete record. If a result cannot be reproduced from what we publish, it does not deserve trust.
We avoid generic leaderboards. A Harness Test answers a narrow question under a frozen contract — and shows its working.
- 1Hypothesis
- 2Exact versions and models
- 3Environment
- 4Prompts
- 5Controlled variables
- 6Metrics
- 7Traces
- 8Results
- 9Failures
- 10Limitations
- 11Reproduction instructions
- 12Raw data
Incident status
Every Unharnessed analysis carries an explicit evidence status. Readers always know how established a claim is.
Confirmed
Verified against primary evidence.
Reproduced
Independently reproduced in our own environment.
Reported
A credible report that has not yet been verified.
Disputed
Contested by the vendor or other parties.
Resolved
Fixed or mitigated in a named version.
Historical
No longer applicable to current versions.
Unverified
Insufficient evidence. A GitHub issue alone does not establish a vulnerability.
Evidence rules
The standing rules behind every claim we publish — and every recommendation we make to clients.
Pinned claims
Teardowns and claims are pinned to versions, commits, retrieval dates and deployment assumptions.
Vendor statements
Treated as self-asserted primary evidence, never as independent confirmation.
Social posts
A discovery input. Stronger evidence is required before anything is asserted.
Model versus harness
Model capability is reported separately from the behaviour of a specific model–harness configuration.
Visible uncertainty
Methods, traces, limitations and open questions are published alongside the findings.
Human gate
Every piece is commissioned and approved for publication by a human editor.
For client teams
The same rigour, applied to your systems.
Advisory engagements follow this methodology end to end — pinned versions, declared evidence, visible uncertainty and recommendations you can inspect.