Skip to content

Methodology

How we examine agent systems.

Every Harness the Agents investigation follows a published method: pinned versions, declared evidence and visible uncertainty. The same method governs our client work.

The Harness Stack

A product name is not an architecture. Every analysis identifies the layers of the system — and which layer a claim actually belongs to.

  1. 01

    Model

    Reasoning, generation and inference — the capability the system is built on.

  2. 02

    Agent loop

    Planning, iteration and stop conditions that turn a model into an agent.

  3. 03

    Tools and skills

    The interfaces through which the agent reads and changes the world.

  4. 04

    Context and memory

    What the agent carries between steps, sessions and tasks — and who can read it back.

  5. 05

    Execution

    Where code actually runs: sandboxes, runtimes, isolation and recovery.

  6. 06

    Permissions

    What the agent is allowed to do, and how those boundaries are enforced.

  7. 07

    Observability

    Traces, logs and telemetry that make behaviour inspectable after the fact.

  8. 08

    Organisational control

    Identity, policy, budgets, approvals and accountability — the layer no single harness provides.

The Harness Card

Every product profile records the same stable snapshot, so two systems can be compared field by field — and a profile can be revisited version by version.

When we evaluate a system for a client, the Harness Card is the baseline artifact: the frozen description everything else refers back to.

  • Category
  • Licence
  • Deployment
  • Model support
  • Tool interfaces
  • State model
  • Sandbox scope
  • Approvals
  • Traceability
  • Extensibility
  • Version / commit
  • Review date

The Harness Test protocol

Every controlled experiment publishes its complete record. If a result cannot be reproduced from what we publish, it does not deserve trust.

We avoid generic leaderboards. A Harness Test answers a narrow question under a frozen contract — and shows its working.

  • 1Hypothesis
  • 2Exact versions and models
  • 3Environment
  • 4Prompts
  • 5Controlled variables
  • 6Metrics
  • 7Traces
  • 8Results
  • 9Failures
  • 10Limitations
  • 11Reproduction instructions
  • 12Raw data

Incident status

Every Unharnessed analysis carries an explicit evidence status. Readers always know how established a claim is.

Confirmed

Verified against primary evidence.

Reproduced

Independently reproduced in our own environment.

Reported

A credible report that has not yet been verified.

Disputed

Contested by the vendor or other parties.

Resolved

Fixed or mitigated in a named version.

Historical

No longer applicable to current versions.

Unverified

Insufficient evidence. A GitHub issue alone does not establish a vulnerability.

Evidence rules

The standing rules behind every claim we publish — and every recommendation we make to clients.

Pinned claims

Teardowns and claims are pinned to versions, commits, retrieval dates and deployment assumptions.

Vendor statements

Treated as self-asserted primary evidence, never as independent confirmation.

Social posts

A discovery input. Stronger evidence is required before anything is asserted.

Model versus harness

Model capability is reported separately from the behaviour of a specific model–harness configuration.

Visible uncertainty

Methods, traces, limitations and open questions are published alongside the findings.

Human gate

Every piece is commissioned and approved for publication by a human editor.

For client teams

The same rigour, applied to your systems.

Advisory engagements follow this methodology end to end — pinned versions, declared evidence, visible uncertainty and recommendations you can inspect.