Independent research into agent systems
Independent research on responsible agent systems.
Harness the Agents is an independent, evidence-led publication and research lab. We study how AI models become agent systems and how those systems are operated responsibly: across harnesses, runtimes, control planes, governance, permissions, memory, execution, observability, and recovery.

Featured research
View all researchInside
DeepSeek Harness Isn't a Claude Code Killer. It's Something More Interesting.
DeepSeek's open-source agent harness passed 168,000 GitHub stars in its first week. I installed it, read the repo, and ran it myself. Here is what holds up.
Above
Amodei's Embedded Evaluators: Access Is the Whole Game
Amodei pledges embedded evaluators with employee-like access and publication rights. What the pledge promises, what it leaves open, and what is on paper today.
Inside
An Agent That Works for Weeks Needs More Than Memory
Salesforce's long-horizon runtime lets an Agentforce agent pursue a goal for weeks. Memory, durable execution and dynamic steering are documented. The authority questions are not.
Advisory
Need this rigour on your own systems?
We evaluate, design and govern agent systems with the same rigour we publish: version-pinned assessments, harness selection and control-plane architecture.
Editorial sections
Inside the Harness
Version-pinned teardowns of agent systems: the loop, tools, memory, permissions and execution boundaries, examined in the source at a named version.
Above the Harness
Enterprise control architecture for agent systems: identity, policy, approvals, budgets, observability and accountability that no single harness provides.
Harness Test
Reproducible controlled experiments: same task, pinned versions, published prompts, environments, traces and limitations.
Unharnessed
Evidence-led analysis of agent system failures: what broke, the mechanism behind it, the blast radius and the mitigation.
Agent Field Notes
A weekly intelligence briefing: what shipped and what it means for the people running these systems.
Reproducible experiments
All experiments are documented end-to-end so results can be verified and built upon.
Methodology
Design & protocol
Versions
Code & deps
Models
Providers & configs
Metrics
What we measure
Traces
Execution logs
Results
Analysis & stats
Limitations
Threats & caveats
Raw data
Downloads
Agent Field Notes
A periodic briefing that summarises our latest articles and links back to the canonical research.
We respect your inbox. Unsubscribe anytime.
