Skip to content
Insights
Inside

Anthropic's Model Hardware Standard: The Hard Part of Physical AI Was Never the Model

Anthropic's Model Hardware Standard gives AI agents a standard way to run lab and factory hardware. How it sits below MCP, and what the research preview doesn't answer.

By Adam Maguire Wilson16 min read
On this page

The robot demos coming out of this city look effortless. A humanoid from Unitree, whose offices are a short ride from my apartment, will do something on a factory floor that would have been science fiction three years ago, and the clip makes the rounds and everyone argues about the model. What the clip doesn't show is the fortnight of integration work before the camera rolled: the robot speaks one protocol, the conveyor speaks another, and the safety PLC speaks a third, and a human being with a laptop had to teach all three to get along.

That unglamorous gap is what Anthropic is now aiming at. On 27 August it opened a research preview of the Model Hardware Standard (MHS), a shared specification that lets AI agents operate physical devices: microscopes, liquid handlers, robotic arms, and further out, the machinery of advanced manufacturing. I've been through the announcement and the early coverage, and I think the interesting story isn't that agents are getting hands. It's where Anthropic has chosen to put the safety, and how much of the hard part still has no answer.

Key Takeaways - MHS is a shared specification, in application-only research preview, for AI agents to operate physical devices. Each device gets a standardised driver with simple read and write primitives plus a generated reference file describing what it can measure, what can be adjusted, and which safety limits are enforced. - It sits below and beside MCP rather than replacing it: MCP can carry the tool call from the agent, MHS makes the machine behind that call understandable and consistently operable. - Safety is designed into the driver layer. Devices declare their bounds, interlocks and emergency stops, and agents operate within them by default, because Anthropic concedes Claude's physical reasoning still needs expert oversight. - Early pilots (Genentech, Carnegie Mellon, QuEra, the University of Washington, HHMI Janelia) are real but partner-reported, and every detailed example runs on Claude. - There is no public spec, SDK, licence or governance model yet, no published security model, and no date for open sourcing. Treat the preview as a direction of travel, not a standard you can build on today.

What Anthropic actually announced

The announcement is narrower than the headlines suggest, and I mean that as praise. Anthropic is not claiming Claude can run your factory. It is opening MHS to a first cohort of scientific research labs and advanced manufacturers, by application, through a dedicated preview site, with the stated intent to open source the standard after the preview.

The origin story is worth a paragraph because it explains the design. MHS began as a collaboration between Alek Kemeny on Anthropic's Beneficial Deployments team and Arco Bast, a postdoctoral scientist at HHMI's Janelia Research Campus. Bast had built a shared memory dictionary to synchronise lasers, motorised focusers and cameras from different vendors for brain-imaging experiments, the kind of plumbing every wet lab reinvents badly. Kemeny and Bast then wired AI models into that interface. So MHS isn't a standards committee's thought experiment. It's a lab hack that grew up.

The claimed prize, in Anthropic's own words: it "typically takes a lab or manufacturing facility weeks, if not months, to set up and integrate their hardware," and MHS "reduces this integration work to hours or minutes." If you've ever watched a postdoc fight a vendor's Windows-only control software for a fortnight, you'll know the weeks-to-months figure is not marketing inflation.

Anthropic opened an application-only research preview of the Model Hardware Standard on 27 August 2026, targeting scientific research labs and advanced manufacturers, per its announcement. The standard grew out of a collaboration between Anthropic's Alek Kemeny and HHMI Janelia postdoc Arco Bast, and Anthropic claims it cuts hardware integration from weeks or months to hours or minutes.

How MHS works, and where MCP fits

The mechanics are refreshingly boring, which is what you want from plumbing. Each device gets a standardised MHS driver built on a small set of primitives, commands like "read" (get temperature) or "write" (set temperature), that make the device discoverable in a standard format. The driver also carries natural-language tags, things like the weight of a robot arm, which generate a per-device reference file. That file tells the agent what the device can measure, what can be adjusted, and which safety limits will be enforced. The agent reads the manual before it touches the machine, and the manual is written in a format every agent reads the same way.

The question everyone in my world asked within minutes was: what does this do to MCP? The answer, from the announcement, is that nothing happens to it. Anthropic lists three ways an agent can control a device through MHS: MCP, the command line, and code APIs. MHS is model-agnostic, and any agent harness can reach it through standard protocols, the Model Context Protocol among them. The cleanest summary I've seen came from one early analysis: MCP can carry the tool call, while MHS makes the machine behind that tool call understandable and consistently operable. Kemeny's own framing, reported by nextaipress, is blunter: "What MCP did for software, MHS will do for the hardware world."

Layer diagram showing how an agent harness connects through MCP to an MHS driver, which wraps a physical device. The MHS driver generates a reference file describing capabilities and enforced safety limits, and the device declares bounds, interlocks and an emergency stop that the agent operates within by default.

The shape of the thing, from Anthropic's announcement and the Strands engineering write-up: the transport layer is unchanged, and the safety contract lives in the driver.

If you want the software-side background, I wrote a plain-English walkthrough of what MCP actually does for a business earlier this year, and the mental model transfers directly. MCP standardised how agents ask for things. MHS standardises what a physical thing looks like when it's asked.

MHS gives each device a standardised driver with read and write primitives plus a generated reference file describing capabilities, adjustments and enforced safety limits, per Anthropic's announcement. It does not replace MCP: agents reach devices through MCP, the command line or code APIs, and MHS is model-agnostic across agent harnesses.

The safety lives in the driver, not the model

This is the section that matters, and it's the one most coverage will skip for the flashier partner numbers. An agent that misreads a spreadsheet costs you a bad report. An agent that misreads a centrifuge costs you something else entirely. So the design question is not "is the model smart enough," it's "where do the hard limits live when the model isn't."

Anthropic's answer is architectural, and the clearest description of it came not from Anthropic but from AWS, whose Strands team is building MHS as a transport backend in its open-source Strands Robots library: "Safety is built into the standard itself. Devices declare their bounds, interlocks, and emergency stops, and agents inherit and operate within them by default." Read that twice. The interlock isn't a prompt instruction the model might forget or a policy it might talk its way around. It's a property of the driver, below the model, enforced whether the model cooperates or not. That's the correct instinct, and it matches how I think about agent governance in software: the guardrail that survives contact with a confused agent is the one the agent can't edit.

Anthropic is also unusually candid about the model's limits. The announcement concedes that Claude "learns about the physical world through text and images, meaning its spatial and physical reasoning have limitations that still require expert oversight." The preview exists, in part, "to build additional safety evaluations," and Anthropic says it is developing a physical safety roadmap to bolster its safeguards policy against misuse, with findings released as deployment guidance when MHS is open sourced.

One partner anecdote shows why that humility is earned. Genentech automated the BCA protein assay, a workhorse chemistry measurement, across a liquid handler, a robotic arm and a plate reader, and researchers had to teach Claude that foaming samples were physical failures, not software bugs. A model that has never had hands doesn't know what bubbles mean. That gap between text knowledge and physical consequence is exactly what the driver-level limits are for, and exactly why I'd keep a human in the loop for anything that spins, heats or pressurises.

MHS puts safety in the driver rather than the prompt: devices declare bounds, interlocks and emergency stops that agents inherit by default, per the AWS Strands write-up. Anthropic concedes Claude's physical reasoning still requires expert oversight and says the preview will inform a physical safety roadmap and deployment guidance, per its announcement.

The pilot results, read carefully

The partner list is genuinely impressive. Genentech ran the BCA assay end to end. Carnegie Mellon ran serial-dilution dose-response experiments roughly three times faster across three computers with, in Anthropic's phrase, "fundamentally incompatible interfaces." QuEra Computing had an agent develop a controller that recovers a quantum laser's frequency lock 99.3% of the time without human intervention. The University of Washington ran agent-supervised qPCR with collision-free plate handoffs between a robot arm and a liquid handler. And at Janelia, Virginie Ruetten in the Ahrens lab used MHS to unify a microscopy rig that had previously needed seven separate vendor programs to run. Seven. Every experimental scientist reading this just winced in recognition.

Bar chart showing QuEra Computing's quantum laser frequency lock recovery: 58 percent of the time before the agent-developed controller and 99.3 percent of the time with it, without human intervention.

QuEra's agent-developed controller recovered the laser's frequency lock 99.3% of the time, per Anthropic's announcement; the 58% baseline is The Register's reporting. Partner-reported, so file under promising, not proven.

Now the discipline. All of these numbers are partner-reported, published in the vendor's own announcement. CMU's dose-response run used colour dye, not a drug candidate. QuEra's controller ran on a testbed, not a live quantum processor. And every detailed example in the launch materials runs on Claude, which makes the "model-agnostic" claim exactly that so far: a claim. None of this makes the results fake; pilot results from serious institutions rarely are. It means the evidentiary standard is "the people who built it say it works," and the honest reader holds that at arm's length until the spec is public and someone outside the friend group reproduces it.

Two things in the ecosystem list did catch my eye. Hugging Face is adding MHS support to LeRobot, its open robotics library, and Raspberry Pi is integrating after its camera driver tests. Those two matter more for adoption than any pharma logo, because they put the standard in front of the hobbyist and research communities that turned ROS into the default robotics stack over a decade. The Open Robotics community is already arguing about where MHS sits relative to ROS 2, and that argument, not the launch post, is where the standard's real shape will get decided.

Pilot results are real but partner-reported: Carnegie Mellon's dose-response experiments ran roughly three times faster, QuEra's agent-developed controller recovered its laser's frequency lock 99.3% of the time, and HHMI Janelia unified a microscopy rig that had required seven vendor programs, all per Anthropic's announcement. Every detailed launch example runs on Claude, so the model-agnostic claim is undemonstrated publicly.

The names that aren't on the list

Here's my one genuinely parochial observation, and I'll flag it as such. Scan the hardware partners Anthropic names: Tecan for liquid handlers, Universal Robots and Doosan for arms, QIAGEN and Automata for lab automation, Danaher exploring, MBF Bioscience for microscopes. A serious list. And not one Chinese robotics company on it.

I live fifteen minutes from the centre of gravity of the world's humanoid and quadruped industry. Unitree and Deep Robotics are both a short ride from my apartment, and the manufacturers who would get the most from a hardware-agent standard, the ones running the factories those robots walk into, are overwhelmingly here. The robots themselves are not the bottleneck; the integration is, which is precisely MHS's pitch. So the absence reads less like a snub and more like a preview cohort drawn from Anthropic's existing compliance-friendly relationships, and it leaves the biggest hardware ecosystem on earth as an open question. Whether Chinese vendors adopt MHS, ignore it, or back something homegrown is, for my money, the most consequential unknown in the whole announcement, and nobody covering it in English seems to have asked.

Anthropic's named hardware partners (Tecan, Universal Robots, Doosan, QIAGEN, Automata, MBF Bioscience, with Danaher exploring) are all from Western or Korean lab and industrial automation, per its announcement. No Chinese robotics manufacturer is in the preview cohort, leaving adoption in the world's largest hardware ecosystem an open question.

What the preview doesn't answer

A research preview is a promise with training wheels, and the honest thing to do is list what hasn't been promised. The Kingy analysis did the useful counting: Anthropic has published no specification, SDK, schema, source repository, licence, conformance suite, version number or governance model. modelhardwarestandard.com confirms the intent to open source, with no date. MCP eventually landed at the Linux Foundation's Agentic AI Foundation; whether MHS gets the same neutral home, or stays Anthropic's to steer, is unannounced, and it matters, because a hardware standard one vendor controls is a product feature, not a standard.

The security model is the bigger hole. Nothing public describes authentication, credential scoping, network isolation, audit logging, driver revocation, or what stops a prompt-injection attack arriving through data a connected device feeds back to the agent. In software that last one is an open research problem; Brave's Comet research showed how little stands between a hostile page and a helpful agent. In hardware it's a safety case. I'd want to see the threat model before I let an MHS driver anywhere near a machine that can hurt someone, and right now there isn't one to read.

There's also a governance critique worth taking seriously. One reaction to the closed preview on Hacker News was blunt: "You shouldn't need permission to read a standard." Standards earn trust by being public before they're blessed, and Anthropic is asking hardware makers to build drivers against something only accepted applicants can see. The Register, never shy, called the whole thing "a somewhat optimistic ambition" from a company that cannot reliably anticipate how its own models behave, and raised the misuse scenario nobody wants printed next to their logo: the same standard that runs a drug-discovery assay could, in principle, run a centrifuge cascade. Anthropic's physical safety roadmap is presumably the answer to that. We haven't seen it yet.

One piece of timing sharpens all of this. The EU Machinery Regulation applies in full from 20 January 2027 and, as one report noted, for the first time explicitly covers AI-based safety functions and machines with self-evolving behaviour. Any standard for AI-actuated machinery sold into Europe will meet that regulation in about sixteen months, whether it's ready or not.

MHS has no public specification, SDK, licence, governance model or security model, per the Kingy analysis and the preview site. Anthropic says it will open source the standard and release deployment guidance from the preview, without dates, per its announcement.

What to do with this

Depends who you are, so briefly, three audiences.

If you run a lab or a pilot line, the move is cheap: identify the one instrument that eats the most scientist-hours in manual babysitting, and check whether its vendor is on the partner list or has a programmable interface at all. MHS only works with hardware that has one, and Anthropic is still courting the manufacturers whose kit doesn't. If you're eligible and curious, the preview application exists. If you build agents for a living, watch the Strands Robots integration and the LeRobot support, because that's where you'll first touch this without anyone's permission. And if you invest or set strategy, the signal isn't "Anthropic does robotics now." It's that the value in physical AI is sliding down the stack from the model to the integration and safety layer, and the company that owns the boring layer tends to own the ecosystem. MCP went from launch to industry default in about a year; the same playbook is visibly being run again, this time with hardware.

For labs, the actionable step is auditing which instruments have programmable interfaces, since MHS requires one; for builders, the AWS Strands Robots and Hugging Face LeRobot integrations are the first public touchpoints, per the Strands write-up and Anthropic's announcement.

Frequently asked questions

Is MHS a replacement for MCP?

No. MHS sits below and beside MCP. The agent's tool call still travels over MCP (or the command line, or a code API), exactly as it does today. MHS standardises the other end: how the physical device describes itself, what it can do, and which limits are enforced.

Can I use the Model Hardware Standard today?

Only if you're accepted into the research preview, which is application-only and aimed at scientific research labs and advanced manufacturers via modelhardwarestandard.com. There's no public spec or SDK yet. Anthropic says it intends to open source MHS after the preview, with no date given.

Does MHS only work with Claude?

Anthropic says MHS is model-agnostic and reachable by any agent harness through standard protocols. But every detailed example in the launch materials uses Claude, so treat cross-model operation as designed-for rather than demonstrated until independent builds appear.

Is it safe to let an AI agent control physical equipment?

The honest answer is that it depends on where the limits live. MHS puts bounds, interlocks and emergency stops in the device driver, enforced below the model, which is the right architecture. But Anthropic itself says Claude's physical reasoning still requires expert oversight, there's no public threat model yet, and The Register's sceptical read on misuse risk is fair. Supervised pilots on declared, interlocked equipment: reasonable. Unsupervised agents on anything that spins, heats or pressurises: not yet.

The bottom line

Every few months the industry rediscovers that intelligence was never the scarce resource. Kitesurf did it for browsers last month; MHS is doing it for hardware now. The models are good enough to run a protein assay or hold a laser lock. What was missing was a way for them to find the machine, read its manual, and know exactly where the walls are, and that's a plumbing problem, which is why a plumbing spec is the right shape for the answer.

What's proven today: the driver pattern works, the integration-time collapse is plausible and partner-witnessed, and putting the safety in the device rather than the prompt is the correct call. What's not: everything you need to trust it in production, starting with a spec you can read without permission, a threat model, a governance home, and one result reproduced by someone who wasn't invited to the launch. I'd bet on the direction. I wouldn't bet a centrifuge on the preview.

If you're weighing up where hardware fits in your own agent plans, that's a conversation I have with clients regularly. Get in touch.

Sources

  • Anthropic, "Previewing the Model Hardware Standard": https://www.anthropic.com/news/model-hardware-standard-research-preview (published 2026-08-27, retrieved 2026-08-29)

  • Model Hardware Standard preview site: https://modelhardwarestandard.com (retrieved 2026-08-29)

  • AWS Strands, "Robots working together: the Model Hardware Standard and Strands Robots": https://strandsagents.com/blog/robots-working-together-model-hardware-standard-strands-robots/ (retrieved 2026-08-29)

  • The Register, "Anthropic proposes plumbing spec to link AI agents to lab kit and robots": https://www.theregister.com/ai-and-ml/2026/08/28/anthropic-proposes-plumbing-spec-to-link-ai-agents-to-lab-kit-and-robots/5293135 (published 2026-08-28, retrieved 2026-08-29)

  • Kingy.ai, "Anthropic's Model Hardware Standard (MHS)": https://kingy.ai/blog/anthropic-model-hardware-standard-mhs/ (published 2026-08-28, retrieved 2026-08-29)

  • nextaipress, "Anthropic Model Hardware Standard (MHS)": https://nextaipress.com/anthropic-model-hardware-standard-mhs/ (published 2026-08-28, retrieved 2026-08-29)

  • Open Robotics Discourse, "Where does an intent validation layer sit relative to Anthropic's Model Hardware Standard and to ROS 2?": https://discourse.openrobotics.org/t/where-does-an-intent-validation-layer-sit-relative-to-anthropics-model-hardware-standard-and-to-ros-2/57725 (published 2026-08-28, retrieved 2026-08-29)

  • Brave, "Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet": https://brave.com/blog/comet-prompt-injection/ (published 2025-08-20, retrieved 2026-08-29)

  • Hacker News discussion, "Previewing the Model Hardware Standard": https://news.ycombinator.com/item?id=49468834 (retrieved 2026-08-29)

  • Flickr, cover image "Robot Arm" by fireflythegreat (CC BY 2.0): https://www.flickr.com/photos/fireflythegreat/5604962354/ (retrieved 2026-08-29)

Keep reading

Agent Field Notes

Get the next issue.

Agent harnesses, runtimes, security and governance, explained for the people who have to operate them.

Facing a decision like this?

We run architecture reviews, governance assessments and version-pinned framework evaluations for teams making consequential agent decisions.

About the author

Adam Maguire Wilson

Founder and independent advisor on AI agent systems.

adam.mw