Cloudflare's vulnerability agents get your code and your traffic, but not your deploy button
Cloudflare's invitation-only Vulnerability Discovery and Remediation service points OpenAI's Daybreak models at customer-authorised codebases, prioritises findings with live WAF and traffic context, and keeps implementation authority with the customer. A look at where each control actually sits.
On this page
What Cloudflare announced
On 3 September, Cloudflare opened early access to Vulnerability Discovery and Remediation, an invitation-only service inside its Managed Defense offering that hunts vulnerabilities in customer codebases. Through the OpenAI Daybreak Defense Network, the service runs OpenAI's Daybreak models, including the purpose-trained GPT-5.6-Cyber, for reconnaissance, hunting and validation against code the customer has authorised Cloudflare to inspect. Findings come back as proposed code patches and mitigations, each automatically checked before a human ever sees it. The customer decides what gets implemented.
The discovery machinery is not new. Cloudflare described the internal version in June in Build your own vulnerability harness: a model-agnostic pipeline that scans its own fleet, forces every finding through adversarial validation, and turns raw model output into patches its engineers can trust. That post is worth reading on its own terms, and we will not retell it here. What the September announcement adds is two things the internal system never needed: someone else's code, and a boundary around who is allowed to touch it.
Context is the actual product
Every security vendor now has a model that can read code and find bugs. What none of them can do from a repository alone is answer the question the Cloudflare post opens with: your scanner flagged thousands of vulnerabilities, dozens of them critical, so which one do you fix first? A static finding cannot tell you whether the vulnerable handler is deployed, whether anyone is hitting that route, or what protections already stand in front of it.
Cloudflare's answer is the telemetry it already holds. An investigation starts with a snapshot from Web Assets and the Web Application Firewall: which routes are active, how much traffic they carry, what security events surround them, and which WAF rules are already blocking attacks. Routes carrying heavy request volume get treated as hot paths and profiled more strictly. For Workers applications, the service retrieves the latest source version and its configured routes, then ties them to request metadata from Workers Observability, so the exact code under review is matched to the endpoints it serves in production.
The discipline in the design is easy to miss, and it is the part we find genuinely encouraging. Network context guides where the hunter agents look, but it cannot create a finding. Every vulnerability has to be corroborated by evidence in the source code itself. Production data can raise a finding's risk rating, for example when the affected endpoint carries significant traffic or shows signs of active probing, but a quiet route with a real bug stays a real bug, and a loud route with no bug stays clean. Evidence about exposure decides the order of the queue. It is not allowed to manufacture entries in it.
That is what makes agent findings actionable rather than merely numerous. Not a cleverer model, but a prioritisation layer grounded in how the application is actually being used and attacked.
The authority ladder
The most carefully drawn part of the announcement is who may do what, and it is worth laying out as the ladder it is.
The model sits at the bottom with no authority at all. The harness runs on Cloudflare's infrastructure, prompts travel from Workers through Cloudflare AI Gateway to OpenAI's servers, and no inference runs at the edge. The model cannot apply any patch or rule it proposes. Before its context reaches OpenAI, the service strips what the investigation does not need and applies the engagement's redaction controls, and the harness is explicitly built to treat source code, logs and request metadata as evidence to inspect rather than instructions to follow. Tool calls are logged and checked against the investigation's access policy before they run.
Above the model sits deterministic code. Every patch or rule proposal must pass checks implemented outside the model, and if a check fails, the workflow stops before anything reaches customer review. Above that sits Cloudflare's own team, which validates the output before presentation: proposed WAF rules are checked for syntax and run against synthetic fixtures representing expected requests, never against live customer traffic, and ambiguous results are held back for diagnosis rather than shipped with a shrug. Only then does the customer see anything, and the customer decides whether to test or deploy.
There is one rung that breaks the propose-only pattern, and it is the one to interrogate. If the customer has authorised the service to defend the zone, Cloudflare will deploy the proposed WAF custom rule itself, scoped conservatively around the method, path and request details needed to reach the vulnerable code. Where a route pattern contains only variables and wildcards, it refuses to suggest a rule at all, on the stated grounds that it would rather miss a possible connection than claim one the evidence cannot support. We covered Harness's security agents in August, which stop at the pull request and let a developer merge or not. Cloudflare's instinct is the same, one rung further up: not a code merge, but a reversible mitigation at an edge the vendor already operates, gated behind a separate and explicit opt-in. That is a defensible place to put the line. It is also a line customers should read twice before signing, because "defend your zone" is doing a great deal of work in that sentence.
A model trained to stop saying no
The choice of model makes the surrounding controls more interesting, not less. GPT-5.6-Cyber is the top tier of OpenAI's expanded Daybreak programme, announced on 10 August: a variant built on GPT-5.6 Sol and trained to reduce refusals on higher-risk, dual-use security tasks. On OpenAI's internal completion evaluation, which measures how often a model responds to requests involving exploit chains, authentication bypass and privilege escalation, GPT-5.6-Cyber answers 95 percent of the time. The general-purpose Sol answers 1.5 percent, and even Sol with Daybreak Blue's defensive-work safeguards answers just 2 percent.
Read those figures from the defender's side and they are the point of the exercise: a security model that refuses nineteen requests in twenty is not much use for hunting real bugs. Read them from the control-design side and they say something else. Cloudflare has deliberately put the most willing version of a frontier model it can get inside this pipeline. The refusal layer, which is normally the last backstop against a model doing something unwise, has been intentionally weakened in exchange for capability. Everything the model is no longer willing to refuse has to be refused by the harness instead: the scoping, the redaction, the policy-checked tool calls, the external validation, the human review. Cloudflare's architecture reads as if it understood exactly that trade, and engineered for it.
What the announcement cannot tell you
Some candour about the evidence, because the launch deserves it.
Everything above traces to Cloudflare's own description of its own service. There are no named customers, no pricing, and no published detection or fix-quality numbers for the customer-facing offering. The June post did publish real internal figures, including a validation rejection rate that fell from 40 to 11 percent as context injection improved, but those numbers describe Cloudflare scanning Cloudflare. A paying customer's sprawling monorepo is a different evidence base, and the honest answer about fix quality will come from early engagements, not from the announcement.
Availability is narrow by design. Early access is invitation-only through the Managed Defense team, each engagement starts with a single authorised application codebase, and connecting findings to production requires granting read access to the Web Assets inventory, the relevant WAF controls and Workers Trace Events Logpush where available. The investigation is described as semi-automated, with Cloudflare's team validating output before review, which raises a plain throughput question nobody can answer yet: how many applications, for how many customers, at what pace.
The last limitation is structural. The prioritisation advantage exists only where Cloudflare can see the workload: Workers applications and proxied traffic behind its WAF. The further your estate sits from that edge, the more this resembles a well-instrumented scanner and the less it resembles the demo. The telemetry is both the moat and the lock-in, and the code you authorise also leaves your boundary for OpenAI's servers, redacted but real. Any security team considering early access should ask its account team exactly what the redaction removes, what is retained, and precisely which actions the "defend your zone" authorisation permits before anything is signed.
The scarce resource in vulnerability management was never the ability to find bugs. Models made finding cheap a year ago. What stays scarce is knowing which three findings matter tonight, and having a trusted path from that knowledge to a fix in production. Cloudflare is betting that the path runs through the edge it already operates, with a model that proposes, machinery that checks, and a customer who disposes. The design is right. Whether customers will hand one vendor both their source code and their traffic data to run it is what early access will actually decide.
Keep reading
Agent Field Notes
Get the next issue.
Agent harnesses, runtimes, security and governance, explained for the people who have to operate them.
Facing a decision like this?
We run architecture reviews, governance assessments and version-pinned framework evaluations for teams making consequential agent decisions.