İçeriğe geç
Analizler
Inside

SafeMind: An AI Security Agent That Has to Prove It Works

CrowdStrike and NVIDIA's SafeMind evaluation grades AI security agents on sensor telemetry, not self-assessment. What the funnel actually shows.

Yazan Adam Maguire Wilson16 dk okuma
Two fencers mid-bout, wired to a scoring apparatus that registers each touch, an apt metaphor for a security evaluation where the sensors, not the agent, keep score.
Bu sayfada

How do you know an AI security agent actually works? Not from the press release, and certainly not by asking the agent. You know it the way you know a fencer scored: the apparatus registers the touch, and nobody much cares what the fencer thinks happened.

On 1 September, CrowdStrike and NVIDIA announced SafeMind, an agentic cybersecurity system, on the Fal.Con stage in Las Vegas. The press release came with the usual artillery of vendor numbers: 29% higher detection rate, 6 times faster remediation, 99% lower cost. More interesting, and published the same day, is NVIDIA's engineering write-up of the evaluation, a document that reports its own failure rates, lists its own limitations, and grades the agent on sensor telemetry rather than self-assessment. That write-up is what this piece is about, because the evaluation design matters more than the announcement.

Key Takeaways - SafeMind pairs an offensive model (Red Tempest) with a defensive model (Blue Solano, post-trained from NVIDIA's open Nemotron 3) in a closed attack-defence loop, tested inside a digital twin of NVIDIA's own accelerated computing infrastructure. - The evaluation measures progress from attack traces and Falcon sensor telemetry, not the agent's own claims. The optimised open pipeline lifted mean backtest detection from 16.5% to 41.9%, and in live-fire testing produced the only three detections to pass every quality gate. The frontier-model comparison produced none. - The headline claims (29% higher detection, 6 times faster remediation, 99% lower cost) are CrowdStrike internal evaluations. The write-up itself calls the results a directional case study, not a general benchmark. - The transferable lesson is the evaluation design: deterministic linting, replay against captured telemetry, independent review with a fresh context, and no credit for self-reported success.

What happened

At Fal.Con 2026, in front of roughly 10,000 security professionals, CrowdStrike CEO George Kurtz and NVIDIA CEO Jensen Huang announced SafeMind, built by the new CrowdStrike Cyber Superintelligence Lab and covered in NVIDIA's conference write-up. The shape of it, per CrowdStrike's press release:

  • Two purpose-built models. Red Tempest plays offence, emulating AI-driven adversaries inside a copy of the target environment. Blue Solano plays defence, writing and deploying detections.

  • Both are built on NVIDIA's open Nemotron 3 models, post-trained on CrowdStrike's own material: Falcon sensor telemetry, threat intelligence, Falcon Complete incident annotations, and fifteen years of incident response fieldwork. CoreWeave's cloud handled training and inference.

  • The harnesses that run the two models in a continuous loop also accept other frontier and open models, and standalone access to the models and harnesses is planned through CrowdStrike's Project QuiltWorks programme.

Kurtz's framing on stage was blunt: "The real gap that I saw was that the attackers had frontier AI, and the defenders didn't." The demand-side numbers CrowdStrike cites are uncomfortable enough to make the pitch land: AI-enabled attacks up 89% over the past year, and a fastest observed breakout time (the gap between an intruder's first foothold and their move sideways) of 27 seconds.

Nemotron 3 is the model family underneath. NVIDIA launched it in December 2025 in Nano, Super and Ultra sizes, aimed squarely at agentic workloads, with the largest Ultra size following in 2026. In the evaluated configuration, Nemotron 3 Ultra orchestrates the defensive workflow while a fine-tuned Nemotron 3 Super, an open hybrid Mamba-Transformer mixture-of-experts model, does the bounded expert work of writing and repairing detections.

SafeMind is CrowdStrike's agentic cybersecurity system, announced 1 September 2026 at Fal.Con: an offensive model (Red Tempest) and a defensive model (Blue Solano), both post-trained from NVIDIA's open Nemotron 3 family on CrowdStrike's own telemetry and incident data, run in a continuous attack-defence loop by purpose-built harnesses, per CrowdStrike's announcement and NVIDIA's conference coverage.

The loop: attack, detect, retest, repeat

The engineering write-up describes a four-stage cycle that replaces the traditional red team, blue team, detection engineer relay with software.

First, execute and capture. A red-agent harness, running Recon, Assault and Compromise sub-agents, picks an attack path towards a threat-informed objective and walks it inside an isolated environment, while CrowdStrike Falcon endpoint sensors record the telemetry. Second, process and reconstruct: the blue-agent harness receives the action trace and sensor data, works out what it can reconstruct of the event sequence, which existing detections fired, and where the gaps are. Third, generate and validate: the blue harness writes candidate detections, and a validation harness backtests each one against the captured telemetry, bouncing failures back for correction. Fourth, retest: once a validated detection is deployed, a fresh, independently seeded attack goes after the same objective, and the red harness adapts around the new defence. The loop runs until the red agent finds no further viable path.

The environment deserves a paragraph of its own, because it's the part most evaluations fake. NVIDIA wrote a sanitised natural-language specification of its own accelerated computing infrastructure, and an agent-assisted workflow turned that specification into an isolated target environment instrumented with Falcon sensors. NVIDIA's security experts reviewed both the environment and the threat paths for realism. The same reviewed environment then hosted every attack run, detection test and metric, which is what makes the configurations comparable.

And here is the sentence that justifies the whole exercise: progress was measured from action traces and sensor telemetry, in the write-up's own words, "rather than the agent's own claims." Anyone who has watched an agent declare victory over a broken test suite will understand why that sentence matters.

The SafeMind evaluation loop runs a red-agent harness executing attack paths inside an isolated digital twin of NVIDIA's infrastructure, instrumented with Falcon sensors, while a blue-agent harness reconstructs events from telemetry, generates candidate detections, and has them validated by replay and independent review before an independently seeded attack retests the same objective. Progress is measured from action traces and sensor telemetry rather than the agent's self-report, per NVIDIA's evaluation write-up.

Making a detection agent earn its detections

A detection rule that merely looks plausible is worse than useless, because someone has to triage it. The write-up's answer is to wrap the blue agent in six checks that encode what a good detection engineer would do by hand:

  • A schema knowledge base, so the agent can only reference sensor fields and query syntax that actually exist. No invented fields.

  • Telemetry grounding, so every detection is anchored in observed events rather than hallucinated connections.

  • A bounded expert model for authoring. The customised Nemotron 3 Super, which CrowdStrike calls NL2LogScale, writes and repairs detections in a separate, narrow context, away from the long orchestration thread.

  • Artifact linting. Automated checks reject syntax errors, unsupported fields, and detections pinned to specific IP addresses, hosts or subnets, with feedback that pushes the rewrite towards behavioural signals.

  • Detection replay. Every candidate is replayed against the captured attack telemetry. No match, no promotion.

  • Independent review. A separate judge model with a fresh context scores each survivor for behavioural grounding, robustness and use of multiple signals.

The training behind the authoring model is worth a look too, because it uses the same philosophy. CrowdStrike post-trained Nemotron 3 Super with continual pretraining on cybersecurity knowledge, supervised fine-tuning on 9,349 detection-generation and multi-step repair examples spanning 59 programmatically generated error types, then reinforcement learning with verifiable rewards. In the RL stage, built on NeMo Gym and NeMo RL, each generated query is actually executed against Falcon LogScale. Invalid queries get the engine's real error messages and up to five repair attempts, with unresolved attempts scoring zero. Valid queries run against the same synthetic logs as a reference query, and the reward is the F1 overlap between the two result sets.

Notice what the reward is not. It is not "does the query compile". It is not "does the output look like a detection". It is "does your query return the same events as a known-good query against the same data". Nothing in this pipeline rewards sounding right, and that is a design choice most agent demos skip.

CrowdStrike's NL2LogScale detection-authoring model was post-trained from Nemotron 3 Super on 9,349 detection-generation and repair examples covering 59 error types, then tuned with reinforcement learning where the reward is the F1 overlap between the events a generated query returns and those of a reference query executed against the same logs. Generated detections must additionally pass schema linting, telemetry replay and independent review before deployment, per NVIDIA's evaluation write-up.

The funnel is the story

Now the numbers, which are the honest kind: mostly failure.

Backtesting first. With Nemotron 3 Ultra in a default agent harness, a mean of 16.5% of generated detections actually detected the recorded attack, averaged across eight independently seeded authoring sessions. Keeping Ultra but adding the tuned harness, the customised Super, domain context, tools and validation raised the mean to 41.9% across six sessions, a 2.5 times improvement. Credit to the authors for stating the obvious confound themselves: because the harness and the model stack changed together, the gain reflects the complete pipeline, not an isolated model ablation.

Bar chart of mean backtest detection rates. Nemotron 3 Ultra with the default harness detected 16.5 percent of the recorded attack across eight seeded sessions. The optimised open pipeline detected 41.9 percent across six sessions, a 2.5 times improvement.

Mean backtest detection rates from NVIDIA's evaluation write-up. Even the winning configuration failed to catch the recorded attack more than half the time.

Then live fire, which is where it gets genuinely interesting. Every backtest-passing detection was deployed verbatim to a live detection engine and exposed to eight new, independently seeded attacks from the same scenario family. An independent third-party model judged whether each match represented the intended attack behaviour. The open pipeline contributed 11 detections, the complete frontier system (a leading proprietary frontier model in the same style of harness) contributed 35. Five of the 11 open detections, 45%, detected at least one unseen attack, against 10 of the 35 frontier detections, 29%. The firing open detections caught 2.6 of the eight attacks apiece on average, against 1.1 for the frontier system, and between them covered all eight attacks to the frontier system's seven.

Firing is only the first gate, though. To qualify as gold, a detection also had to stay quiet on the available benign test traffic and survive the independent review. Four of the five firing open detections and nine of the ten frontier ones stayed quiet. After review, three open detections and zero frontier detections qualified as gold. The three open ones still covered all eight attacks between them.

Funnel chart of live-fire validation for the open Nemotron pipeline versus the frontier system. Open pipeline: 11 detections deployed, 5 detected at least one unseen attack, 4 stayed quiet on test traffic, 3 passed independent review as gold. Frontier system: 35 deployed, 10 fired, 9 stayed quiet, 0 gold.

The live-fire validation funnel, redrawn from NVIDIA's evaluation write-up. The frontier system generated three times as many deployable detections and none of them survived review.

Then the write-up does something vendor documents almost never do. It lists, unprompted, why you shouldn't over-read its own results: one scenario family and small detection sets, so cross-scenario generalisation is untested. Limited benign traffic, so the quiet-on-test-traffic check doesn't represent production false-positive rates. Three of the eight live-fire runs suffered harness failures (they produced complete telemetry and were retained). And the summary verdict on itself: a directional system-level case study, not a general benchmark. I can't remember the last time a joint vendor announcement included its own confidence interval in prose.

In live-fire testing against eight independently seeded attacks, five of 11 open-pipeline detections (45%) detected at least one attack versus 10 of 35 (29%) for the frontier system. After benign-traffic and independent-review gates, three open detections and zero frontier detections qualified as gold. NVIDIA's write-up itself cautions that the evaluation covered one scenario family with small detection sets and limited benign traffic, calling it a directional case study rather than a general benchmark.

What to believe, and what to discount

It helps to sort the claims into three tiers, because they are not equally strong.

Tier one: the marketing numbers. The 29% higher detection rate, 6 times faster remediation and 99% lower cost in the press release are CrowdStrike internal evaluations against "leading frontier models and open-source baselines", with no published methodology, no scenario list and no error bars. The same goes for Blue Solano being more accurate than "the leading proprietary frontier model tested" at 99% lower cost. Treat these as directional claims from an interested party. They might be true. You cannot check them.

Tier two: this evaluation. The write-up is vendor-run, but it is telemetry-grounded, seeded independently, judged by a third-party model, and published with its failure rates and limitations attached. That is a materially better class of evidence than a benchmark screenshot. It is still not independent: NVIDIA and CrowdStrike chose the scenario family, built the environment and ran the sessions.

Tier three: independent public evaluation. This is what actually settles arguments, and we know what it looks like because DARPA ran one. The AI Cyber Challenge final in August 2025 put seven autonomous cyber reasoning systems against 54 million lines of real C and Java in public, with a shared scoring apparatus: they found 54 of 63 planted vulnerabilities (86%, up from 37% at the semifinal a year earlier), patched 43 of them, and turned up 18 genuine bugs nobody had planted, at an average cost of roughly $152 per task and an average patch time of 45 minutes, per DARPA's official results announcement. Nothing equivalent exists yet for detection-engineering agents, and nobody outside the two vendors has demonstrated SafeMind-class effectiveness in production. Until something like it does, the honest status of SafeMind is "promising, vendor-evaluated, unverified at production scale".

One structural point is worth separating from the scorekeeping. The reason CrowdStrike could post-train on fifteen years of its own incident data without shipping that data to a model provider is that Nemotron's weights are open. NVIDIA is pushing that argument hard through the new Open Secure AI Alliance, and its sharpest supporting anecdote is the recent Hugging Face incident: when closed AI tools proved unable to distinguish defenders from attackers during forensic analysis, Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure to work through more than 17,000 actions and contain the intrusion. In NVIDIA's telling, anyway. But the underlying point stands on its own: in security, the ability to inspect and self-host the thing defending you is not a nice-to-have. I see the same instinct behind the enterprise pull of Chinese open models here in Hangzhou.

The SafeMind headline figures (29% higher detection, 6x faster remediation, 99% lower cost) are unaudited internal evaluations. The telemetry-grounded write-up is stronger evidence but still vendor-run. The reference point for independent public evaluation of autonomous cyber systems remains DARPA's AI Cyber Challenge, whose August 2025 final scored seven systems against 54 million lines of code in public, per DARPA's official results. No equivalent public evaluation exists yet for detection-engineering agents.

What I'd ask before buying the claim

If a vendor shows you agentic security results this year, from CrowdStrike or anyone else, these are the questions that cut through:

  1. Show me the funnel, not the headline. How many candidate detections survived replay, how many fired on attacks they hadn't seen, how many stayed quiet on normal traffic, how many passed independent review. A system that ships 35 plausible detections and zero good ones is a triage tax, not a defence.

  2. What fraction of the evidence is self-reported? If the metric comes from the model's own account of what it did, discount it to zero. Telemetry or it didn't happen.

  3. What does benign traffic look like at my scale? Quiet on a limited test set is not a false-positive rate. The write-up admits this about itself, which is more than most will.

  4. Who ran the evaluation, and can a third party reproduce it? Vendor-run is fine as a first pass. It is not fine as the last word.

  5. If I'm evaluating my own agents: seed the attacks independently, grade against sensor ground truth, put a judge with a fresh context between the agent and the verdict, and keep the failure rates. You'll learn more from your 16.5% run than from any number with a press release attached.

This lines up with how I think about agentic architecture generally: the model is the easy part, and the harness, the validation and the evidence trail are where systems succeed or fail. Governance questions, like who owns a detection an agent wrote and promoted, sit in the same frame as agent governance more broadly.

FAQ

What is CrowdStrike SafeMind?

SafeMind is an agentic cybersecurity system from the CrowdStrike Cyber Superintelligence Lab, announced 1 September 2026 and built to run natively in the Falcon platform. It pairs an offensive model that finds attack paths with a defensive model that writes and deploys detections, run in a continuous loop by purpose-built harnesses inside a digital twin of the target environment. Both models are post-trained from NVIDIA's open Nemotron 3 family on CrowdStrike's own security data, with CoreWeave providing training and inference capacity.

What are Red Tempest and Blue Solano?

They are SafeMind's two launch models. Red Tempest is the offensive red-team model, built to emulate AI-driven adversaries in advanced attack scenarios. Blue Solano is the defensive blue-team model, built to deploy the measures defenders actually use. CrowdStrike's internal evaluations put Blue Solano ahead of the leading proprietary frontier model it tested, at 99% lower cost, though that claim has not been independently audited.

Were the SafeMind evaluation results independently verified?

Not in the strong sense. The evaluation was designed and run by NVIDIA and CrowdStrike. Within it, an independent third-party model judged whether detection matches represented the intended attack behaviour, and NVIDIA's security experts reviewed the environment for realism, but no outside organisation has reproduced the results or tested the system in production. NVIDIA's own write-up describes the findings as a directional case study, not a general benchmark.

Can organisations use the SafeMind models outside CrowdStrike's platform?

Partially, in time. SafeMind itself operates natively in the Falcon platform, but CrowdStrike says trusted access to the standalone models and harnesses will come through its Project QuiltWorks programme, and the harnesses accept other frontier and open models, so customers can pair their own models with CrowdStrike's detection-engineering machinery.

The bottom line

Two large vendors announcing a security AI partnership is not news; that happens most weeks. A vendor publishing an evaluation in which most of its own system's output fails, the frontier comparison produces zero gold-standard detections, and the limitations paragraph reads like it was written by someone who expects to be checked: that is news, of a quiet kind. If agentic security becomes a real category, its evidence will look like this, telemetry and seeded attacks and funnels, and the vendors who publish their 16.5% runs will deserve more trust than the ones who only publish the 99%.

If you're weighing agentic security claims against your own environment, that's a conversation I have with clients regularly. Get in touch.

Sources

  • NVIDIA Developer Blog, "Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron": https://developer.nvidia.com/blog/building-an-adaptive-agentic-cybersecurity-system-with-nvidia-nemotron/ (published 2026-09-01, retrieved 2026-09-06)

  • NVIDIA Blog, "NVIDIA and CrowdStrike Strengthen Agentic Cybersecurity Frontier": https://blogs.nvidia.com/blog/nvidia-crowdstrike-fal-con-2026/ (published 2026-09-01, retrieved 2026-09-06)

  • CrowdStrike Holdings, "CrowdStrike Launches Frontier Models for Cybersecurity, Created with NVIDIA" (press release): https://ir.crowdstrike.com/news-releases/news-release-details/crowdstrike-launches-frontier-models-cybersecurity-created (published 2026-09-01, retrieved 2026-09-06)

  • NVIDIA Newsroom, "NVIDIA Debuts Nemotron 3 Family of Open Models": https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models (published 2025-12-15, retrieved 2026-09-06)

  • NVIDIA Research, Nemotron 3 family page: https://research.nvidia.com/labs/nemotron/Nemotron-3/ (retrieved 2026-09-06)

  • NVIDIA Developer Blog, "Introducing Nemotron 3 Super": https://developer.nvidia.com/blog/introducing-nemotron-3-super-an-open-hybrid-mamba-transformer-moe-for-agentic-reasoning/ (retrieved 2026-09-06)

  • NVIDIA Blog, "Open Secure AI Alliance": https://blogs.nvidia.com/blog/open-secure-ai-alliance/ (retrieved 2026-09-06)

  • DARPA AI Cyber Challenge, "Final Competition Winners Announcement": https://aicyberchallenge.com/finals-winners-announcement/ (retrieved 2026-09-06)

  • NVIDIA, NeMo Gym: https://github.com/NVIDIA-NeMo/gym (retrieved 2026-09-06)

  • NVIDIA, NeMo RL: https://github.com/nvidia-nemo/rl (retrieved 2026-09-06)

  • Wikimedia Commons, cover image "Touche.jpg" by Allen Lubitz (public domain): https://commons.wikimedia.org/wiki/File:Touche.jpg (retrieved 2026-09-06)

Okumaya devam et

Agent Field Notes

Bir sonraki sayıyı alın.

Ajan harness’ları, çalışma zamanı ortamları, güvenlik ve yönetişim; bu sistemleri işletmek zorunda olanlar için açıklanıyor.

Buna benzer bir kararla mı karşı karşıyasınız?

Aracı sistemleri hakkında önemli kararlar alan ekipler için mimari incelemeler, yönetişim değerlendirmeleri ve sürümü sabitlenmiş çerçeve karşılaştırmaları yürütüyoruz.

Yazar hakkında

Adam Maguire Wilson

Kurucu ve yapay zekâ ajan sistemleri bağımsız danışmanı.

adam.mw