For Frontier AI Labs: Train Models That Defend | Simbian
Your model can attack. Train it to defend.
Ask a frontier model to write an exploit and it delivers. Ask it to catch the breach and it goes blind. Not one has passed our Cyber Defense Benchmark. That skill was never in the pretraining data, so we build it: the real investigations you can grade, plus a deterministic environment that scores whether the model actually stopped the attack.
Trusted by leading enterprises and MSSPs
Cyber Defense Benchmark
The Frontier Model Cyber-Defense Checklist
HOW THE DATA IS BUILT 01 Real attack telemetry the behavior we start from 02 Modeled at org scale composed in Holodeck 03 Verified ground truth machine-checked where it can be 04 Licensed data, our IP clean to train on Customer /production logs never resold The moat is the modeling and the verification.
Can you fine-tune your way to cyber defense?
No. Defense is abductive work. You infer an attacker's intent from logs that look completely normal, and that reasoning barely shows up in pretraining. You can't fine-tune a skill the data never held. It has to be grown with RL against a real reward: data you can grade, and an environment that knows a good investigation from a bad one.
How do you get a verifiable reward for defense?
Holodeck. It's a deterministic, Gymnasium-compatible environment where the reward comes from machine-checkable truth, not a model's opinion. Same seed, same run, so nothing drifts and there's no LLM judge to game. The calls a machine can't make, like severity and escalation, go to expert human graders, and a held-out set catches anything trying to hack the reward.
Will it plug into our RLVR stack?
Yes. Tasks ship as JSONL with verifiers and Inspect-compatible formats, split into train, test, and held-out. Holodeck speaks Gymnasium, so your base model drops in without a custom harness. Deterministic and assertion graders give you an RL-usable reward; rubric and human graders sit alongside as acceptance and regression signal, not the reward itself.
0 of 25
frontier models pass the Cyber Defense Benchmark
45.1% best model's MITRE ATT&CK coverage (passing bar: 50%)
13 MITRE ATT&CK tactics measured end to end
Everything you need to grow the capability.
Verified data, a deterministic environment, expert validation, and a published eval. Take them on their own or together. Every piece is built for the RLVR toolchain you already run.
Verified security training data
Holodeck, a deterministic RL environment
Expert human validation
The Cyber Defense Benchmark
Real investigations, decomposed into gradeable tasks
We take real attack chains and investigations across threat hunting, detection engineering, triage, and offensive security, then break them into tasks a model can be scored on. Each one carries verified ground truth wherever the work can be checked by machine. It ships as JSONL and Inspect-compatible tasks, split into train, test, and held-out.
Verified security training data
A gym where a base model learns to hunt
Drop a base model into the loop. It observes, acts, and earns a reward, learning to run an investigation the way an analyst would. The reward comes from machine-checkable ground truth, so the same seed always produces the same run. Nothing drifts, and there's no LLM judge to talk your agent past. It plugs straight into the RLVR toolchain you already run.
Expert human validation
Career analysts grade what a machine can't
Severity, escalation, the judgment that separates a real analyst from a checklist: none of it comes from a generalist crowd. Career threat hunters and detection engineers grade those tasks and your model's outputs, and a held-out, human-graded set catches the reward hacking an environment can miss on its own.
The Cyber Defense Benchmark
The eval that hasn't been solved
A published benchmark, built to resist memorization, across the 13 MITRE ATT&CK tactics that make up a real breach. No frontier model has cleared the bar yet. Run it before you train to find the gaps, then run it again after to see exactly what moved.
Deterministic ground truth. No reward-model drift.
Most reward models start to drift the moment your agent learns to game them. Holodeck's doesn't. Its reward comes from machine-checkable truth, and the same seed always produces the same run. A correct investigation is a correct investigation. The moat is the modeling and the verification, not a pile of hoarded data.
Why frontier labs choose Simbian
You can't crowdsource security data
Generalist vendors staff doctors, lawyers, and bankers. We staff career threat hunters and detection engineers. Security is the one domain where the ground truth has to come from people who've actually run operations, not a crowd.
Deterministic, verifiable reward
Holodeck scores every investigation against machine truth. There's no reward model to drift and no LLM judge to sweet-talk. Same seed, same result, so the reward is exact and reproducible.
Verified and clean to license
Every trajectory is our own IP with verified ground truth, informed by operating security through the world's largest MDRs, not sourced from their logs. We don't resell customer or production data.
Bring a base model, measure the lift
We won't claim a number we haven't earned. Bring a base model, baseline it on the Cyber Defense Benchmark, train it in Holodeck, and we measure the lift together.
Questions from research teams
Where does the training data come from?
We build it. Our attack and organization models are built on real attacker behavior, composed into org-scale campaigns and verified, informed by operating security through the world's largest MDRs, not sourced from their logs. Every trajectory is our own IP with verified ground truth. We do not resell customer or production logs.
Do you sell data, environments, or both?
Both, and the published benchmark too. The verified training data, the Holodeck RL environment, and the Cyber Defense Benchmark are available independently or together.
Will this drop into our RLVR stack?
Yes. JSONL plus verifiers and Inspect-compatible task formats, with train, test, and held-out splits. Holodeck is Gymnasium-compatible, so a base model drops in with no bespoke harness. Deterministic and assertion graders give an RL-usable reward; rubric and human graders provide acceptance and regression signal, not the RL reward.
How do you stop reward hacking?
Holodeck's reward is deterministic and machine-checkable, so there's no reward model to drift and no LLM judge to game. Judgment tasks that can't be mechanically verified go to expert human graders, and a held-out human-graded set is there to catch Goodharting the verifiable reward.
Have you trained a model to a measured improvement?
The data, environment, and benchmark are built and published. Training a model on them is the joint experiment we propose: baseline on the Cyber Defense Benchmark, train in Holodeck, and measure the delta together.
What Our Customers Say
Simbian's AI Agents consistently deliver precise and accurate responses, significantly easing our workload. What used to take days now takes minutes, and we're thrilled with how seamlessly it integrates into our existing processes. It's not just about saving time; it's about maintaining the highest standards of security and accuracy, which is exactly what Simbian enables us to do.
Matillion
Suchit Mishra
Director of Information Security
Security is a domain of ever-increasing complexity. Every day a security incident brings new variables. Simbian is building a fully autonomous security platform. We are excited to partner with them as it allows us to be strategic in our security goals, leaving mechanics of security to Simbian.
Axelar
Sergey Gorbunov
Co-founder
Security partners, especially MSSPs and MDRs, are at a critical juncture. Attacks are getting accelerated with AI. We must use AI on defense side too. We have gotten great support from Simbian with its fully autonomous security. It allows us to do more with less, directly impacting both our top and bottom lines.
Cybalt
Khirodra Mishra
CEO
Simbian's platform takes a straightforward approach to solving core problems we see every day in the SOC. The power in the platform, their AI agents, is in its simplicity. They are not adding steps and processes to achieve results. The Security Accelerator platform drives efficiency without sacrificing efficacy. It allows us to shift the role of the analyst; to give them the time to use human insight, because well trained AI that we can review, and audit, is immensely powerful. It sets a whole new bar for security operations.
SMT
Mohammad Qasas
SOC Lead
Simbian's AI agents augment and automate many security services resulting into better efficiencies and increased precision.
Wipro
Siva VRS
Vice President
What Simbian's doing in that space has really been a differentiator and a game changer for how my team's thinking about these problems. We're no longer thinking about a pipeline of work that we've got to have 20 people to solve.
Bottomline
Blaine Brennecke
Director of Security Operations