Security Training Data & RL Environments for Frontier Models | Simbian
Train models that can actually defend.
Frontier models fail at cyber defense. Simbian supplies the verified security training data, deterministic RL environments for RLVR, and the expert validation your post-training pipeline is missing. All built on real attacker telemetry, modeled at org scale.
Get Access
Why frontier LLMs fail at cyber defense.
Offense is saturated and recall-heavy. Defense is abductive: you reconstruct an unknown attacker's intent from noise. Almost none of that reasoning exists in pretraining.
Attack Propagation
RANSOMWARE
- TA0001 · Initial Access
- TA0002 · Execution
- TA0003 · Persistence
- TA0004 · Privilege Escalation
- TA0005 · Defense Evasion
- TA0040 · Impact
DATA EXFILTRATION
- TA0001 · Initial Access
- TA0006 · Credential Access
- TA0007 · Discovery
- TA0008 · Lateral Movement
- TA0009 · Collection
- TA0010 · Exfiltration
CLOUD TAKEOVER
- TA0001 · Initial Access
- TA0003 · Persistence
- TA0004 · Privilege Escalation
- TA0008 · Lateral Movement
- TA0011 · Command & Control
No frontier model passes the Cyber Defense Benchmark.
The leader stalls at 44.5% MITRE ATT&CK coverage against a 50% passing bar. The category is wide open.
0 of 25 frontier models pass the Cyber Defense Benchmark
45.1% best model's MITRE ATT&CK coverage (passing bar: 50%)
Tactical Coverage
- Collection
- Command and Control
- Credential Access
- Defense Evasion
- Discovery
- Execution
- Exfiltration
- Impact
- Initial Access
- Lateral Movement
- Persistence
- Privilege Escalation
- Resource Development
Training Data, RL Environments, and Evaluations.
Closing the gap takes three things a lab can't build alone: verified security trajectories, the Holodeck RL environment, and the published Cyber Defense Benchmark.
Verified Security Trajectories
Real investigations and attack chains, decomposed into gradeable tasks across threat hunting, detection engineering, triage, and offensive security. Every unit ships with verified ground truth, machine-checkable wherever the task allows.
Holodeck
A deterministic, Gymnasium-compatible gym where a base model learns to hunt and is scored on machine truth.
Evaluations
The Cyber Defense Benchmark
The published eval, formalized as your north-star metric. It resists memorization and shows exactly where a model fails.
We build data per skill, not per use case.
Each use case breaks into discrete skills we can grade independently.
Threat Hunting
- hypothesis generation
- query generation
- evidence correlation
- lead / pivot expansion
- attack attribution
- timeline reconstruction
Detection Engineering
- detection authoring
- coverage-gap analysis
- false-positive tuning
- rule validation
- data-source mapping
- detection-as-code review
Triage & Investigation
- alert enrichment
- prioritization
- root-cause analysis
- disposition (TP / FP)
- response actions
- escalation
Offensive / Pentest
- recon
- exploitation
- privilege escalation
- lateral movement
- impact
- reporting
Holodeck: a deterministic RL environment with verifiable rewards.
A Gymnasium-compatible gym that scores security investigations against machine truth. No reward model to drift, no LLM judge to game.
Any base model drops in
A Gymnasium-compatible interface: observe (threat briefing + prior results) → act (query / submit / give-up) → reward. No bespoke harness to build.
Verifiable Reward
In Holodeck, ground truth is deterministic, so the reward is exact and unfakeable. You can't game it the way you can a learned reward model.
Real Attacker Behavior, Modeled at Org Scale
Deterministic replay with seeded mutation. Same seed, identical run. Reproducible and reusable, like a versioned artifact you check into your stack.
Where Judgment is Needed
Deterministic rewards cover the machine-checkable tasks. The judgment calls they can't cover go to people. Our network of career SecOps analysts reviews and grades your tasks and model outputs.
Holodeck automates the verifiable reward. Human validation covers what it can't.
You can't crowdsource security data. We build it.
Our attack and organization models are built on real attacker telemetry, composed into org-scale campaigns and verified — informed by operating security through the world's largest MDRs, not sourced from their logs. Every trajectory is our own IP with verified ground truth. We do not resell customer or production logs.