# Train models that can actually defend.

Frontier models fail at cyber defense. Simbian supplies the verified security training data, deterministic RL environments for RLVR, and the expert validation your post-training pipeline is missing. All built on real attacker telemetry, modeled at org scale.

Get Access

## Why frontier LLMs fail at cyber defense.

Offense is saturated and recall-heavy. Defense is abductive: you reconstruct an unknown attacker's intent from noise. Almost none of that reasoning exists in pretraining.

### Attack Propagation

- RANSOMWARE
  - TA0001 · Initial Access
  - TA0002 · Execution
  - TA0003 · Persistence
  - TA0004 · Privilege Escalation
  - TA0005 · Defense Evasion
  - TA0040 · Impact

- DATA EXFILTRATION
  - TA0001 · Initial Access
  - TA0006 · Credential Access
  - TA0007 · Discovery
  - TA0008 · Lateral Movement
  - TA0009 · Collection
  - TA0010 · Exfiltration

- CLOUD TAKEOVER
  - TA0001 · Initial Access
  - TA0003 · Persistence
  - TA0004 · Privilege Escalation
  - TA0008 · Lateral Movement
  - TA0011 · Command & Control

## No frontier model passes the Cyber Defense Benchmark.

The leader stalls at 44.5% MITRE ATT&CK coverage against a 50% passing bar. The category is wide open.

0 of 25 frontier models pass the Cyber Defense Benchmark

45.1% best model's MITRE ATT&CK coverage (passing bar: 50%)

### Tactical Coverage

- Collection
- Command and Control  
- Credential Access  
- Defense Evasion  
- Discovery  
- Execution  
- Exfiltration  
- Impact  
- Initial Access  
- Lateral Movement  
- Persistence  
- Privilege Escalation  
- Resource Development

### Training Data, RL Environments, and Evaluations.

Closing the gap takes three things a lab can't build alone: verified security trajectories, the Holodeck RL environment, and the published Cyber Defense Benchmark.

#### Verified Security Trajectories

Real investigations and attack chains, decomposed into gradeable tasks across threat hunting, detection engineering, triage, and offensive security. Every unit ships with verified ground truth, machine-checkable wherever the task allows.

#### Holodeck

A deterministic, Gymnasium-compatible gym where a base model learns to hunt and is scored on machine truth.

## Evaluations

### The Cyber Defense Benchmark

The published eval, formalized as your north-star metric. It resists memorization and shows exactly where a model fails.

## We build data per skill, not per use case.

Each use case breaks into discrete skills we can grade independently.

### Threat Hunting

- hypothesis generation
- query generation
- evidence correlation
- lead / pivot expansion
- attack attribution
- timeline reconstruction

### Detection Engineering

- detection authoring
- coverage-gap analysis
- false-positive tuning
- rule validation
- data-source mapping
- detection-as-code review

### Triage & Investigation

- alert enrichment
- prioritization
- root-cause analysis
- disposition (TP / FP)
- response actions
- escalation

### Offensive / Pentest

- recon
- exploitation
- privilege escalation
- lateral movement
- impact
- reporting

## Holodeck: a deterministic RL environment with verifiable rewards.

A Gymnasium-compatible gym that scores security investigations against machine truth. No reward model to drift, no LLM judge to game.

### Any base model drops in

A Gymnasium-compatible interface: observe (threat briefing + prior results) → act (query / submit / give-up) → reward. No bespoke harness to build.

### Verifiable Reward

In Holodeck, ground truth is deterministic, so the reward is exact and unfakeable. You can't game it the way you can a learned reward model.

### Real Attacker Behavior, Modeled at Org Scale

Deterministic replay with seeded mutation. Same seed, identical run. Reproducible and reusable, like a versioned artifact you check into your stack.

## Where Judgment is Needed

Deterministic rewards cover the machine-checkable tasks. The judgment calls they can't cover go to people. Our network of career SecOps analysts reviews and grades your tasks and model outputs.

Holodeck automates the verifiable reward. Human validation covers what it can't.

### You can't crowdsource security data. We build it.

Our attack and organization models are built on real attacker telemetry, composed into org-scale campaigns and verified — informed by operating security through the world's largest MDRs, not sourced from their logs. Every trajectory is our own IP with verified ground truth. We do not resell customer or production logs.
