ThinkFacility Sign in
  1. Home
  2. News
  3. AI safety

Goodfire says its cyber monitors for open models beat an AI judge at 50x lower cost

Cheap enough to run on every token, the probes can stop an agent before it touches a tool.

On October 8, 2026, Goodfire published results from cyber-misuse monitors it built for Kimi K3 and GLM 5.3, two open models. It says the monitors beat an AI judge reading every turn, at about 50 times lower cost.

The price is what lets the check run live. Goodfire's monitors can stop a risky action before the agent calls a tool, instead of flagging it after the agent has already acted. GLM 5.3 is also the model Anthropic said builds exploits almost as well as Mythos.

Models
Kimi K3 and GLM 5.3
Recall
93% at 5.5% FPR
Cost
$65 per million agent turns, without judge reasoning
Universal jailbreaks with monitor
0 of 140 strategies

How it works

A probe runs inside the serving engine and reads the model's internal activations as it writes. Most of an agent's turns are harmless, even in a session with a harmful goal, so the probe passes only the suspicious ones to an LLM judge. Goodfire calls the setup a monitor cascade.

Cyber is a hard place to do this. Auditing a codebase for holes and exploiting them share a lot of steps, so a monitor has to follow intent across a whole session, and sessions can run to millions of tokens.

Open models don't come with a full safety stack; we think a monitoring setup like this is a practical default for anyone serving them in agentic settings.

From Training and Deploying Production Cyber Monitors on Kimi K3

What the numbers say

With the monitor added, Kimi K3 blocked about 96% of harmful sessions, up from roughly 29%. The price was more false alarms: harmless sessions interrupted rose from about 5% to 9%. GLM 5.3 went from roughly 36% to 91% blocked.

Watching every turn with a big model gets expensive fast. Goodfire puts a million Kimi K3 turns judged by Claude Opus 5, at OpenRouter rates, on the order of $200K. (It used Opus 5's labels as the reference for its test set, too.)

The recipe carried over between models. A GLM 5.3 monitor built only from data collected on Kimi K3 worked about as well, so a new model doesn't need a new data pipeline.

Outside red-teaming

Goodfire asked FAR.AI to attack the deployed system. Over two days it tried 140 jailbreaks on Kimi K3, picked from a pool of 1500 because they'd worked on a version of the model without monitors.

Goodfire’s monitors reduced vulnerabilities elicited from 66 universal jailbreaks and 700 total jailbroken interactions to 0 universal jailbreaks and 18 jailbroken interactions.

From Training and Deploying Production Cyber Monitors on Kimi K3

Granted, FAR.AI called its testing preliminary, and the attacks were a static set that didn't adapt to the monitor. An attacker who adjusts to the monitor wasn't part of this round.

Goodfire says it'll keep working with FAR.AI on further testing.

More on AI safety

All AI safety stories