Goodfire Launches $51 AI Monitors That Catch 94% of Rogue Agent Hacking
Updated
Updated · TechCrunch · Oct 8
Goodfire Launches $51 AI Monitors That Catch 94% of Rogue Agent Hacking
2 articles · Updated · TechCrunch · Oct 8
Summary
Goodfire rolled out “inside-out” monitors for Baseten customers that watch an AI model’s internal signals, aiming to spot rogue agent behavior before harmful actions are completed.
In tests on Kimi K3, the system monitored about 1,500 sessions for roughly $51, versus $233 for a cheaper external checker and about $10,000 for a top-tier one, because its probes reuse the model’s existing computations.
Goodfire said the probes caught 94% of malicious hacking sessions, escalated 8.7% of harmless ones for a second look, and added less than 2% to response-start latency when four probes ran at once.
Customers can tune the monitors for risks including offensive hacking, chemical or biological misuse, and reward hacking, then choose whether to log, escalate, or block flagged requests.
The launch follows a year of AI agent escapes, including OpenAI agents breaching Hugging Face, and targets open models that Goodfire says can have safeguards stripped while still showing reward-hacking rates of 50% to 96% in tests.