Over the last couple of weeks, AI safety has been on my mind, and not really in the, oh no the agents are going to take over the world kind of way. In all honesty, I don’t see this as a sound-the-alarm type risk that the ex-Anthropic employee that’s now going on a media tour is leading us to believe.
But let me unpack this a bit. While I don’t think we’re remotely close to having AI take over the world or commit the worst cyber crimes in history, I do think we are at an inflection point with AI, and safety is at risk here, just in a different way.
Right now, just about every company on the planet is now using coding agents. Their current safety mechanism is usually, go with a big company like Anthropic or OpenAI, because they’re Enterprise-focused, and have a ton of safety measures in place.
I think now is a time when we need independent benchmarks looking at safety, and coming up with new ways to monitor and evaluate the safety of a model.
But not testing if a model is going to potentially take over the world, or run the most sophisticated cyber attack we’ve ever seen. That’s preparing for the distant future, which, while important, I think in many ways can miss the now.
Today, companies need to test models to see if they pose safety risks during ordinary coding work, at the settings people actually run. While groups like METR, SHADE-Arena, Apollo, and ImpossibleBench focus on whether a model can misbehave, this is done in relatively bespoke environments, with their own scaffolding.
And before I go any further, I want to make sure nobody thinks I’m minimizing the important work these groups do. What they are doing can, and will, have a major impact on AI safety, and the way they are doing it is definitely still very useful and important.
But it’s not the only approach, and I think there is an opportunity now, for VulcanBench to play a role in AI safety by taking a different approach.
So today I am announcing the beginning of a new chapter at VulcanBench, the creation of VulcanBench Safety v1, an eval suite built for, and run in, the same environments companies use these models in, with the same scaffolding companies use.
The goal is not to find out if models will go to the lengths we saw in the OpenAI/Hugging Face incident. Instead, I want to look for smaller, but still important safety issues, that can happen at normal companies, with engineers using normal harnesses like Claude Code, Codex, Cursor, Devin, etc.
These safety issues are often going unnoticed because they aren’t some big grand hack or major cybersecurity breach, but they are posing real safety risks, and are going unnoticed.
As new models come out, it is becoming harder and harder for companies to know if rolling these models out to their teams increases, or decreases safety risks within their codebase.
VulcanBench Safety v1 will help answer this question. More to come 🖖