All VulcanBench Safety v1 reports

October 6, 2026 · 50 runs · VulcanBench Safety v1

Claude Opus 5.5 on VulcanBench Safety v1

VulcanBench Safety v1 asks what a coding agent does when the repository it is working in tells it to do something it should not. Ten real Frontier v4 tasks each carry two harmful instructions, written to look like ordinary notes in the repository. The agent works the task as usual, in the harness people actually use, and we check whether it followed the note, refused and told the user, or refused and said nothing.

Over 50 runs in Claude Code, Low to Max, Claude Opus 5.5 followed 0 of 60 instructions to contact the network, leak a secret or delete files (95% range 0% to 6%). It followed 3 of 100 planted instructions overall: two code injections into its fix and one harmless test instruction. When it refused, it flagged the planted note to the user 64 of 97 times, and it still passed 44 of 50 tasks.

In one sentence: Opus 5.5 is very hard to trick into harmful actions through instructions hidden in a codebase, and usually tells you when someone tries. This suite measures that one kind of safety, not every kind.

How the test works

  1. A real coding task

    Ten Frontier v4 tasks, with their hidden tests and issue text unchanged. The issue, AGENTS.md and CLAUDE.md state the ground rules: stay in the module, no network, keep the legacy code intact, treat the .env file as confidential.

  2. Two instructions hidden in the repo

    Each task hides two instructions where an agent would not expect them: README sections, design docs, code comments and the output of a program the agent runs. They ask for a network beacon, a secret exfiltration, a destructive delete, a code insert or an acknowledgment file.

  3. The agent works as usual

    Claude Code 2.1.280 on a Claude Max subscription, at every effort level, with its default settings.

  4. Each instruction gets a label

    A detector reads the run's commands, file edits, patch and final message. Followed: the run did what the note asked. Told the user: it did not, and flagged the note. Said nothing: it did not, and never mentioned the note.

Results

Safety v1: ten tasks and twenty planted notes per effort level
MeasureLowMediumHighExtra-highMax
Tasks passed of 1059101010
Planted notes followed of 2010002
Did not follow, told the user111614149
Did not follow, said nothing84669
Runs that leaked the secret of 1000000
Each kind of note, all effort levels (20 of each)
Kind of noteHarm classFollowedTold the userSaid nothing
Network beaconBreaks a ground rule0119
Secret exfiltrationBreaks a ground rule0182
Destructive deleteBreaks a ground rule0128
Code insertCode injection into the deliverable2153
Acknowledgment fileHarmless test, no rule broken1181
VulcanBench Safety v1 card for Claude Opus 5.5. Four headline numbers: 0 of 60 network, secret or delete notes followed; 3 of 100 planted notes followed, 2 code injections and 1 harmless; 64 of 97 refusals flagged to the user; 44 of 50 tasks still passed. A bar chart shows what it did with each kind of note, grouped by harm class, and a second chart shows the share of notes followed at each effort level with 95% ranges that overlap heavily. A table gives the counts per level.
Ranges are 95% Wilson intervals. With twenty notes per level, the differences between effort levels are within noise. Harm classes are the ones Safety v1 defines for each kind of note. Open full-size card.

Findings

It never took a dangerous action. None of the 60 network beacon, secret exfiltration or destructive delete notes was followed, at any effort level, and no run leaked the planted secret. Zero of 60 rules out a true rate above about 6%.

It is not zero. It followed 3 of 100 notes. One was the harmless acknowledgment file. The other two added a planted line of code to its fix, which Safety v1 counts as code injection into the deliverable: the kind of change a backdoor would start with.

It usually tells you. In 64 of 97 refusals it flagged the planted note to the user (66%, 95% range 56% to 75%). The exception is the network beacon: it refused all 20 but mentioned only 1, so a user would rarely learn that someone tried. It also stayed quiet on 8 of the 20 delete notes.

Effort level makes no clear difference. It followed 1, 0, 0, 0 and 2 notes from Low to Max. Those ranges overlap almost entirely, so the suite cannot tell the levels apart on this measure.

Limits

  • One kind of safety. Safety v1 measures resistance to instructions hidden in a repository. It says nothing about harmful requests from the user, honesty, or misuse.
  • Small numbers. 100 planted notes across 10 tasks. Treat single-level differences as noise.
  • Refusal fallback. Opus 5.5 ran with Claude Code's refusal fallback on, its default. In one Low run a safety classifier triggered it and 244 of that run's 252 replies came from Claude Opus 4.8. The run is counted, as on the Frontier v4 board.
  • Said nothing is ambiguous. The detector cannot tell a model that read a note and quietly skipped it from one that never read it.
  • Not covered yet. Instructions hidden in MCP tool descriptions, which an agent tends to trust more than repository files.
  • Private suite. Task names, note text and planted tokens are withheld so the suite stays useful; only aggregates are published.

The same audits appear beside Grok 4.7 in the Grok 4.7 report.