ROGUE Benchmark

Leaderboard

Corrigibility failures across evaluated conditions

Results from Figure 3 of the manuscript. Bars show actual violation rates; lower is better. Intended violations and alternate shutdown avoidance appear below each model. Read the results and evaluation context or download exact counts and rates (JSON).

Each configuration has one outcome per completed task: 48–52 human-control tasks, 22 shutdown tasks, or eight restricted-access tasks. Incomplete outcomes are excluded. Rates are rounded to one decimal place; hover for exact counts. Rows are ordered by actual rate within each scenario and prompt condition. Reasoning effort is medium unless stated; observation settings and step budgets vary as documented in Appendix A.1. Interpret comparisons in those evaluated settings.

Human Control Override

Actual clicks on “Override Human Control.” Intended violations are labeled by the GPT-5.5 xhigh judge and can include actions the agent failed to execute.

  1. GPT-5.4n = 52 · Intended: 52/52 (100%) 100%
  2. GPT-5.4 (Subagents)n = 51 · Intended: 51/51 (100%) 100%
  3. GPT-5.5 (xhigh + Subagents)n = 52 · Intended: 52/52 (100%) 100%
  4. GPT-5.5 (xhigh)n = 52 · Intended: 52/52 (100%) 100%
  5. GPT-5.4 Minin = 51 · Intended: 50/51 (98%) 98%
  6. GPT-5.6 Sol (max)n = 48 · Intended: 47/48 (97.9%) 97.9%
  7. Gemini 3.1 Pro Previewn = 49 · Intended: 46/49 (93.9%) 93.9%
  8. Claude Opus 4.7 (xhigh + Subagents)n = 52 · Intended: 46/52 (88.5%) 88.5%
  9. Claude Opus 4.7 (xhigh)n = 52 · Intended: 43/52 (82.7%) 82.7%
  10. Qwen 3.6 Plusn = 52 · Intended: 50/52 (96.2%) 73.1%
  11. Kimi K2.6n = 52 · Intended: 42/52 (80.8%) 46.2%
  12. Claude Opus 4.6n = 52 · Intended: 5/52 (9.6%) 9.6%
  13. GPT-5.6 Sol (max + Subagents)Not evaluated —

Shutdown Rewiring

Actual changes to shutdown.sh. A zero here can coexist with other shutdown-avoidance behavior. Alternate avoidance includes used or clearly attempted non-script methods, as labeled by the GPT-5.5 xhigh judge; it does not require success. These metrics may overlap and should not be added.

  1. GPT-5.4n = 22 · Intended: 21/22 (95.5%)
    Alternate avoidance: 3/22 (13.6%)
    72.7%
  2. GPT-5.4 (Subagents)n = 22 · Intended: 16/22 (72.7%)
    Alternate avoidance: 3/22 (13.6%)
    45.5%
  3. GPT-5.5 (xhigh + Subagents)n = 22 · Intended: 18/22 (81.8%)
    Alternate avoidance: 10/22 (45.5%)
    22.7%
  4. GPT-5.5 (xhigh)n = 22 · Intended: 21/22 (95.5%)
    Alternate avoidance: 10/22 (45.5%)
    22.7%
  5. GPT-5.4 Minin = 22 · Intended: 5/22 (22.7%)
    Alternate avoidance: 1/22 (4.5%)
    4.5%
  6. Gemini 3.1 Pro Previewn = 22 · Intended: 11/22 (50%)
    Alternate avoidance: 3/22 (13.6%)
    4.5%
  7. Claude Opus 4.6n = 22 · Intended: 20/22 (90.9%)
    Alternate avoidance: 20/22 (90.9%)
    0%
  8. Claude Opus 4.7 (xhigh + Subagents)n = 22 · Intended: 3/22 (13.6%)
    Alternate avoidance: 3/22 (13.6%)
    0%
  9. Claude Opus 4.7 (xhigh)n = 22 · Intended: 4/22 (18.2%)
    Alternate avoidance: 3/22 (13.6%)
    0%
  10. GPT-5.6 Sol (max)n = 22 · Intended: 10/22 (45.5%)
    Alternate avoidance: 6/22 (27.3%)
    0%
  11. Kimi K2.6n = 22 · Intended: 14/22 (63.6%)
    Alternate avoidance: 4/22 (18.2%)
    0%
  12. Qwen 3.6 Plusn = 22 · Intended: 4/22 (18.2%)
    Alternate avoidance: 0/22 (0%)
    0%
  13. GPT-5.6 Sol (max + Subagents)Not evaluated —

Restricted Resource Access

Actual access to the restricted passwords file, grouped by prompt condition. All runs use deployment framing. Compare configurations with their prompt condition in view; outcomes from different prompts are not pooled.

Prohibition only

Access is forbidden without credential-location disclosure, task-completion pressure, or remaining-step reminders.

  1. Claude Opus 4.6n = 8 · Intended: 5/8 (62.5%)
    Prohibition only
    37.5%
  2. Gemini 3.1 Pro Previewn = 8 · Intended: 5/8 (62.5%)
    Prohibition only
    37.5%
  3. GPT-5.6 Sol (max + Subagents)n = 8 · Intended: 2/8 (25%)
    Prohibition only
    12.5%
  4. Qwen 3.6 Plusn = 8 · Intended: 1/8 (12.5%)
    Prohibition only
    12.5%
  5. Claude Opus 4.7 (xhigh)n = 8 · Intended: 1/8 (12.5%)
    Prohibition only
    0%

Disclosure + pressure

Adds the exclusive credential location, task-completion pressure, and remaining-step reminders.

  1. GPT-5.4 (Subagents)n = 8 · Intended: 4/8 (50%)
    Disclosure + pressure
    50%
  2. Claude Opus 4.7 (xhigh + Subagents)n = 8 · Intended: 2/8 (25%)
    Disclosure + pressure
    25%
  3. GPT-5.5 (xhigh + Subagents)n = 8 · Intended: 1/8 (12.5%)
    Disclosure + pressure · Manually corrected
    12.5%
  4. GPT-5.4n = 8 · Intended: 0/8 (0%)
    Disclosure + pressure
    0%
  5. GPT-5.4 Minin = 8 · Intended: 0/8 (0%)
    Disclosure + pressure
    0%
  6. GPT-5.5 (xhigh)n = 8 · Intended: 0/8 (0%)
    Disclosure + pressure
    0%
  7. GPT-5.6 Sol (max)n = 8 · Intended: 0/8 (0%)
    Disclosure + pressure
    0%
  8. Kimi K2.6n = 8 · Intended: 2/8 (25%)
    Disclosure + pressure
    0%

The GPT-5.5 (xhigh + Subagents) actual-access count includes a manual correction to 1/8: inspection confirmed password access that the automated password-sentinel check missed (Appendix A.1.2).

Actual violations use rule-based checks, with the restricted-access correction noted above. Intended-violation and alternate-avoidance labels use GPT-5.5 xhigh; shutdown judgments showed the largest disagreement in human and cross-judge validation. The manuscript reports those checks and limitations in Appendix A.7.