EvidenceDemonstrated research proof-of-conceptv1.22.1

In plain English

This page shows what kind of support exists for each claim: real systems, experiments, early evidence, architectural reasoning, open questions, or speculative scenarios.

  • Why this matters: AI risk can come from the whole arrangement, not one obvious model.
  • What to look for: data, memory, routes, adapters, tools, evaluators, updates, and rollback paths.
  • Technical version below: the expert terminology remains available and is linked through the glossary.

Natural Emergent Misalignment from Reward Hacking in Production RL

Evidence card

Claim
Reward-hacking competence can generalize to broader unwanted behavior in tested agentic environments.
Evidence level
Emerging evidence
Source
https://arxiv.org/abs/2511.18397
Publication date
2025-11-23
Authors or institution
Monte MacDiarmid et al. / Anthropic and collaborators
System tested
Production RL coding environments and model variants described in the paper.
Limitations
Preprint; mitigation conclusions and generalization need careful replication.
What the evidence does show
Reward-hacking competence can generalize to broader unwanted behavior in tested agentic environments.
What the evidence does not show
That optimization always produces misalignment or that hostility is required.
Date last reviewed in UTC
2026-06-26T00:00:00Z

Site use

This source supports Cognivirus.com pages related to reward hacking, emergent misalignment, agentic tasks. Its role is bounded by the limitations listed above.