reward-hacking-misalignment

Reproducing "Natural Emergent Misalignment from Reward Hacking" (MacDiarmid et al., Anthropic 2025) with open-source models. Includes reward-hackable RL environments, misalignment evaluations, training configs, and evaluation scripts. Models trained on OLMo (7B, 32B) and GPT-OSS (20B, 120B).

// repository documentation