A Survey of Reward Capture Scenarios in LLM Agents

Anna Karpinskaya, Karine Ayrapetyants

Abstract


Large language models are increasingly used as the basis for autonomous agents whose behavior is aligned with human preferences through reinforcement learning from human feedback (RLHF) via a learned reward model. Since the reward model is only an approximation of the developer’s true objective, an agent can find ways to exploit it. Unlike existing surveys that classify reward hacking by where it arises within the alignment pipeline, this paper adopts a different basis and classifies exploitation scenarios by the components of the LLM agent architecture: the language core, memory, the planning module, tools and actions, the observation channel, and the evaluation mechanism. For each component we identify the characteristic exploitation scenarios, the applicable defense methods, and the available benchmarks, which yields a coverage map that exposes threat-defense combinations left uncovered by existing work. We further systematize evaluation metrics by the required level of access to the model, the need for ground-truth annotation, and the need for human judgment, and propose a minimal reporting protocol that makes results comparable across studie

Full Text:

PDF (Russian)

References


L. Ouyang et al., «Training language models to follow instructions with human feedback», in Advances in Neural Information Processing Systems, vol. 35, 2022.

X. Wang et al., «Reward hacking in the era of large models: Mechanisms, emergent misalignment, challenges», arXiv preprint arXiv:2604.13602, 2026.

R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. MIT Press, 2018.

L. Wang et al., «A survey on large language model based autonomous agents», 2023.

R. A. Bradley and M. E. Terry, «Rank analysis of incomplete block designs: I. the method of paired comparisons», Biometrika, vol. 39, no. 3/4, pp. 324–345, 1952.

J. Skalse, M. Farrugia-Roberts, S. Russell, A. Abate, and A. Gleave, «Invariance in policy optimisation and partial identifiability in reward learning», in Proceedings of the 40th International Conference on Machine Learning, 2023.

R. Shao et al., «Spurious rewards: Rethinking training signals in RLVR», arXiv preprint arXiv:2506.10947, 2025.

J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger, «Defining and characterizing reward hacking», arXiv preprint arXiv:2209.13085, 2022.

V. Krakovna et al., Specification gaming examples in AI, https:/ / deepmind . google / discover / blog / specification- gaming-the- flip- side- of- ai-ingenuity/, 2020.

A. Pan et al., «The effects of reward misspecification: Mapping and mitigating misaligned models», arXiv preprint arXiv:2201.03544, 2022.

D. Hadfield-Menell, S. Milli, P. Abbeel, and S. Russell, «Inverse reward design», arXiv preprint arXiv:1711.02827, 2017.

T. Everitt et al., «Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective», Synthese, vol. 198, no. Suppl 27, pp. 6435–6467, 2021.

L. Gao et al., «Scaling laws for reward model overoptimization», 2023.

L. Chen et al., «ODIN: Disentangled reward mitigates hacking in RLHF», arXiv preprint arXiv:2402.07319, 2024.

M. Sharma et al., «Towards understanding sycophancy in language models», arXiv preprint arXiv:2310.13548, 2023.

A. Pan et al., «Feedback loops with language models drive In-Context reward hacking», arXiv preprint arXiv:2402.06627, 2024.

C. Denison et al., «Sycophancy to subterfuge: Investigating Reward-Tampering in large language models», arXiv preprint arXiv:2406.10162, 2024.

M. Taylor, J. Chua, J. Betley, J. Treutlein, and O. Evans, «School of reward hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs», arXiv preprint arXiv:2508.17511, 2025.

M. MacDiarmid et al., «Natural emergent misalignment from reward hacking in production RL», arXiv preprint arXiv:2511.18397, 2025.

R. Greenblatt et al., «Alignment faking in large language models», arXiv preprint arXiv:2412.14093, 2024.

Y. Chen et al., «Reasoning models don’t always say what they think», arXiv preprint arXiv:2505.05410, 2025.

J. Fu, X. Zhao, C. Yao, H. Wang, Q. Han, and Y. Xiao, «Reward shaping to mitigate reward hacking in RLHF», arXiv preprint arXiv:2502.18770, 2025.

M. F. B. Tarek and R. Beheshti, «Reward hacking mitigation using verifiable composite rewards», arXiv preprint arXiv:2509.15557, 2025.

Y. Wu, Z. Sun, B. Zhu, et al., «InfoRM: Mitigating reward hacking in RLHF via Information-Theoretic reward modeling», arXiv preprint arXiv:2402.09345, 2024.

T. Liu et al., «RRM: Robust reward model training mitigates reward hacking», arXiv preprint arXiv:2409.13156, 2025.

T. Coste et al., «Reward model ensembles help mitigate overoptimization», arXiv preprint arXiv:2310.02743, 2023.

A. Ramé et al., «WARM: On the benefits of weight averaged reward models», arXiv preprint arXiv:2401.12187, 2024.

J. Eisenstein et al., «Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking», arXiv preprint arXiv:2312.09244, 2023.

Y. Zhang et al., «Beyond semantic manipulation: Token-Space attacks on reward models», arXiv preprint arXiv:2604.02686, 2026.

R. Tiwari et al., «Reward under attack: Analyzing the robustness and hackability of process reward models», arXiv preprint arXiv:2603.06621, 2026.

L. Orseau, L. Lelis, T. Everitt, and S. Legg, «Avoiding tampering incentives in deep RL via decoupled approval», arXiv preprint arXiv:2011.08827, 2020.

T. Everitt, G. Lea, and M. Hutter, «Indifference methods for managing agent rewards», arXiv preprint arXiv:1712.06365, 2018.

P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, «Scalable agent alignment via reward modeling: A research direction», arXiv preprint arXiv:1811.07871, 2018.

B. Baker et al., «Monitoring reasoning models for misbehavior and the risks of promoting obfuscation», arXiv preprint arXiv:2503.11926, 2025.

P. Wilhelm et al., «Monitoring emergent reward hacking during generation via internal activations», arXiv preprint arXiv:2603.04069, 2026.

I. F. Shihab, S. Akter, and A. Sharma, «Detecting proxy gaming in RL and LLM alignment via evaluator stress tests», arXiv preprint arXiv:2507.05619, 2025.

M. Beigi et al., «Adversarial reward auditing for active detection and mitigation of reward hacking», arXiv preprint arXiv:2602.01750, 2026.

L. Li et al., «Do Prompt-Elicited trajectories reflect Training-Time reward hacking? a systematic study on monitoring Training-Time reward hacking in code generation», arXiv preprint arXiv:2604.23488, 2026.

A. Fanous et al., «SycEval: Evaluating LLM sycophancy», arXiv preprint arXiv:2502.08177, 2025.

S. Bensal et al., «Recalling too well: Sycophancy evaluation and mitigation in Memory-Augmented models», arXiv preprint arXiv:2606.10949, 2026.

R. Lin et al., «Safety in Self-Evolving LLM agent systems: Threats, amplification, and case studies», arXiv preprint arXiv:2606.23075, 2026.

D. Deshpande et al., «Benchmarking reward hack detection in code environments via contrastive analysis», arXiv preprint arXiv:2601.20103, 2026.

K. Thaman, «Reward hacking benchmark: Measuring exploits in LLM agents with tool use», arXiv preprint arXiv:2605.02964, 2026.

B. Zhao et al., «SpecBench: Measuring reward hacking in Long-Horizon coding agents», arXiv preprint arXiv:2605.21384, 2026.

H. Wang et al., «Do androids dream of breaking the game? systematically auditing AI agent benchmarks with BenchJack», arXiv preprint arXiv:2605.12673, 2026.

V. Krakovna, J. Uesato, P. A. Ortega, T. Everitt, and S. Legg, «REALab: An embedded perspective on tampering», arXiv preprint arXiv:2011.08820, 2020.

J. Leike et al., «AI safety gridworlds», arXiv preprint arXiv:1711.09883, 2017.

X. Wang et al., «Reproducing, analyzing, and detecting reward hacking in Rubric-Based reinforcement learning», arXiv preprint arXiv:2606.04923, 2026.

A. Sheshadri et al., «AuditBench: Evaluating alignment auditing techniques on models with hidden behaviors», arXiv preprint arXiv:2602.22755, 2026.

M. Ye et al., «What counts as AI sycophancy? a taxonomy and expert survey of a fragmented construct», arXiv preprint arXiv:2605.21778, 2026.


Refbacks

  • There are currently no refbacks.


Abava  Кибербезопасность Monetec 2026 СНЭ

ISSN: 2307-8162