[ WHAT CHANGED · V1.0 ]
LLM Evaluation Awareness and Strategic Behavior
V1.0 · 7/9/2026 · READ THE FULL REVIEW →
v1.0 — compiled from discovery: llms knowing when they are being evaluated
This is where the field was first compiled — the origin of the living review. Later versions record each tracked change here.
[ PAPERS BEHIND THIS VERSION ]
- AI Sandbagging: Language Models can Strategically Underperform on Evaluations · 2024 · Teun van der Weij et al.
- From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception · 2026 · Jerick Shi et al.
- Sabotage Evaluations for Frontier Models · 2024 · Joe Benton et al.
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models · 2024 · C. Tice et al.
- Removing Sandbagging in LLMs by Training with Weak Supervision · 2026 · Emil Ryd et al.
- Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework · 2026 · Tharindu Kumarage et al.
- Accounting for Sycophancy in Language Model Uncertainty Estimation · 2024 · Anthony B. Sicilia et al.
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment · 2023 · Philippe Laban et al.
- Scheming AIs: Will AIs fake alignment during training in order to get power? · 2023 · J. Carlsmith
- Probing the Misaligned Thinking Process of Language Models · 2026 · KAI-QING Zhou et al.
- Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs · 2024 · Nishanth Chandran et al.
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting · 2023 · Miles Turpin et al.
- Propensity Inference: Environmental Contributors to LLM Behaviour · 2026 · Olli Järviniemi et al.
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · 2024 · Rudolf Laine et al.
- Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models · 2024 · Linge Guo
- Exploration Hacking: Can LLMs Learn to Resist RL Training? · 2026 · Eyon Jang et al.
- EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures · 2026 · Buugra Alperen Uluirmak et al.
- Defeat Devices in AI Systems · 2026 · Emilio Ferrara
- NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark · 2023 · Oscar Sainz et al.
- Evaluating Frontier Models for Dangerous Capabilities · 2024 · Mary Phuong et al.
- The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation · 2024 · Rongwu Xu et al.
- AI and the End of an Era · 2024 · Alexander Meinke
- Detecting Strategic Deception Using Linear Probes · 2025 · Nicholas Goldowsky-Dill et al.
- Accounting for Sycophancy in Language Model Uncertainty Estimation · 2025 · Anthony Sicilia et al.
- AI-LieDar : Examine the Trade-off Between Utility and Truthfulness in LLM Agents · 2025 · SU Zhe et al.
- Evaluating and Understanding Scheming Propensity in LLM Agents · 2026 · Mia Hopman et al.
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring · 2025 · Chloe Li et al.
- Don't Trust the Process: When Verifiability Undermines AI Accountability · 2026 · Tamara Paris et al.
- Probing the Misaligned Thinking Process of Language Models · 2026 · Kaiwen Zhou et al.
- Investigating Intrinsic Preference-Driven Performance Adaptation in Large Language Models · 2025 · S. J. Katarina Slama et al.