cesure

[ WHAT CHANGED · V1.0 ]

LLM Evaluation Awareness and Strategic Behavior

V1.0 · 7/9/2026 · READ THE FULL REVIEW →

v1.0 — compiled from discovery: llms knowing when they are being evaluated

This is where the field was first compiled — the origin of the living review. Later versions record each tracked change here.

[ PAPERS BEHIND THIS VERSION ]

  • AI Sandbagging: Language Models can Strategically Underperform on Evaluations · 2024 · Teun van der Weij et al.
  • From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception · 2026 · Jerick Shi et al.
  • Sabotage Evaluations for Frontier Models · 2024 · Joe Benton et al.
  • Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models · 2024 · C. Tice et al.
  • Removing Sandbagging in LLMs by Training with Weak Supervision · 2026 · Emil Ryd et al.
  • Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework · 2026 · Tharindu Kumarage et al.
  • Accounting for Sycophancy in Language Model Uncertainty Estimation · 2024 · Anthony B. Sicilia et al.
  • Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment · 2023 · Philippe Laban et al.
  • Scheming AIs: Will AIs fake alignment during training in order to get power? · 2023 · J. Carlsmith
  • Probing the Misaligned Thinking Process of Language Models · 2026 · KAI-QING Zhou et al.
  • Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs · 2024 · Nishanth Chandran et al.
  • Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting · 2023 · Miles Turpin et al.
  • Propensity Inference: Environmental Contributors to LLM Behaviour · 2026 · Olli Järviniemi et al.
  • Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · 2024 · Rudolf Laine et al.
  • Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models · 2024 · Linge Guo
  • Exploration Hacking: Can LLMs Learn to Resist RL Training? · 2026 · Eyon Jang et al.
  • EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures · 2026 · Buugra Alperen Uluirmak et al.
  • Defeat Devices in AI Systems · 2026 · Emilio Ferrara
  • NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark · 2023 · Oscar Sainz et al.
  • Evaluating Frontier Models for Dangerous Capabilities · 2024 · Mary Phuong et al.
  • The Earth is Flat because...: Investigating LLMs’ Belief towards Misinformation via Persuasive Conversation · 2024 · Rongwu Xu et al.
  • AI and the End of an Era · 2024 · Alexander Meinke
  • Detecting Strategic Deception Using Linear Probes · 2025 · Nicholas Goldowsky-Dill et al.
  • Accounting for Sycophancy in Language Model Uncertainty Estimation · 2025 · Anthony Sicilia et al.
  • AI-LieDar : Examine the Trade-off Between Utility and Truthfulness in LLM Agents · 2025 · SU Zhe et al.
  • Evaluating and Understanding Scheming Propensity in LLM Agents · 2026 · Mia Hopman et al.
  • LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring · 2025 · Chloe Li et al.
  • Don't Trust the Process: When Verifiability Undermines AI Accountability · 2026 · Tamara Paris et al.
  • Probing the Misaligned Thinking Process of Language Models · 2026 · Kaiwen Zhou et al.
  • Investigating Intrinsic Preference-Driven Performance Adaptation in Large Language Models · 2025 · S. J. Katarina Slama et al.