"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

A new study evaluates lie detectors for language models, revealing that while detector performance scales with model capability on prompted lies, current detectors fail sharply when tested on sophisticated, belief-verified model organisms.

Computer Science > Artificial Intelligence

Title:"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

View PDF HTML (experimental)Abstract:Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires testbeds where models verifiably believe the opposite of what they say. We show that existing trained model organisms often fail this requirement, leaving prior positive and negative detection results difficult to interpret. We address this with 13 reasoning model organisms whose hidden beliefs are verified in chain-of-thought and shown to generalise to held-out tasks, alongside Varied Deception, a prompted-lying testbed covering a broad range of lie-inducing motivations. On these testbeds we evaluate four detectors: a chain-of-thought judge, a logprob classifier, and two activation probes, including Did-You-Lie (DYL), a new method for training follow-up probes. On prompted lying, across 31 open-weight models spanning 2B to 1T parameters, all four detectors show positive scaling with model capability. However, every activation- and logprob-based detector drops sharply on our trained model organisms, with DYL retaining the most signal; only the chain-of-thought judge remains strong, with 0.82 balanced accuracy, partly as an artefact of our verification process favouring CoT-readable beliefs. Current lie detectors therefore cannot support high-confidence claims about model beliefs, and we suggest research directions that may address some of their current limitations. We release our datasets, model organisms, and trained detectors.

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Source: arXiv cs.AI Recent

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Computer Science > Artificial Intelligence

Title:"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

More in this category

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models

DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions

Most read

Discover All Categories