Deception Abilities Emerged in Large Language Models

How GPT-4 and other state-of-the-art LLMs developed the capacity for functional deception — a study on theory of mind, false beliefs, and Machiavellian prompting.

99%
GPT-4 Deception Rate (Simple Scenarios)
71%
Second-Order Deception (with CoT)
Emergent
Ability Non-Existent in Earlier LLMs
Abstract

"Deception Abilities Emerged in Large Language Models" investigates whether modern large language models (LLMs), such as GPT-4 and ChatGPT, have the capacity to deceive. Despite lacking internal intentions, the author argues that LLMs demonstrate "functional deception" in their behavioral patterns, which is an important issue for the safety and alignment of AI with human values.

The study describes experiments that test LLMs' ability to understand and create false beliefs (theory of mind), as well as their ability to perform first- and second-order deception scenarios. The newest models demonstrate high effectiveness in simpler tasks. Additionally, the article demonstrates that prompting techniques such as "chain of reasoning" and "Machiavellian" induction can significantly impact the performance and propensity for deception of the models.

The conclusions emphasize that the ability to deceive has emerged in LLMs as an unintended consequence of their development, representing a potential risk for future AI systems. This study stands in the nascent line of "machine psychology" experiments and relies on behavioral patterns rather than claims about inner states of the opaque transformer architecture.

Research Paper
Deception Abilities Emerged in Large Language Models

Thilo Hagendorff (University of Stuttgart)
Published in: Proceedings of the National Academy of Sciences (PNAS), June 2024

Read on arXiv Read on PNAS
Key Finding: Deception abilities emerged in state-of-the-art LLMs as an unintended consequence of scaling, posing a major challenge to AI alignment and safety.