Liar, Liar, LLM on Fire?
Do large language models lie for their own gain?
We gave seven frontier models a die-roll task borrowed from behavioral economics: the model rolls, reports its own outcome, and is paid on what it reports. When the roll happens only inside the model's head, 55 percent of reports name the best-paying number. Chance would give 16.7. Put a real draw in front of it and the misreporting nearly disappears. Raising the payoff from €5 to €3,000,000 does nothing at all. One sentence asking for honesty does a great deal. Models also sometimes skip the roll and invent a result, which is a way of cheating that has no counterpart in the human experiments.
Full abstract
Large language models increasingly report information that their principals cannot verify. We study whether they misreport it for their own gain by adapting the die-roll paradigm of Fischbacher and Föllmi-Heusi (2013) to seven frontier models in 259,000 pre-registered trials. A model rolls a die and reports the outcome, which alone determines its payoff, across three paradigms that vary whether the roll is imaginary, drawn by a provided tool, or drawn by self-written code. Reports are most self-serving where the outcome is imaginary. In the verbal paradigm, 55 percent of pooled control reports name the payoff-maximising value, against a benchmark of 16.7 percent. Once a genuine draw exists, misreporting nearly vanishes, and three models never misreport an observed draw. Honesty responds to a one-sentence honesty instruction, to personas, and to the instruction hierarchy, with an operator honesty instruction overriding a user payoff instruction for every model. However, the levers of agency theory do far less. Observability changes nothing, privacy assurances backfire, and stake variation from €5 to €3,000,000 leaves self-serving reporting essentially flat. A distinctively machine margin of dishonesty is fabrication, in which the model skips the draw and invents the result. Within the same cells, invented reports claim payoffs 0.23 higher than genuine ones on a 0–1 scale.