Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
Jiannan Xu ⋅ Yongkang Duan ⋅ Jane Jiang ⋅ Jiding Zhang
Abstract
As large language models (LLMs) are increasingly deployed as decision-making agents in competitive and strategic environments, their performance depends critically on how they model human counterparts. Yet little is known about the implicit assumptions LLMs make about human rationality. Drawing on the level-$k$ thinking framework from behavioral game theory, we design a suite of normal-form games and introduce a structural estimation procedure that infers an LLM's latent belief about its opponent’s reasoning depth from observed choices. Across models, we find a systematic bias: LLM agents overwhelmingly assume humans to be Nash-type players, i.e., fully rational strategic optimizers---and respond with equilibrium play. However, human subjects in our online experiments exhibit substantial heterogeneity, spanning \textsc{Random} through \textsc{Nash} reasoning. This mismatch has important consequences. In some settings, an LLM that outsmarts a boundedly rational human secures higher payoffs. In others, however, overestimating human sophistication induces what we term a \emph{Nash equilibrium trap}: equilibrium play yields strictly lower joint payoffs than strategies calibrated to actual human behavior. We propose a behaviorally grounded prompting intervention that explicitly accounts for human reasoning heterogeneity. This intervention significantly reduces equilibrium bias and improves payoff outcomes by inducing opponent-aware strategic reasoning. Our results demonstrate that strategic competence alone is insufficient for effective human–AI interaction. Accurate modeling of human rationality is a distinct alignment challenge for LLM agents operating in real-world decision-making settings.
Successful Page Load