Do Humans and LLMs Diverge in Belief Revision? Evidence from a Bayesian Analysis
Abstract
Large language models are increasingly used as proxies for human cognition, but do they reason like humans when beliefs conflict with new evidence? We study this question through \emph{belief revision}, that is, how agents update their beliefs in light of contradictions. Cognitive psychology has shown that humans favor \emph{explanation-based revision}: they first generate explanations for the conflict, such as identifying conditions under which a general rule admits exceptions, and revise accordingly. In this paper, we develop a Bayesian model that decomposes this process into two orthogonal components: \emph{what} to revise and \emph{how broadly} to generalize to similar instances. We first run three human-subject experiments on simple, everyday abstract and grounded scenarios to evaluate how people revise their beliefs, and use this data to test our model's predictions. We then evaluate three frontier LLMs (GPT, Claude, Gemini) on the same scenarios and show that while LLMs match human rule-targeting rates, they consistently fail to generalize: both humans and LLMs target the directly contradicted beliefs, but humans propagate revisions to related beliefs while LLMs revise only the directly contradicted instance.