Creativity and Hallucinations Share Mechanisms in LLMs
Abstract
A potential connection between creativity and hallucinations in large language models has been previously discussed, but their relationship is still not well understood and current evidence is primarily correlational. In this work, we study this question using mechanistic interpretability and targeted internal interventions. We localize components associated with creativity by predicting creativity metrics across established creativity benchmarks. We then intervene on these components and measure the resulting changes in factuality benchmarks. Across settings, we observe a component-level trade-off: selectively attenuating a small subset of creativity-linked components reduces creativity metrics while improving factual performance under fixed decoding. These results provide intervention-based evidence that internal computations supporting creative generation can also contribute to hallucination-prone behavior. We also observe strong alignment between creativity steering directions and hallucination-associated directions across models, providing complementary geometric evidence for a shared representational axis underlying this trade-off. Together, these findings suggest that the creativity–hallucination trade-off is not merely a behavioral tendency but a structural property of how these models internally represent and generate text.