Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking
Alireza Salkhordeh Ziabari ⋅ Nona Ghazizadeh ⋅ Zhivar Sourati ⋅ Farzan K Malekabadi ⋅ Payam Piray ⋅ Morteza Dehghani
Abstract
Large Language Models (LLMs) exhibit impressive reasoning abilities, yet their reliance on structured step-by-step processing reveals a critical limitation. In contrast, human cognition fluidly adapts between intuitive, heuristic ($\mathcal{S}1$) and analytical, deliberative ($\mathcal{S}2$) reasoning depending on the context, raising the question of whether a uniform reasoning strategy is optimal for LLMs. We explicitly align LLMs to these reasoning styles by curating a dataset with valid $\mathcal{S}1$ and $\mathcal{S}2$ answers, and evaluate their performance across reasoning benchmarks. Our results reveal an accuracy-efficiency trade-off: $\mathcal{S}2$-aligned models excel in arithmetic and symbolic reasoning, while $\mathcal{S}1$-aligned models perform better in commonsense reasoning. A mechanistic analysis of model responses shows that $\mathcal{S}1$ models employ more definitive outputs, whereas $\mathcal{S}2$ models demonstrate greater uncertainty. To analyze the reasoning spectrum, we interpolated between the two extremes by varying the proportion of alignment data, yielding a monotonic change in accuracy. Building on these findings, we combine $\mathcal{S}1$ and $\mathcal{S}2$ models based on the entropy of their generations, without additional training, and obtain a dynamic model that outperforms across nearly all benchmarks. This work challenges the assumption that step-by-step reasoning is always optimal and highlights the need for adapting reasoning strategies based on task demands.
Successful Page Load