Reach Into The Choir: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
Abstract
Recent work documents that frontier language models are converging toward increasingly similar outputs. Most evidence for this, however, relies on single-pass responses, leaving open what models can stably generate when systematically probed at greater depth. This question has direct consequences for ensemble methods, model selection, and LLM-based evaluation. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles, prompting models to generate ranked lists under temperature variation and prompt perturbation to distinguish stable high-frequency responses from rarer outputs that emerge only under sustained elicitation. Applied across seven models and 27 questions spanning five domains (social, memory, welfare, introspective, and meta-inquiry), CHOIR reveals that individual models produce stable, distinctive profiles, with cross-model agreement substantially lower than surface comparisons suggest. Matched question pairs show that apparent consensus often reflects echo of question vocabulary rather than independently convergent content. Conditioning models on distinct synthetic persona profiles substantially increases cross-model agreement on social and welfare questions, but on introspective questions, models retain distinct profiles regardless of conditioning. A blind ranking evaluation, in which two independent sets of LLMs ranked responses without source attribution, confirms that both sets prefer conditioned outputs in most domains, with no persona-level self-preference bias. These results indicate that apparent model homogeneity is in part an artefact of shallow elicitation, and that structured depth probing recovers meaningful diversity beneath it.