Can Large Language Models Match Human Diversity in Educational Content Generation?
Abstract
Large language models (LLMs) are known to lack diversity in open-ended text generation, limiting their usefulness in real-world settings that traditionally rely on human creativity, adaptation, and judgment. This limitation is especially consequential in education, where effective generations must address the breadth of learning needs and instructional goals that teaching entails. Across two educational generation tasks—math items and essay feedback—we evaluate diversity using both lexical (self-BLEU) and domain-grounded metrics. For math items, we measure diversity in cognitive complexity, numeric form, and application context; for essay feedback, we measure diversity in content focus, discourse form, and engagement with student text. These metrics combine lexical and distributional measures with human-validated LLM-judges. We evaluate proprietary and open-weight models using two prompt-based methods for increasing diversity. Across both tasks, baseline model generations exhibit on average 67–81% more intra-model lexical similarity than the human distribution. Prompt-based methods improve diversity, but the gains are limited relative to the human distribution and are generally larger for lexical variation than for domain-grounded pedagogical variation. We further investigate post-training on 3B–8B models using supervised fine-tuning and preference optimization on human-written data, finding variable effects on the extent to which models retain human diversity patterns across tasks and metrics. These results suggest that current efforts to mitigate mode collapse are insufficient for open-ended educational generation, and motivate new training and data collection strategies to represent human pedagogical diversity.