Uniform Information Density-Based Preference Optimization
Abstract
To enable large language models to generate outputs that align with human preferences, optimization algorithms based on the comparison between chosen and rejected output pairs have been developed and been proven effective. Some recent endeavors have utilized certain linguistically-derived assumptions about the preference pairs to guide the design of more precise optimization directions. However, human preferences are complex and hard to measure by simplistic criteria, and it is important to incorporate broader and more sophisticated insights about preferred vs. not-preferred language. In this study, we adopt a well-established psycholinguistic theory, Uniform Information Density, to develop a novel preference optimization algorithm, UIDPO. The basic assumption is that human prefers to produce language in which information is uniformly distributed, because uniformity induces less cognitive processing effort. To implement this idea into the algorithm, we derive a compound metric for information uniformity by combining the global standard deviation and the local second-order difference, and integrate it to the objective of preference optimization. We compare UIDPO with SimPO, a previous state-of-the-art method, obtaining comparable and sometimes better performance on AlpacaEval2, ArenaHard, and MT-Bench. Additional linguistic analysis on complement clauses also shows that UIDPO can better reflect the tendencies of keeping information uniform in natural language than baselines. Our results show that the subtle human preference over the information rate in language is a helpful source for achieving better human alignment for LLMs.