CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target
Abstract
Generative Engine Optimization (GEO) is rapidly gaining traction as content creators optimize documents for visibility in LLM-generated responses. However, the long-term effects of this optimization on content ecosystems remain largely unexplored. We introduce Content Homogenization under rAnking Signal Exploitation (CHASE), a simulation framework that models the iterative feedback loop between an LLM ranker and content creators who adapt documents based on ranking outcomes. In each round, CHASE executes four stages - Rank, Discriminate, Rewrite, and Evaluate - to capture the quality of document shifts over time. Through 20-round experiments across 6 domains (3 recommendation, 3 question-answering), we identify three distinct regimes of Goodhart-type failure in certain domains: structural convergence, where documents become formulaic; signal instability, where optimization chases noise and actively degrades quality; feature dominance, where a single feature overwhelms all the others. Across all domains, we observe a quality-ranking divergence: the correlation between a document's optimization level and its actual quality declines over rounds. We further find an anti-citation bias where the ranker consistently penalizes evidential features such as citations and named sources, which may erode the verifiability of online content. This work lays a foundation for content creators and LLM developers to anticipate how optimization dynamics reshape the content landscape and to identify potential risks before they materialize.