REFRAMED: Towards Realistic Audio Description Generation for Movies
Abstract
Audio Description (AD) provides verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps between dialogue and must convey only the events needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation for movies in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,052 movie scenes from 207 films, with professional AD transcripts (both US and UK), professional subtitles, and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that, while models outperform trivial baselines, they fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on accessible video understanding.