GLAZE: A Gaze-Language Benchmark for Grounding Human Gaze on Web User Interfaces
Abstract
Eye gaze offers a grounding signal for multimodal language models when understanding user interfaces, such as inferring which element a user intends to interact with. But natural gaze on dense interfaces rarely maps cleanly to element boxes or semantic groups. Users scan across distant regions, revisit earlier elements, and spread attention over large layouts. Existing research either studies UI eye tracking without aligned language supervision or uses gaze to guide models in natural image and egocentric settings. We introduce the first benchmark for gaze-language grounding on web user interfaces. GLAZE pairs more than 10,000 UI elements with eye-tracking traces and aligned verbal descriptions from 27 participants with UI/UX experience. Given a screenshot and a gaze segment, a model must infer the interface content or function that the individual attends to at that moment. We evaluate open and proprietary multimodal language models across multiple gaze representations. Current models often identify the general region being viewed but miss the precise attended elements and their functional role, leaving a substantial gap between natural human gaze and model grounding on real-world UI interfaces. Our evaluation of gaze visualization also reveals which forms of gaze information are most beneficial to models across different UI types. The benchmark supports behavior-grounded evaluation of multimodal language models, with potential applications in GUI agents, gaze-based accessibility tools, and automated usability analysis.