FOCUS: Closed-Loop Attention Feedback for Efficient Vision-Language Understanding
Abstract
Processing every video frame through Vision-Language Models (VLMs) adds significant inference cost and limits their ability to comprehend real-time videos. Existing approaches to this problem either sub-sample frames uniformly, ignoring content and missing important events, or prune visual tokens inside the VLM to reduce per-call cost while still invoking the model on every frame. We introduce FOCUS, a three-layer framework that operates before the VLM, reducing the number of invocations entirely and steering each one towards what matters. At the first layer, temporal gating measures inter-frame dissimilarity directly in the VLM's patch embedding space, triggering burst captures only when the scene changes in a semantically meaningful way. At the second layer, patch-level dissimilarity signals are accumulated across burst frames and fused with a top-down attention prior to produce a spatial focus map that encodes where both motion and semantic relevance coincide. At the third layer, this focus map is used to calibrate the VLM's value cache during generation; vision tokens in task-relevant regions are progressively amplified, steering the model's output without discarding any visual context. The attention feedback loop closes after each response: text-to-vision cross-attention is extracted from the decoder and fed back as the prior for the next burst, allowing the system to converge on semantically important regions over successive observations. All three layers operate in the same embedding space and are linked by the attention feedback signal, keeping gating, spatial focus, and generation steering aligned with the downstream task. On egocentric video captioning across four domains, FOCUS reduces VLM invocations by ~57% (compared to uniform sampling) while maintaining caption quality. On the EgoSchema long-form video QA benchmark subset, FOCUS achieves 73.20% accuracy operating in its native streaming mode.