TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
Abstract
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. We propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances video temporal understanding. TEMPURA first applies masked event prediction to reconstruct missing events and generate step-by-step causal explanations from dense event annotations, which draws inspiration from infilling techniques for language modeling. TEMPURA then learns to perform video segmentation and dense captioning to decompose videos into non-overlapping events with detailed, timestamp-aligned descriptions. We train TEMPURA on VER, a large-scale dataset curated by us that comprises 1M training instances and 500K videos with temporally aligned event descriptions and structured reasoning steps. Experiments on temporal grounding, highlight detection, and dense video captioning benchmarks demonstrate that TEMPURA outperforms strong baseline models. We further show that TEMPURA's training pipeline is model-agnostic and generalizes to other video Large Multi-modal Models, confirming that incorporating causal reasoning with fine-grained temporal segmentation leads to improved video understanding.