Sci-VBench: Evaluating Knowledge- and Reasoning- Intensive Video Generation in Science Domains
Abstract
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge-grounded scientific reasoning in video generation. It contains 1,500 expert-annotated examples spanning 78 scientific domains across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluaters and judge systems are able to achieve high agreement with expert judgments, enabling reliable and reproducible evaluation. We conduct a comprehensive evaluation of 11 models frontier proprietary and open-source models. Our experiment results reveal a substantial gap between perceptual realism and scientific validity: even visually coherent videos frequently violate core domain constraints or drift from the specified setup.