SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
Abstract
Large language models (LLMs) are increasingly used for long-form scientific question answering, where responses must synthesize evidence across interacting factors. However, they frequently hallucinate, producing unsupported or inconsistent claims. Retrieval-Augmented Generation (RAG) can improve trustworthiness by grounding generation in external sources. In this context, scientific simulators are especially compelling because they can validate quantitative hypotheses and capture evolving dynamics. Yet, simulation-based RAG is non-trivial due to two challenges: how to retrieve from scientific simulators, and how to efficiently verify and update long-form answers. To overcome these challenges, we propose SimulRAG, a simulator-based RAG framework with a generalized retrieval interface that translates between text and simulator parameters/outputs. SimulRAG further introduces claim-level generation with uncertainty estimation and simulator boundary assessment (UE+SBA) to selectively verify and update claims. We also release a long-form scientific QA benchmark spanning climate science, epidemiology, and urban planning, with ground truth verified by simulations and human annotators. Experiments show SimulRAG improves informativeness by 30.4% and factuality by 16.3% over traditional RAG, while UE+SBA enhances claim-level efficiency and quality.