FLY-EVAL++: Evidence-Based Evaluation for Safety-Constrained Modeling in Embodied Systems
Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics. Predictions that are numerically close to ground truth may still violate operational constraints, exhibit physically inconsistent dynamics, or fail to produce usable structured outputs. Despite this, existing evaluation protocols largely ignore safety and constraint compliance. We propose FLY-EVAL++, an evidence-based evaluation framework that formalizes assessment as a two-stage process. First, deterministic verification extracts structured evidence on protocol compliance, physical feasibility, and safety constraints. Second, rubric-guided aggregation produces interpretable multi-dimensional judgments based on this evidence. We apply this framework to Flight Trajectory and Attitude Prediction (FTAP) using an extended FLY-BENCH dataset with multi-step prediction tasks. Across 66 LLMs, we find that safety compliance is the most discriminative dimension of model capability. Models with comparable predictive performance can differ by over 28 points in safety scores. We further identify systematic failure patterns, including safety violations under physically plausible predictions and instability in multi-step rollouts, which are not captured by conventional evaluation. These findings suggest that evaluation in safety-critical domains must explicitly account for constraint satisfaction and structured validity. They motivate a shift from accuracy-centric metrics to evidence-driven, multi-dimensional assessment.