Are LLMs Robust Enough for Operations Research? An Adversarial Attack and Defense Study via AttackOR
Abstract
Operations research (OR) solves complex real-world decision problems, from supply chain optimization to resource allocation, where errors carry real costs. LLMs are increasingly deployed to solve such problems by translating natural-language descriptions into optimization code. But how reliable are LLMs when problem descriptions are noisy, ambiguous, or deliberately misleading? We propose AttackOR, a unified framework for adversarial evaluation and robustness training in OR. We introduce four attack strategies that generate domain-consistent, instance-specific distractors to simulate the noisy and ambiguous inputs encountered in real-world OR: attribute injection, phantom options, arithmetic equivalence, and unit conversion. Experiments show that even a strong OR-tuned model is highly vulnerable, dropping from 59% to 40% Pass@1 under attack. To improve robustness, we combine supervised fine-tuning with GRPO-based reinforcement learning on adversarially augmented data. Our approach not only defends against adversarial inputs, achieving 71% Pass@1 under attack, but also improves performance on standard OR benchmarks, from 59% to 74%. Findings reveal that clean-benchmark performance alone is insufficient to guarantee the reliability of LLM-based OR solvers in practice.