Can Coding Agents Reproduce Findings in Computational Materials Science?
Abstract
Large language model agents are increasingly marketed as autonomous coding agents, as evidenced by their strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only coding and execution but also navigating complex domain-specific procedures and interpreting results in the context of scientific claims. To address this question, we present RepliCan, a benchmark for evaluating LLM-based agents' ability to reproduce computational materials science claims. The benchmark poses three compounding challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject-matter experts, we curate a set of claims to test whether coding agents can recover and execute the end-to-end workflow needed to support or undermine the claim's credibility. We then evaluate multiple representative coding-agent settings across several foundation models. Our results show that current LLM-based agents achieve low overall success rates on RepliCan, with the best-performing setting achieving a success rate of 54.2%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position RepliCan as both a benchmark for computational scientific reproducibility and a diagnostic testbed for understanding the current limitations of agentic systems in AI-for-science settings.