CooperBench: Why coding agents cannot be your teammates yet
Abstract
Resolving team conflicts requires not only task-specific competence, but also social intelligence to find common ground and build consensus. Similarly, as AI agents increasingly collaborate on complex work, they must develop coordination capabilities to function as effective teammates. We hypothesize that current agents lack these capabilities. To test it, we introduce CooperBench, a benchmark of 600 collaborative coding tasks spanning 12 libraries and 4 languages. Each task assigns agents independently implementable features that may conflict without coordination. Tasks are grounded in real open-source repositories with expert-written tests, which makes the cooperation outcomes fully verifiable. Evaluating SOTA coding agents, we observe the curse of coordination: at least 30\% average drop in success when agents work together versus a single agent completing both tasks, especially when the tasks are not extremely easy or hard. We identify three failure modes: (1) communication channels become jammed with vague, ill-timed, and inaccurate messages; (2) even with good communication, agents deviate from their commitments and hold incorrect expectations about others; and (3) agents have troubles using version control tools to perform real-time coordination. Beyond an open-source benchmark, this work contributes a novel understanding of what agents need to learn to become effective teammates.