ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Abstract
LLM agents are increasingly embedded in productivity and enterprise domains. Agents such as OpenClaw operate across email, calendars, documents, and messaging, but evaluating them on live services risks irreversible errors and lacks reproducibility. We introduce smolclaws, a framework providing five high-fidelity simulated services (Gmail, Google Calendar, Google Docs, Google Drive, and Slack) with full state management and deterministic replay. We design 40+ structured tasks spanning multi-step workflows, cross-service coordination, and safety-critical scenarios. We further propose a data-driven skill improvement loop: collect agent trajectories at scale, extract failure patterns, refine skill specifications, and verify improvements reproducibly. Across 1,500+ evaluation trials, frontier agents solve only 30% of tasks, with failures dominated by incorrect API usage and safety violations. Systematic trajectory analysis and skill refinement yield +35% improvement in average reward, demonstrating that skill quality is a practical proxy for model improvement. Smolclaws and its composable task framework democratize the creation of RL environments for improving agent skills and model training.