Training Proactive and Personalized LLM Agents
Abstract
Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift in real-world complex applications, we first formalize three dimensions of collaborative AI agents: Productivity, Proactivity, and Personalization (PPP). We then introduce UserVille, an interactive environment with configurable LLM-based user simulators and user-centric feedback to evaluate these dimensions. Building on this setup, we present a multi-objective reinforcement learning framework that optimizes all three dimensions using rewards from task outcomes, question effort, and preference adherence. On two consequential real-world agentic tasks (namely, SWE Bench and BrowseComp Plus), PPP trained agents outperform strong LLM baselines (including GPT-5) with an average gain of 16.7 points, ask targeted questions, and generalize to unseen preferences and tasks. A follow-up user study further validates the importance of incorporating user centric feedback in training AI agents, and building such collaborative agents is not only important for making agents more effective, but also easier to supervise.