The 2nd Workshop on Lifelong Agent: Learning, Aligning, Evolving
Artificial intelligence is entering a new phase: from one-shot assistants to persistent agents that remember, act, and adapt across extended interactions. Recent work increasingly studies agents as stateful systems with memory, planning, tool use, and environment interaction, rather than static models evaluated only on short-horizon benchmarks. This shift makes the central challenge more concrete: how can agents continuously improve while remaining reliable, efficient, and aligned over time?
The notion of a lifelong agent offers a natural lens for this challenge. A lifelong agent should not only acquire new knowledge and skills, but also manage memory, personalize safely, interact with evolving tool ecosystems, and withstand long-term deployment without drift or brittle failures. This workshop brings these emerging directions under a unified agenda centered on agents that learn, align, and evolve throughout their lifespan.
As the second edition of the Lifelong Agents workshop, this event builds on the strong community interest from the first edition while moving the conversation toward next-stage questions: how learning changes alignment, how new tools alter reliability, how personalization affects oversight, and how persistent deployment demands new forms of evaluation and governance. By bringing together language agents, reinforcement learning, multimodal and embodied systems, human-AI interaction, evaluation, and AI safety, we aim to shape a coherent roadmap for agents that can be built responsibly and studied rigorously under real-world conditions.
COLM 2026 Workshop on Efficient Reasoning
Posters are allowed to be posted at any time within Poster Session 1 and Poster Session 2
Reasoning is rapidly moving beyond text-only benchmarks into multimodal, spatial, and embodied settings, where models must interpret long contexts, reason over visual scenes, plan actions, use tools, and interact with dynamic environments. This shift is especially visible in code agents, scientific assistants, embodied systems, and multimodal decision-making pipelines, where strong reasoning quality must now coexist with tight constraints on latency, memory, throughput, and serving cost. At the same time, many of the strongest gains in reasoning still come from more test-time compute, longer trajectories, or deeper search, which makes deployment substantially harder. The central question is therefore no longer only how to make models reason better, but how to make them reason efficiently, robustly, and at scale in the wild.
Building on the success of the 1st Workshop on Efficient Reasoning at NeurIPS 2025 (1,000+ attendees, 275 submissions), this second edition at COLM 2026 will bring together researchers and practitioners working across data, algorithms, systems, and real-world applications to tackle the challenges of making reasoning practical under real-world constraints.
COLM Workshop on Context Beyond the Window: Persistent Knowledge in Language Models
COLM 2026 Workshop on long-context modeling, memory, retrieval, and knowledge internalization in language models. October 9, 2026, San Francisco.
Workshop on Agent Behavior
The dominant paradigm in AI agent research is heavily capability-centric. Agents are primarily evaluated on what they achieve, and progress is measured by rigidly closing the gap between observed and desired performance. Though this has undoubtedly driven remarkable advances, it represents an incomplete account of agency that completely sidesteps a critical question: how do agents achieve what they achieve? A system that achieves a goal through opaque, brittle, or socially harmful processes raises deep concerns that traditional test metrics simply do not capture.
We propose a workshop centered on this complementary question, which we refer to as the study of agent behavior. This encompasses the full range of both observable and latent processes governing an agent's actions, including its decision-making strategies, its interaction patterns, its internal representations, and its responses to interventions.
2nd Workshop on Human-Centered Privacy and Security for Language Models
Language models are rapidly becoming embedded in the interaction fabric of our society — as assistants, tutors, collaborators, and autonomous agents across domains from healthcare to software engineering to creative work. Their versatility makes them uniquely powerful but also uniquely hard to secure: privacy and security risks arise not just from model internals but from the rich, open-ended ways people interact with and through these systems.
Workshop on Responsibly Enabling Data for Foundation Models
As foundation models scale, available training data sources have rapidly depleted. However, several forms of valuable data artifacts such as medical records, legal, and financial documents are restricted from use in model training due to their sensitive nature. In addition, the strong reasoning capabilities in current generative models have opened the possibility for highly personalizable AI applications but these remain bottlenecked by limited access to high quality user data. Hence, it is of immense value to responsibly unlock these data sources (for example: using data transformation or constrained training paradigms) or to generate synthetic alternatives. In this workshop, we aim to bring together domain experts in data, privacy, model training, and legal policy, to advance the frontier of responsibly leveraging such sensitive data with foundation models.
Topics of interest include (but are not limited to):
Generative AI for the World: The First Workshop on Globalizing Tasks, Evaluations, and Systems
Generative AI holds an unprecedented promise: for the first time, a single technology may enable communication and mutual understanding across the world’s languages, cultures, and communities. Yet the global deployment of these systems reveals how far we remain from achieving this vision. AI systems built on assumptions reflecting a narrow slice of the world’s languages, literatures, and cultures are being used by everyone, everywhere—and problems are surfacing across text, images, audio, and video (Khanuja et al., 2024). New failure modes are being revealed by novel global usage patterns. Models struggle with tasks and interaction patterns (e.g., code-switching) never anticipated by predominantly monolingual development pipelines (Oh et al., 2026; Tamkin et al., 2024). These usability issues remain invisible to standard evaluation metrics (Oh et al., 2025).
Countless variations of register, context, and cultural expectations across languages and communities exert pressure on systems operating under Western-centric design assumptions. The boundaries of these systems are always discovered not by their developers, but by users—and with the rising popularity of AI, that is the whole population of the world. That users are finding prompting in Chinese reduces token costs, or that switching languages altogether yields better results than carefully engineered prompts, should not be surprising. Bringing together these perspectives, including deployment failures observed by companies, users’ creative workarounds, and researchers’ analyses of deeper structural limitations, is critical to understanding and improving these systems.
This workshop brings together industry practitioners, researchers, and global users at a critical moment to ask: what would it take for AI to truly serve everyone? Tasks, evaluations, and systems themselves must be rethought at a global scale. We advocate for globalization, not localization—not merely adapting systems to fit diverse contexts, but allowing global diversity to shape how they are built from the ground up. Across modalities—including language, vision, and audio—we seek to identify shared patterns and transferable insights across real-world use. What global weaknesses and failure modes remain undiscovered? Which tasks require different evaluation methods across cultures, and where do standard metrics fall short? This workshop creates a space to surface, study, and extend emerging innovations and vulnerabilities across languages, modalities, and communities.
Methods and Opportunities at Small Scale (MOSS)
As the current LLM research community pushes towards massive scale, this workshop takes a complementary approach by focusing on the under-explored small-scale regime, where scale encompasses compute, data, and model size. We ask: to what extent is scale necessary, and how far can we push toward smaller settings while maintaining competitive performance and enabling scientifically meaningful discoveries?
Beyond transferring insights from small- to large-scale settings, we emphasize the intrinsic value of understanding small-scale regimes. Practically, studying small-scale limits can yield more efficient algorithms and system designs. Small-scale settings also enable controlled, systematic experimentation with rapid iterations, thereby facilitating scientific progress.
This workshop aims to highlight the methods, opportunities, and insights enabled by small-scale experimentation, and to foster discussion on how such approaches can broaden participation and accelerate progress in LLM research.
Workshop on Scientific Understanding of Foundation Models
Moving from empirical scaling phenomena toward predictive science for foundation models.
Despite the extraordinary capabilities of modern foundation models, our scientific understanding of these systems remains remarkably shallow. We can observe that scaling works — but we cannot yet predict when capabilities will grow, why certain representations form, or how reasoning behavior arises from training dynamics.
This workshop aims to catalyze a shift from capability demonstration to formal, testable theory. We seek to uncover laws, invariants, and causal structures — and to develop rigorous evaluation methodologies that can make foundation models more controllable, reliable, and interpretable.
By bringing together researchers from theory, empirical ML, interpretability, optimization, evaluation, and scientific methodology, we aim to lay groundwork for a genuine science of foundation models — one built on predictive understanding, not post-hoc narrative.
Actionable Interpretability
The workshop on Actionable Interpretability@COLM2026 aims to foster discussions on leveraging interpretability insights to drive tangible advancements in AI across diverse domains. We welcome contributions that move beyond theoretical analysis, demonstrating concrete improvements in model alignment, robustness, and real-world applications. Additionally, we seek to explore the challenges inherent in translating interpretability research into actionable impact.
Second Tokenization Workshop
Tokenization–the process of converting raw data into discrete units for model input and output–has emerged as a critical component across machine learning domains. Originally central to natural language processing (NLP), tokenization is now equally essential in multimodal learning, computer vision, speech processing, and other areas. Recent research has shown that tokenization strategies significantly impact model utility, efficiency, and generalization, sparking a surge of interest in this foundational topic.
LM4Sci 2.0: The Second Workshop on Language Models for Scientific Discovery
Significant advancements in Large Language Models (LLMs) have spurred interest in using these frontier AI models to assist researchers in various scientific tasks, such as:
Accelerating the research ideation process through AI-assisted brainstorming and concept exploration.
Searching, synthesizing literature reviews and enabling literature-based question-answering.
Using AI for data-driven discovery and complex scientific data analysis.
End-to-end research pipeline including experiment execution, ML engineering, and paper generation.
Non-Autoregressive Language Models for Fast and Flexible Text Generation
A COLM 2026 workshop on non-autoregressive language modeling — diffusion, flow matching, and any-order autoregression. San Francisco, October 9, 2026.
LLM/VLM Deployment Opportunities and Risks in Healthcare
Benchmark performance is a poor proxy for real-world readiness. DAIH brings together machine learning researchers, physicians, health-system leaders, and policy researchers to tackle safety, equity, privacy, regulation, workflow fit, and post-deployment monitoring.
Learning from Situated and Embodied Interaction
Social Sim'26: Fidelity in Applications
In an era where digital interactions increasingly shape our social fabric, LLM-based social simulation offers a powerful lens to understand complex societal dynamics. As the scope and ambition of these simulations expand, however, important methodological and conceptual challenges become increasingly salient. Beyond compelling demonstrations, questions remain around evaluation, robustness, interpretability, and empirical grounding. Now in its second iteration, our workshop invites researchers, practitioners, and thought leaders to explore rigorous and responsible approaches to LLM-based social simulation.
AI Measurement Science: Toward Rigorous AI Evaluation
A one-day workshop on the science of evaluating frontier AI systems: measurement under interaction, strategic optimization, and non-stationarity.
AdvML-Frontiers × CoTMA: From Model Security to Compositional Threats in Multi-Agent AI Systems
The AdvML-Frontiers × CoTMA workshop is a cross-community effort that unifies two complementary workshop brands: AdvML-Frontiers and CoTMA (Compositional Threats in Multi-Agent AI Systems). As foundation models evolve from standalone predictors into reusable assets and interconnected multi-agent systems, the adversarial surface of AI has expanded far beyond traditional input-output robustness. Building on the AdvML-Frontiers theme, the workshop focuses on securing frontier models both as valuable assets and as complex systems. This includes emerging security challenges surrounding provenance, watermarking, fingerprinting, unauthorized distillation, supply-chain attacks, model integrity, reasoning-time intervention, and system-level robustness. Complementing this perspective, CoTMA focuses on compositional and interaction-layer threats in multi-agent AI systems, where vulnerabilities emerge through communication, delegation, shared memory, tool use, and coordination among agents. By bridging these two themes, AdvML-Frontiers × CoTMA aims to advance a broader vision of AI security and safety for interconnected, evolving, and deployable AI ecosystems, while fostering collaboration across adversarial ML, foundation models, systems security, and agentic AI communities.