Skip to yearly menu bar Skip to main content


Oral

Source-Modality Monitoring in Vision-Language Models

Etha Tianze Hua ⋅ Tian Yun ⋅ Ellie Pavlick
Oct 6, 10:00 AM - 10:15 AM Grand Ballroom
We define and investigate *source-modality monitoring* -- the ability of multimodal models to track and communicate the input source from which pieces of information originate. We consider source-modality monitoring as an instance of the more general *binding problem*, and evaluate the extent to which models exploit syntactic vs. semantic signals in order to bind words like *image* in a user-provided prompt to specific components of their input and context (i.e., actual images). Across experiments spanning 11 vision-language models (VLMs) performing target-modality information retrieval tasks, we find that both syntactic and semantic signals play an important role, but that the latter tend to outweigh the former in cases when modalities are highly distinct distributionally. We discuss the implications of these findings for model robustness, and in the context of increasingly multimodal agentic systems.
Show more
View full details
Oral

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

Yijia Shao ⋅ Zora Wang ⋅ Neel Ahuja ⋅ Yicheng Wang ⋅ Bowen Liu ⋅ Diyi Yang
Oct 6, 10:15 AM - 10:30 AM Grand Ballroom
AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working sessions contributed by 93 human workers, our analysis yields insights on two fronts: on the agent side, rankings on CollabSkill diverge meaningfully from those of existing fully autonomous benchmarks where Codex leads, with Claude Code ranking first; on the human side, CollabSkill reveals that practical experience emerges as the primary driver of collaboration skill, with hands-on collaboration meaningfully shifting workers' AI literacy. Together, we hope CollabSkill enables the community to invest in systematic evaluation of human-agent collaboration and spurs development efforts aimed at building AI agents that genuinely augment human workers.
Show more
View full details
Oral

Reasoning about Intent for Ambiguous Requests

Irina Saparina ⋅ Mirella Lapata
Oct 6, 10:30 AM - 10:45 AM Grand Ballroom
Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong. We propose generating a single structured response that enumerates the different ways an ambiguous request can be interpreted, each coupled with a corresponding answer. Our models are trained with reinforcement learning using a dual reward objective: recall on ambiguous inputs to maximise coverage of valid interpretations, and precision on unambiguous ones to suppress spurious alternatives. Training requires only multiple valid answers per input as supervision, no clarification questions or explicit interpretations are needed. Experiments on conversational question answering and semantic parsing demonstrate that our method achieves higher coverage of valid answers than baseline approaches. Human evaluation confirms that predicted interpretations are meaningful and explain their corresponding answers. Our approach promotes transparency with explicit interpretations, achieves efficiency by requiring only one generation step, and supports downstream applications through its structured output format.
Show more
View full details
Oral

Extracting memorized pieces of (copyrighted) books from open-weight language models

A. Feder Cooper ⋅ Mark Lemley ⋅ Allison Casasola ⋅ Ahmed M Ahmed ⋅ Aaron Gokaslan ⋅ Amy B Cyphert ⋅ Christopher De Sa ⋅ Daniel Ho ⋅ Percy Liang
Oct 6, 10:45 AM - 11:00 AM Grand Ballroom
Plaintiffs and defendants in copyright lawsuits make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected expression from books in their training data. We show that these polarized positions dramatically oversimplify the relationship between memorization and copyright. To do so, we develop a technique to measure memorization of books, which we apply to 200 books and 14 open-weight LLMs. Through over 3000 experiments, we show that memorization varies both by model and book. While we find that most LLMs do not memorize most books either in whole or in part, there are notable exceptions. Llama 3.1 70B entirely memorizes some books, like Harry Potter and the Sorcerer's Stone; memorization is so extensive that one can deterministically extract the whole book almost verbatim using the book's first few words as an initial prompt. We discuss why our results have significant implications for copyright cases, though not ones that unambiguously favor either side.
Show more
View full details
Oral

A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Large Language Model Training

Zihan Qiu ⋅ Zeyu Huang ⋅ Kaiyue Wen ⋅ Peng Jin ⋅ Bo Zheng ⋅ Yuxin Zhou ⋅ Haofeng Huang ⋅ Zekun Wang ⋅ Xiao Li ⋅ Huaqing Zhang ⋅ Yang Xu ⋅ Haoran Lian ⋅ Siqi Zhang ⋅ Rui Men ⋅ jianwei zhang ⋅ Ivan Titov ⋅ Dayiheng Liu ⋅ Jingren Zhou ⋅ Junyang Lin
Oct 6, 3:30 PM - 3:45 PM Grand Ballroom
We investigate the functional role of emergent outliers in large language models, specifically attention sinks (a few tokens that consistently receive large attention logits) and residual sinks (a few fixed dimensions with persistently large activations across most tokens). We hypothesize that these outliers, in conjunction with the corresponding normalizations (\textit{e.g.}, softmax attention and RMSNorm), effectively rescale other non-outlier components. We term this phenomenon \textit{outlier-driven rescaling} and validate this hypothesis across different model architectures and training token counts. This view unifies the origin and mitigation of both sink types. Our main conclusions and observations include: (1) Outliers function jointly with normalization: removing normalization eliminates the corresponding outliers but degrades training stability and performance; directly clipping outliers while retaining normalization leads to degradation, indicating that outlier-driven rescaling contributes to training stability. (2) Outliers serve more as rescale factors rather than contributors, as the final contributions of attention and residual sinks are significantly smaller than those of non-outliers. (3) Outliers can be absorbed into parameters or reduced by explicit gated rescaling, which in turn yield better training and quantization performance.
Show more
View full details
Oral

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

Yuqing Yang ⋅ Robin Jia
Oct 6, 3:45 PM - 4:00 PM Grand Ballroom
We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in their previously generated false assertions. Using model-specific testbeds, we find that while LLMs are capable of retraction, they do so only rarely, even when they can recognize their mistakes when asked in a separate interaction. We identify a reliable predictor of retraction: the model's \emph{momentary belief}, as measured by a linear probe on its internal representation. The probe is trained to predict the correctness of answers on external datasets unrelated to retraction, then applied to settings where models should retract. A model retracts only when it ``believes'' its answers to be incorrect \emph{during generation}; these beliefs frequently diverge from models' parametric knowledge as measured by factoid questions. Steering experiments further demonstrate that model belief causally drives retraction. In particular, when the model believes its answer to be incorrect, this not only encourages the model to attempt further verification, but also alters attention dynamics to promote retraction. Finally, we show that supervised fine-tuning re-uses this existing mechanism linking belief with retraction, and primarily improves retraction performance by helping the model learn more accurate internal beliefs.
Show more
View full details
Oral

Phonological Perception of Sign Language Models

Kayo Yin ⋅ Jessica Carter ⋅ Alex X Lu ⋅ Annemarie Kocab
Oct 6, 4:00 PM - 4:15 PM Grand Ballroom
Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates the phonological perception of SLR models by probing phonological sensitivity using minimal pairs and evaluating representational alignment with human behavioral data. Our results reveal that SLR models exhibit emergent phonological sensitivity, but with clear architectural trade-offs: pose-based models are sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlates with human perceptual similarity judgments (r ≈ 0.51). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to scale them beyond their architectural inductive biases.
Show more
View full details
Oral

Attribution Bias in Large Language Models

Eliza H Berman ⋅ Bella Chang ⋅ Daniel B Neill ⋅ Emily Black
Oct 6, 4:15 PM - 4:30 PM Grand Ballroom
As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce \textsc{AttriBench}, the first fame- and demographically-balanced quote attribution benchmark dataset. Through explicitly balancing author fame and demographics, \textsc{AttriBench} enables controlled investigation of demographic bias in quote attribution. Using this dataset, we evaluate 11 widely used LLMs across different prompt settings and find that quote attribution remains a challenging task even for frontier models. We observe large and systematic disparities in attribution accuracy between race, gender, and intersectional groups. We further introduce and investigate \textit{suppression}, a distinct failure mode in which models omit attribution entirely, even when the model has access to authorship information. We find that suppression is widespread and unevenly distributed across demographic groups, revealing systematic biases not captured by standard accuracy metrics. Our results position quote attribution as a benchmark for representational fairness in LLMs.
Show more
View full details
Oral

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

Qihan Ren ⋅ Wang Peng ⋅ Ruikun Cai ⋅ Shuai Shao ⋅ Dadi Guo ⋅ Yuejin Xie ⋅ Yafu Li ⋅ Quanshi Zhang ⋅ Xia Hu ⋅ Jing Shao ⋅ Dongrui Liu
Oct 7, 10:00 AM - 10:15 AM Grand Ballroom
A prevailing narrative in LLM post-training holds that supervised fine-tuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and find that cross-domain generalization is not absent but conditional, jointly shaped by optimization dynamics, training data, and base-model capability. Some reported failures are under-optimization artifacts: cross-domain performance first degrades before recovering and improving with extended training (a dip-and-recovery pattern), so short-training checkpoints can underestimate generalization. Data quality and structure both matter: low-quality solutions broadly hurt generalization, while verified long-CoT traces yield consistent cross-domain gains. Model capability is essential: stronger models internalize transferable procedural patterns (e.g., backtracking) even from a toy arithmetic game, while weaker ones imitate surface verbosity. This generalization is asymmetric, however: reasoning improves while safety degrades, reframing the question from whether reasoning SFT generalizes to under what conditions and at what cost.
Show more
View full details
Oral

What Do Language Models Learn and When? The Implicit Curriculum Hypothesis

Emmy Liu ⋅ Kaiser Sun ⋅ Millicent Li ⋅ Isabelle Lee ⋅ Lindia Tjuatja ⋅ Jen-tse Huang ⋅ Graham Neubig
Oct 7, 10:15 AM - 10:30 AM Grand Ballroom
Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining remain poorly understood. Scaling laws on validation loss tell us how much a model improves with additional compute, but not what skills it acquires in which order. To remedy this, we propose the **Implicit Curriculum Hypothesis**: pretraining follows a compositional and predictable curriculum across models and data mixtures. We test this by designing a suite of simple, composable tasks spanning retrieval, morphological transformations, coreference, logical reasoning, and mathematics. Using these tasks, we track emergence points across four model families spanning sizes from 410M-13B parameters. We find that **emergence orderings** of when models reach fixed accuracy thresholds are strikingly consistent ($\rho = .81$ across 45 model pairs), and that composite tasks most often emerge after their component tasks. Furthermore, we find that this structure is encoded in model representations: tasks with similar function vector representations also tend to follow similar trajectories in training. By using the space of representations derived from our task set, we can effectively predict the training trajectories of simple held-out compositional tasks throughout the course of pretraining ($R^2 = .68$--$.84$ across models) without previously evaluating them. Together, these results suggest that pretraining is more structured than loss curves reveal: skills emerge in a compositional order that is consistent across models and readable from their internals.
Show more
View full details
Oral

BugScope: Learn to Detect Bugs Like Human

Jinyao Guo ⋅ Chengpeng Wang ⋅ Dominic DeLuca ⋅ Jinjie Liu ⋅ Zhuo Zhang ⋅ Xiangyu Zhang
Oct 7, 10:30 AM - 10:45 AM Grand Ballroom
Software auditing is an increasingly critical task in the era of rapid code generation. While LLM-based auditors have demonstrated strong potential, their effectiveness remains limited by misalignment with the highly complex, domain-specific nature of bug detection. In this work, we introduce BugScope, a framework that mirrors how human auditors learn specific bug patterns from representative examples and apply this knowledge during code auditing. BugScope structures auditing into three steps: seed identification, context retrieval, and bug detection, and aligns LLMs to each step by analyzing real bug reports and mutated examples, and distilling concise, reusable guidelines. On a curated dataset of 33 real-world bugs from 21 widely used open-source projects, BugScope achieves 86.05\% precision and 87.88\% recall, corresponding to an F1 score of 0.87. By comparison, leading industrial tools such as Claude Code (with Claude Opus 4.6) and Cursor BugBot achieve F1 scores of only 0.51 and 0.43, respectively. Beyond benchmarks, large-scale evaluation on real-world projects such as the Linux kernel uncovered 184 previously unknown bugs, of which 78 have already been fixed and 7 explicitly confirmed by developers. Our code is available at (https://anonymous.4open.science/r/BugScope-40C4).
Show more
View full details
Oral

Do Humans and LLMs Diverge in Belief Revision? Evidence from a Bayesian Analysis

Stylianos Loukas Vasileiou ⋅ Antonio Rago ⋅ Maria Vanina Martinez ⋅ William Yeoh
Oct 7, 10:45 AM - 11:00 AM Grand Ballroom
Large language models are increasingly used as proxies for human cognition, but do they reason like humans when beliefs conflict with new evidence? We study this question through \emph{belief revision}, that is, how agents update their beliefs in light of contradictions. Cognitive psychology has shown that humans favor \emph{explanation-based revision}: they first generate explanations for the conflict, such as identifying conditions under which a general rule admits exceptions, and revise accordingly. In this paper, we develop a Bayesian model that decomposes this process into two orthogonal components: \emph{what} to revise and \emph{how broadly} to generalize to similar instances. We first run three human-subject experiments on simple, everyday abstract and grounded scenarios to evaluate how people revise their beliefs, and use this data to test our model's predictions. We then evaluate three frontier LLMs (GPT, Claude, Gemini) on the same scenarios and show that while LLMs match human rule-targeting rates, they consistently fail to generalize: both humans and LLMs target the directly contradicted beliefs, but humans propagate revisions to related beliefs while LLMs revise only the directly contradicted instance.
Show more
View full details
Oral

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

Joel Niklaus ⋅ Atsuki Yamaguchi ⋅ Michal Štefánik ⋅ Guilherme Penedo ⋅ Hynek Kydlíček ⋅ Elie Bakouch ⋅ Lewis Tunstall ⋅ Edward E Beeching ⋅ Thibaud Frere ⋅ Colin Raffel ⋅ Leandro von Werra ⋅ Thomas Wolf
Oct 7, 3:30 PM - 3:45 PM Grand Ballroom
Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent. We conduct extensive controlled experiments, generating over one trillion tokens, to identify critical factors in rephrasing web text into synthetic pretraining data. Our results reveal that structured output formats, such as tables, math problems, FAQs, and tutorials, consistently outperform both curated web baselines and prior synthetic methods. Notably, increasing the size of the generator model beyond 1B parameters provides no additional benefit. Our analysis also demonstrates that the selection of the original data used for mixing substantially influences performance. By applying our findings, we develop FinePhrase, a 486-billion-token open dataset of rephrased web text. We show that FinePhrase outperforms all existing synthetic data baselines while reducing generation costs by up to 30 times. We provide the dataset, all prompts, and the generation framework to the research community.
Show more
View full details
Oral

Beyond Distribution Sharpening: The Importance of Task Rewards

Sarthak Mittal ⋅ Leo Gagnon ⋅ Guillaume Lajoie
Oct 7, 3:45 PM - 4:00 PM Grand Ballroom
Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their training pipelines, enabling systems to evolve from pure reasoning models into sophisticated agents. However, debate persists regarding whether RL genuinely instills new skills within a base model or merely sharpens its existing distribution to elicit latent capabilities. To address this dichotomy, we present an explicit comparison between distribution sharpening and task-reward-based learning, utilizing RL as a tool to implement both paradigms. Our analysis reveals the inherent limitations of distribution sharpening, demonstrating from first principles how and why the optima can be unfavorable and the approach fundamentally unstable. Furthermore, our experiments using Llama-3.2-3B-Instruct, Qwen2.5-3B-Instruct and Qwen3-4B-Instruct-2507 on math datasets confirm that sharpening yields limited gains, whereas incorporating task-based reward signal can greatly help achieve robust performance improvements and stable learning.
Show more
View full details
Oral

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

Carter Blum ⋅ Katja Filippova ⋅ Ann Yuan ⋅ Asma Ghandeharioun ⋅ Julian Zimmert ⋅ Fred Zhang ⋅ Jessica Hoffmann ⋅ Tal Linzen ⋅ Martin Wattenberg ⋅ Lucas Dixon ⋅ Mor Geva
Oct 7, 4:00 PM - 4:15 PM Grand Ballroom
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to study the causes and training dynamics of this phenomenon by training small Transformer models from scratch on synthetic multilingual datasets. Depending on (1) the correlation between the facts and the language they were learned in (informativeness), and (2) the ease of language identification (extractability), models either develop unified representations across languages or separate representations; only when the representations are unified do facts transfer across languages. Based on these insights, we propose a unifying perspective which explains a range of prior observations concerning cross-lingual transfer in multilingual LLMs. Our work shows how controlled settings can shed light on pre-training dynamics and suggests methods to encourage representational unification as part of training that would improve cross-lingual transfer in LLMs.
Show more
View full details
Oral

Distributed Attacks in Persistent-State AI Control

Josh Hills ⋅ Ida Caspary ⋅ Asa Cooper Stickland
Oct 7, 4:15 PM - 4:30 PM Grand Ballroom
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute infrastructure across pull requests (PRs) and time its payload for the PR with the best natural cover. To study the resulting dynamics, we introduce Iterative VibeCoding, a control setting where an AI coding agent builds CLI tools over a sequence of PRs in a persistent codebase while pursuing a covert side task. In our main experiments, we use Claude Sonnet 4.5 as the attack agent and GPT-4o as the monitor. We compare gradual attacks, which distribute the side task across PRs, against non-gradual attacks concentrated in a single PR. No single monitor catches both strategies: peak evasion (success while evading the monitor) on a single monitor reaches 94%, where the optimal attack strategy depends on the monitor type and structure of the side task. High evasion (≥65%) generalizes across model attack agent backends (Claude Sonnet 4.5, Gemini 3.1 Pro, Kimi K2.5), confirming this is a property of the persistent-state attack surface rather than a single model's capability. Evasion also remains high across state-of-the-art monitor models; the gap between gradual and non-gradual evasion widens for more capable models. A new stateful monitor that tracks suspicious buildup across PRs substantially outperforms simply showing more raw history, reducing gradual-attack evasion to 65%. Combining this stronger diff monitor with trajectory monitors in a four-monitor ensemble reduces gradual-attack evasion to 59%.
Show more
View full details
Oral

More Than Words: Compositional Tokenization for Efficient Language Models

Yuval Reif ⋅ Guy Kaplan ⋅ Roy Schwartz
Oct 8, 10:00 AM - 10:15 AM Grand Ballroom
Language models process and generate text sequentially in *token* units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as “On the table.” is usually produced as four separate predictions for the preposition (*on*), article (*the*), noun (*table*), and period (*.*), affecting the overall context length, and accordingly—inference costs. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a single lexical base token (*table*) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled from-scratch pretraining at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves downstream performance by 1.2 points relative to standard BPE under matched training compute. More broadly, our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.
Show more
View full details
Oral

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

Xinyue Liu ⋅ Niloofar Mireshghallah ⋅ Jane C Ginsburg ⋅ Tuhin Chakrabarty
Oct 8, 10:15 AM - 10:30 AM Grand Ballroom
Frontier LLM companies have repeatedly assured courts and regulators that their models do not store copies of training data. They further rely on safety alignment strategies via RLHF, system prompts, and output filters to block verbatim regurgitation of copyrighted works, and have cited the efficacy of these measures in their legal defenses against copyright infringement claims. We show that finetuning bypasses these protections: by training models to expand plot summaries into full text, a task naturally suited for commercial writing assistants, we cause GPT-4o, Gemini-2.5-Pro, and DeepSeek-V3.1 to reproduce up to 85-90\% of held-out copyrighted books, with single verbatim spans exceeding 460 words, using only semantic descriptions as prompts and no actual book text. This extraction generalizes across authors: finetuning exclusively on Haruki Murakami's novels unlocks verbatim recall of copyrighted books from over 30 unrelated authors. The effect is not specific to any training author or corpus: random author pairs and public-domain finetuning data produce comparable extraction, while finetuning on synthetic text yields near-zero extraction, indicating that finetuning on individual authors' works reactivates latent memorization from pretraining. Three models from different providers memorize the same books in the same regions ($r \ge 0.90$), pointing to an industry-wide vulnerability. Our findings offer compelling evidence that model weights store copies of copyrighted works and that the security failures that manifest after finetuning on individual authors' works undermine a key premise of recent fair use rulings, where courts have conditioned favorable outcomes on the adequacy of measures preventing reproduction of protected expression.
Show more
View full details
Oral

Message Passing Enables Efficient Reasoning

Xuecheng Liu ⋅ Daman Arora ⋅ Gokul Swamy ⋅ Andrea Zanette
Oct 8, 10:30 AM - 10:45 AM Grand Ballroom
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scaling techniques instead use fork and join (FJ) primitives to divide work across multiple LLM threads. Although conceptually promising, this necessitates a centralized controller, limiting scalability. We introduce Message Passing Language Models (MPLMs), a framework for LLM reasoning in which threads communicate directly via lightweight send and receive primitives. MPLMs enable efficient scaling through two key mechanisms: (1) reduced communication costs, achieved by avoiding redundant context sharing, and (2) preemption, which allows threads to terminate early based on partial information from their peers. In Sudoku puzzles, we show that MPLMs require an asymptotically smaller context than both serial CoT and parallel FJ . Most importantly, we fine-tune a single model to solve 25 $\times$ 25 puzzles that remain challenging for standard CoT and FJ approaches, as well as frontier reasoning models without tools. We also demonstrate that on 3-SAT puzzles, where search dominates, the capability of preemption allows termination of unpromising branches, which results in improved efficiency. Together, these results demonstrate that MPLMs could provide a principled and scalable alternative to existing sequential and centrally coordinated parallel reasoning paradigms.
Show more
View full details
Oral

A Mirage of Coherence: How Metaphor Impacts Language Models' Discourse Coherence Assessment

Wanting Ning ⋅ Brian Choi ⋅ Eun Jin Paek ⋅ Hyeju Jang
Oct 8, 10:45 AM - 11:00 AM Grand Ballroom
Recent approaches to discourse coherence assessment increasingly rely on language models, but it remains unclear whether they evaluate underlying discourse structure or merely surface-level lexical patterns. To investigate, we introduce metaphor as a controlled linguistic probe. Using a rigorous paraphrase framework, we generate meaning-preserving metaphorical rewrites for a corpus of human-annotated texts. Across multiple architectures, we uncover a pervasive tendency: models paradoxically reward figurative language, consistently assigning higher coherence scores to texts with novel metaphors, provided the metaphors remain semantically compatible with their context. Furthermore, layer-wise analyses reveal that while early model layers are highly sensitive to surface-level figurative features, higher layers successfully recover the underlying semantics. Ultimately, our findings demonstrate that coherence judgments are influenced by local stylistic variation, indicating that current models often evaluate linguistic realization rather than relying on discourse-level reasoning alone.
Show more
View full details
Oral

Diversity or Precision? A Deep Dive into Next Token Prediction

Haoyuan Wu ⋅ Hai Wang ⋅ jiajia wu ⋅ Jinxiang Ou ⋅ keyaowang ⋅ Weile Chen ⋅ Zihao Zheng ⋅ Bei Yu
Oct 8, 3:30 PM - 3:45 PM Grand Ballroom
Recent advancements have shown that reinforcement learning (RL) can substantially improve the reasoning abilities of large language models (LLMs). The effectiveness of such RL training, however, depends critically on the exploration space defined by the pre-trained model's token-output distribution. In this paper, we revisit the standard cross-entropy loss, interpreting it as a specific instance of policy gradient optimization applied within a single-step episode. To systematically study how the pre-trained distribution shapes the exploration potential for subsequent RL, we propose a generalized pre-training objective that adapts on-policy RL principles to supervised learning. By framing next-token prediction as a stochastic decision process, we introduce a reward-shaping strategy that explicitly balances diversity and precision. Our method employs a positive reward scaling factor to control probability concentration on ground-truth tokens and a rank-aware mechanism that treats high-ranking and low-ranking negative tokens asymmetrically. This allows us to reshape the pre-trained token-output distribution and investigate how to provide a more favorable exploration space for RL, ultimately enhancing end-to-end reasoning performance. Contrary to the intuition that higher distribution entropy facilitates effective exploration, we find that imposing a precision-oriented prior yields a superior exploration space for RL.
Show more
View full details
Oral

MEMENTO: Teaching LLMs to Manage Their Context

Vasilis Kontonis ⋅ Yuchen Zeng ⋅ Shivam Garg ⋅ Lingjiao Chen ⋅ Hao Tang ⋅ Ziyan Wang ⋅ Ahmed H Awadallah ⋅ Eric Horvitz ⋅ John Langford ⋅ Dimitris Papailiopoulos
Oct 8, 3:45 PM - 4:00 PM Grand Ballroom
Reasoning models think in long, unstructured streams with no mechanism for compressing or organizing their own intermediate state. We introduce MEMENTO: a method that teaches models to segment reasoning into blocks, compress each block into a \emph{memento}, i.e., a dense state summary, and reason forward by attending only to mementos, reducing context, KV cache, and compute. To train Memento models, we release OpenMementos, a public dataset of 228K reasoning traces derived from OpenThoughts-v3, segmented and annotated with intermediate summaries. We show that a two-stage SFT recipe on OpenMementos is effective across different model families (Qwen3, Phi-4, OLMo3) and scales (8B--32B parameters). Trained models maintain strong accuracy on math, science, and coding benchmarks while achieving ${\sim}2.5\times$ peak KV cache reduction. We extend vLLM to support our inference method, achieving up to $2\times$ throughput improvement while also enabling us to perform RL and further improve accuracy. Finally, we identify a \emph{dual information stream}: information from each reasoning block is carried both by the memento text and by the corresponding KV states, which retain implicit information from the original block. Removing this channel drops accuracy by 15pp on AIME24.
Show more
View full details
Oral

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Meera Desai ⋅ Sang Truong ⋅ Hanna Wallach ⋅ Alex Chouldechova ⋅ A. Feder Cooper ⋅ Jean Garcia-Gathright ⋅ Daniel Ho ⋅ Abigail Z Jacobs ⋅ Sanmi Koyejo ⋅ Nicholas Pangakis ⋅ Angelina Wang
Oct 8, 4:00 PM - 4:15 PM Grand Ballroom
Evaluation benchmarks play a central role in the development and governance of AI systems, yet persistent concerns about val concerns about their validity remain. It is unclear that such benchmarks actually measure the concepts they claim to assess. Drawing inspiration from the principles of $\textit{convergent and discriminant validity}$, we analyze 56 safety and capability benchmarks across 53 models using correlation analyses and item-level prediction models. When studying benchmarks purporting to measure the same construct, we often see inconsistent consistent rankings on the models. This suggests that the constructs, such as safety detection and exaggerated refusal, may not be sufficiently well-defined. When comparing across constructs, we find that some concepts are distinguishable from others (e.g., unsafe generation is different from bias), but many concepts (e.g., reasoning, knowledge, and comprehension) are not well separated. In some cases, evaluation format and task structure drive correlations between benchmarks more than the concepts they aim to measure. Finally, we investigate individual benchmarks to identify benchmarks that likely measure a different concept than reported. In particular, we find that BBQ, one of the most widely used bias benchmarks, is more correlated with reasoning benchmarks than with other bias benchmarks. Overall, we encourage benchmark developers and practitioners to evaluate benchmarks through the lens of convergent and discriminant validity when developing and using benchmarks. We release our extensive item-level dataset of model predictions to support future empirical work on benchmark validity.
Show more
View full details
Oral

Introspective Diffusion Language Models

Yifan Yu ⋅ Yuqing Jian ⋅ Junxiong Wang ⋅ Zhongzhu Zhou ⋅ Donglin Zhuang ⋅ Xinyu Fang ⋅ Xiaoxia Wu ⋅ Qingyang Wu ⋅ Shuaiwen L Song ⋅ Tri Dao ⋅ Ben Athiwaratkun ⋅ James Zou ⋅ Fan Lai ⋅ Chenfeng Xu
Oct 8, 4:15 PM - 4:30 PM Grand Ballroom
Diffusion language models (DLMs) offer a compelling promise: parallel token generation could break the sequential bottleneck of autoregressive (AR) decoding. Yet in practice, DLMs consistently lag behind AR models in quality. We argue that this gap stems from a fundamental failure of introspective consistency: AR models agree with what they generate, whereas DLMs often do not. We formalize this via the *introspective acceptance rate*, which quantifies whether a model internally accepts its previously generated tokens. Through this lens, we uncover a key structural advantage of AR models: causal masking combined with logit shifting implicitly enforces introspective consistency during training. Motivated by this insight, we introduce the Introspective Diffusion Language Model (I-DLM), a new paradigm that preserves the introspective consistency of AR training while retaining the diffusion-style parallelism. I-DLM uses a novel introspective strided decoding (ISD) algorithm, which enables the model to verify previously generated tokens while advancing new ones in the same forward pass. This yields a new quality–efficiency frontier unavailable to either AR or prior diffusion models, with stride providing a controllable tradeoff between verification depth and parallel progress. Empirically, I-DLM-8B is the first DLM to match the quality of its same-scale AR counterpart while surpassing all prior DLMs in both quality and practical serving efficiency across 15 benchmarks. It attains 72.5 on AIME-24 and 45.1 on LiveCodeBench-v6, outperforming LLaDA-2.1-mini (16B) by more than 29 and 14 points, respectively. Finally, with a single-pass self-speculative decoding pipeline and gated LoRA, ISD enables lossless acceleration and better efficiency at high concurrency compared to speculative decoding.
Show more
View full details