Weekly TMLR digest for Jul 19, 2026

9 views
Skip to first unread message

TMLR

unread,
Jul 19, 2026, 12:00:11 AMJul 19
to tmlr-annou...@googlegroups.com


New certifications
==================

Expert Certification: Self-Supervised Learning via Flow-Guided Neural Operator on Time-Series Data

Duy Nguyen, Jiachen Yao, Jiayun Wang, Julius Berner, Anima Anandkumar

https://openreview.net/forum?id=YAYW9Y173z

---


Survey Certification: Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

Aggelos Psiris, Vasileios Argyriou, Evangelos K. Markakis, Panagiotis Sarigiannidis, Stratis Gavves, Kostas Bekris, Arash Ajoudani, Georgios Th. Papadopoulos

https://openreview.net/forum?id=qF7vdPrpPk

---


Featured Certification, J2C Certification: Achieving Adaptivity and Optimality for Multi-armed Bandits using Exponential-Kullback Leibler Maillard Sampling

Hao Qin, Kwang-Sung Jun, Chicheng Zhang

https://openreview.net/forum?id=IuVkRmecVp

---


Survey Certification: A Survey on Efficient Protein Language Models

Shouren Wang, Debargha Ganguly, Vinooth Rao Kulkarni, Van Yang, Zhuoran Qiao, Daniel Blankenberg, Vipin Chaudhary, Xiaotian Han

https://openreview.net/forum?id=PTReuOwsXz

---


Survey Certification: Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang, Shuaicheng Niu, Taki Hasan Rafi, Jihun Hamm, Marco Pedersoli, Jose Dolz, Yunhui Guo

https://openreview.net/forum?id=mM3r03Xw1V

---


J2C Certification: MatchEx: Model-Level GNN Explanations with Multi-Granular Insights

Sayan Saha, Sanghamitra Bandyopadhyay

https://openreview.net/forum?id=YMETLG2WvM

---


Accepted papers
===============


Title: A Descriptive and Normative Theory of Human Beliefs in RLHF

Authors: Sylee Dandekar, Shripad V. Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum

Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human beliefs about the capabilities of the agent being trained also play a key role in preference generation. We examine two questions related to this hypothesis, one descriptive and one normative, respectively: Do human labelers' beliefs about agent capabilities affect the preferences that they provide? And what is the ideal set of beliefs about an agent---and resulting preferences---for humans to have? We propose a new preference model that incorporates human beliefs and provide a normative theory that bounds the error on the final learned policy based on the _mismatch_ between the human's beliefs and an idealized set of beliefs. We then confirm via a human study that beliefs about agent capabilities do, in fact, significantly affect preferences and can be influenced through simple interventions. Additionally, we empirically show through synthetic experiments that it is often suboptimal for human preference labelers to assume agent optimality. Collectively, these results theoretically and empirically demonstrate how reducing the mismatch between human beliefs and agent capabilities can lead to more performant RLHF and point toward new best practices for RLHF practitioners.

URL: https://openreview.net/forum?id=YdW0KZwPeT

---

Title: M3Ret: Unleashing Zero-shot Multi-Modal Medical Image Retrieval via Self-Supervision

Authors: Che Liu, Zheng Jiang, Chengyu Fang, Heng Guo, Yan-Jie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu

Abstract: Medical image retrieval is essential for clinical decision-making and translational research, relying on discriminative visual representations. Yet, current methods remain fragmented, relying on separate architectures and training strategies for 2D, 3D, and video-based medical data. This modality-specific design hampers scalability and inhibits the development of unified representations.
To enable unified learning, we curate a large-scale hybrid-modality dataset comprising 867,653 medical imaging samples, including 2D X-rays and ultrasounds, RGB endoscopy videos, and 3D CT scans. Leveraging this dataset, we train M3Ret, a unified visual encoder without any modality-specific customization. It successfully learns transferable representations using both generative (MAE) and contrastive (SimDINO) self-supervised learning (SSL) paradigms.
Our approach sets a new state-of-the-art in zero-shot image-to-image retrieval across all individual modalities, surpassing strong baselines such as DINOv3 and the text-supervised BMC-CLIP. More remarkably, strong cross-modal alignment emerges without paired data, and the model generalizes to unseen MRI tasks, despite never observing MRI during pretraining, demonstrating the generalizability of purely visual self-supervision to unseen modalities.
Comprehensive analyses further validate the scalability of our framework across model and data sizes. These findings deliver a promising signal to the medical imaging community, positioning M3Ret as a step toward foundation models for visual SSL in multimodal medical image understanding.

URL: https://openreview.net/forum?id=VgmIrgbzkX

---

Title: NeuMoSync: End‑to‑End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning

Authors: Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei, Mo Chen

Abstract: Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability. Drawing high-level inspiration from global neuromodulatory mechanisms in the brain, we introduce $\textbf{Neu}$ro$\textbf{Mo}$dulation and $\textbf{Sync}$hronization ($\texttt{NeuMoSync}$), a novel architecture that integrates dynamic, neuron-specific modulation into deep neural networks to enhance their adaptability and plasticity. $\texttt{NeuMoSync}$ extends standard neural network architectures with learnable feature vectors per neuron that tracks network-wide historical context and incorporates a module operating at a higher level of abstraction. This module synthesizes neuron-specific signals, conditioned on both current inputs and the network’s evolving state, to adaptively regulate activation dynamics and synaptic plasticity. Evaluated on diverse CL benchmarks, including memorization (Random Label CIFAR-10, Random Label MNIST), concept drift (Shuffle CIFAR-10, Shuffle Mini-ImageNet), class-incremental (Class Split T-ImageNet/ImageNet, Class Split CIFAR-100) and domain-incremental (Permuted MNIST), $\texttt{NeuMoSync}$ demonstrates strong performance in terms of retention of plasticity and achieves improvements in both forward and backward adaptation compared to existing methods. Ablation studies validate the necessity of each component, while the analysis of the learned modulatory signals reveals interpretable coordination patterns across tasks. Our work underscores the potential of integrating global coordination mechanisms into deep learning systems to advance robust, adaptive continual learning.

URL: https://openreview.net/forum?id=i6cb6CGMtY

---

Title: Simplifying Multi-Task Architectures Through Task-Specific Normalization

Authors: Mihai Suteu, Ovidiu Serban

Abstract: Multi-task learning (MTL) aims to leverage shared knowledge across tasks to improve generalization and parameter efficiency, yet balancing resources and mitigating interference remain open challenges. Architectural solutions often introduce elaborate task-specific modules or routing schemes, increasing complexity and overhead. In this work, we show that normalization layers alone are sufficient to address many of these challenges. Simply replacing shared normalization with task-specific variants already yields competitive performance, questioning the need for complex designs. Building on this insight, we propose Task-Specific Sigmoid Batch Normalization (TS$\sigma$BN), a lightweight mechanism that enables tasks to softly allocate network capacity while fully sharing feature extractors. TS$\sigma$BN improves stability across CNNs and Transformers, matching or exceeding performance on NYUv2, Cityscapes, CelebA, and PascalContext, while remaining highly parameter-efficient. Moreover, its learned gates provide a natural framework for analyzing MTL dynamics, offering interpretable insights into capacity allocation, filter specialization, and task relationships. Our findings suggest that complex MTL architectures may be unnecessary and that task-specific normalization offers a simple, interpretable, and efficient alternative. Code is available at https://github.com/xapharius/TSsNorm

URL: https://openreview.net/forum?id=QNO893OXrS

---

Title: A Player Selection Network for Scalable Game-Theoretic Prediction and Planning

Authors: Tianyu Qiu, Eric Ouano, Fernando Palafox, Christian Ellis, David Fridovich-Keil

Abstract: While game-theoretic planning frameworks are effective at modeling multi-agent interactions, they require solving large optimization problems where the number of variables increases with the number of agents, resulting in long computation times that limit their use in large-scale, real-time systems. To address this issue, we propose i) PSN Game—a learning-based, game-theoretic prediction and planning framework that reduces game size by learning a Player Selection Network (PSN); and ii) a Goal Inference Network (GIN) that makes it possible to use the PSN in incomplete-information games where other agents’ intentions are unknown to the ego agent. A PSN outputs a player selection mask that distinguishes influential players from less relevant ones, enabling the ego player to solve a smaller, masked game involving only selected players. By reducing the number of players included in the game, PSN shrinks the corresponding optimization problems, leading to faster solve times. The PSN Game framework is more flexible than existing player selection methods as it i) relies solely on observations of players’ past trajectories, without requiring full state, action, or other game-specific information; and ii) requires no online parameter tuning. Experiments in both simulated scenarios and real-world pedestrian trajectory datasets show that PSN is competitive with, and often improves upon, the evaluated explicit game-theoretic selection baselines in i) prediction accuracy and ii) planning safety. Across scenarios, PSN typically selects substantially fewer players than are present in the full game, thereby reducing game size and planning complexity. PSN also generalizes to settings in which agents’ objectives are unknown, via the GIN, without test-time fine-tuning. By selecting only the most relevant players for decision-making, PSN Game provides a practical mechanism for reducing planning complexity that can be integrated into existing multi-agent planning frameworks.

URL: https://openreview.net/forum?id=YvvB78ILSP

---

Title: Harmonizing Gradient Matching For Fairness

Authors: Ziwei Wu, Yikun Ban, Jingrui He

Abstract: Ensuring fairness across demographic groups is critical for machine learning systems deployed in high-stakes applications. Most existing approaches enforce fairness by directly minimizing disparities in predefined fairness metrics between groups, focusing primarily on the final model outcome. However, differences in distributions across groups can lead to heterogeneous optimization signals during training, resulting in imbalanced parameter updates and unstable fairness–performance trade-offs. In this work, we propose Fair Gradient Matching (FairGM), a fairness-aware optimization framework that harmonizes group-conditioned optimization signals. Instead of focusing solely on fairness metric disparities, FairGM aligns gradient signals of the fairness objective across groups at multiple levels of moments, including the zeroth-moment fairness metric itself, the first-moment mean gradients, and the second-moment gradient variances. These regularizations encourage similar optimization behavior across groups and lead to more stable fairness outcomes. To balance predictive performance and fairness objectives, we further formulate training as a multi-objective optimization problem and solve it using a Pareto-based optimization scheme. The resulting framework is compatible with a range of differentiable fairness metrics and gradient-based classifiers, supported by theoretical analysis connecting gradient alignment for fairness. Experiments on synthetic and real-world datasets demonstrate that FairGM achieves favorable fairness–accuracy trade-offs compared with existing fairness-aware learning methods, demonstrating the effectiveness and scalability of our approach.

URL: https://openreview.net/forum?id=HMlK36YWWt

---

Title: Self-Supervised Learning via Flow-Guided Neural Operator on Time-Series Data

Authors: Duy Nguyen, Jiachen Yao, Jiayun Wang, Julius Berner, Anima Anandkumar

Abstract: Self-supervised learning (SSL) is a powerful paradigm for learning from unlabeled time-series data. However, popular methods such as masked autoencoders (MAEs) rely on reconstructing inputs from a fixed, predetermined masking ratio. Instead of this static design, we propose treating the corruption level as a new degree of freedom for representation learning, enhancing flexibility and performance. To achieve this, we introduce the Flow-Guided Neural Operator (FGNO), a novel framework combining operator learning with flow matching for SSL training. FGNO learns the frequency functions represented through the Short-Time Fourier Transform (STFT). We extract a rich hierarchy of features by tapping into different network layers and flow times that apply varying strengths of noise to the input data. This enables the extraction of versatile representations, from low-level patterns to high-level global features, using a single model adaptable to specific tasks. Unlike prior generative SSL methods that use noisy inputs during inference, we propose using clean inputs for representation extraction while learning representations with noise; this eliminates randomness and boosts accuracy. We evaluate FGNO across three biomedical domains, where it consistently outperforms established baselines. Our method yields up to 39% relative AUROC gains over the MAE baseline in neural signal decoding (BrainTreeBank), 18% RMSE reductions in skin temperature prediction (DREAMT), and over 20% improvement in accuracy and macro-F1 on SleepEDF under low-data regimes. Furthermore, experiments under aggressive downsampling further show robustness to severe bandlimiting, highlighting FGNO's capability to learn in the function space. These results highlight FGNO's robustness to data scarcity and its superior capacity to learn expressive and adjustable representations for time series.

URL: https://openreview.net/forum?id=YAYW9Y173z

---

Title: Tensor-Decomposed RNNs for Marked Temporal Point Processes

Authors: Timothy Mulumba

Abstract: We study parameter-efficient neural Marked Temporal Point Processes (MTPPs) for high-dimensional mark and exogenous feature spaces. Building on tensor-train (TT) factorization of recurrent kernels, we propose mark-aware TT shaping that aligns TT cores with known multi-way domain structure (e.g., asset/venue/side in finance). We provide a conditional intensity function-consistent training recipe and evaluate both accuracy and calibration (mark reliability and time-rescaling diagnostics). Across finance and public MTPP benchmarks, TT-compressed RNNs reduce parameters by 40-70% while matching dense baselines and remaining competitive with attention-based and state-space models. Capacity-controlled and shaping ablations show that the observed calibration gains are not solely due to smaller parameter count: parameter-matched dense models improve calibration partly, while mark-aware TT ordering further improves likelihood, expected calibration error (ECE), and time-rescaling diagnostics over random or unstructured TT orderings.

URL: https://openreview.net/forum?id=Y0up92GcjF

---

Title: SinGLU: Sinusoidal Gated Linear Units Improve Classification Accuracy of Small Vision Transformers

Authors: Luke Byrne, Paul Murray

Abstract: Gated Linear Unit (GLU) variants such as SwiGLU are now widely used in modern Transformers. However, the GLU functions explored in the recent literature represent only a small fraction of the possible GLU design space. Starting from a systematic enumeration of a restricted family of zeroth-, first-, and second-order GLU-type formulas, we conduct a controlled study on ViT‑Tiny across CIFAR‑10, CIFAR‑100, SVHN and ImageNet‑64, instantiating each GLU formula with Sigmoid, Tanh and Sin activations. Under identical training recipes and matched parameter counts, our proposed first-order variant \textbf{SinGLU} achieves higher mean accuracy than SwiGLU across the datasets tested in this ViT-Tiny setting. Inference latency differs by <0.1\% on an NVIDIA A100 GPU, confirming cost parity.

URL: https://openreview.net/forum?id=qq4yipldw2

---

Title: Foundation Models in Robotics: A Comprehensive Review of Methods, Models, Datasets, Challenges and Future Research Directions

Authors: Aggelos Psiris, Vasileios Argyriou, Evangelos K. Markakis, Panagiotis Sarigiannidis, Stratis Gavves, Kostas Bekris, Arash Ajoudani, Georgios Th. Papadopoulos

Abstract: Over the recent years, the field of robotics has been undergoing a transformative paradigm shift from fixed, single-task, domain-specific solutions towards adaptive, multi-function, general-purpose agents, capable of operating in complex, open-world, dynamic environments. This tremendous advancement is primarily driven by the emergence of Foundation Models (FMs), i.e., large-scale neural-network architectures trained on massive, internet-scale, heterogeneous datasets that provide unprecedented capabilities in multi-modal understanding/reasoning, long-horizon planning, and cross-embodiment generalization. In this context, the current study provides a holistic, thorough, systematic, and in-depth review of the research landscape of FMs in robotics. In particular, the evolution in the field is initially delineated through five distinct research phases, spanning from the early incorporation of native Natural Language Processing (NLP) and Computer Vision (CV) models to the current frontier of multi-sensory generalization and real-world deployment. Subsequently, a highly-granular, multi-criteria, taxonomic investigation of the literature methods is performed, examining the following key aspects: a) The employed foundation model types (i.e., LLMs, VFMs, VLMs, and VLAs), b) The underlying neural network architectures, c) The adopted learning paradigms, d) The different learning stages of knowledge incorporation, e) The most common robotic tasks (including perception, planning, navigation, manipulation, and human-robot interaction), and f) The main real-world application domains. For each defined criterion/aspect, a methodical comparative analysis of the various categories of approaches and critical insights are provided. Moreover, a synthesis of the publicly available datasets, required for model training and evaluation, is provided, organized around the main recurring dataset families along with their typical uses and current gaps. Furthermore, a comprehensive and hierarchical discussion on the current open challenges and promising future research directions in the field is incorporated.

URL: https://openreview.net/forum?id=qF7vdPrpPk

---

Title: Discovering Generalizable Governing Equations for Graph Dynamical Systems with Interpretable Neural Networks

Authors: Riccardo Cappi, Paolo Frazzetto, Nicolò Navarin, Alessandro Sperduti

Abstract: The discovery of symbolic governing equations is a central goal in science; yet, it remains challenging particularly for graph dynamical systems, where the network topology further shapes the system behavior. While artificial intelligence offers powerful tools for modeling these dynamics, the field lacks a rigorous comparative benchmark to assess the true scientific utility of the discovered laws. To address this challenge, this work proposes a novel evaluation pipeline designed to rigorously assess state-of-the-art symbolic regression models for graph equation discovery. Moving beyond simple fitting metrics, this framework evaluates discovered laws based on their long-term trajectory stability and, critically, their out-of-distribution generalization to unseen graph topologies. We benchmark established methods, including sparse regression and MLP-based architectures, and introduce the Graph Kolmogorov-Arnold Network-ODE (GKAN-ODE) model, a novel adaptation of KANs explicitly tailored for this domain, augmented by hyperparameter-free multiplicative nodes and a new Spline-Wise symbolic regression algorithm. Across a suite of synthetic and real-world graph dynamical systems, we numerically demonstrate through extensive experiments that neural-based approaches, particularly the GKAN-ODE model, recover exact ground-truth equations and achieve trajectory errors up to two orders of magnitude lower than the baseline methods on out-of-distribution test graphs.

URL: https://openreview.net/forum?id=a2mPNSSAYL

---

Title: Realistic Evaluation of Model Merging for Compositional Generalization

Authors: Derek Tam, Yash Kant, Brian Lester, Igor Gilitschenski, Colin Raffel

Abstract: Model merging has emerged as a practical and cost-effective approach for combining multiple pretrained models into a single model that inherits their capabilities and often achieves improved performance. Its growing popularity has led to the rapid development of numerous merging techniques. However, these methods are typically evaluated in disparate experimental settings and make differing assumptions about model architecture, data availability, and computational budget, making direct comparison difficult. In this work, we systematically characterize the relative strengths and limitations of existing merging methods by evaluating them within a unified experimental framework. Our study focuses on compositional generalization --- \ie whether merging can successfully combine distinct skills to generalize to new settings. We also analyze the computational costs of each method and examine how performance scales as the number of merged models increases. Overall, we evaluate eight merging methods in a novel benchmark spanning three distinct cross-modal settings, resulting in 12,000 unique merge configurations. Our findings reveal the absence of a one-size-fits-all merging strategy and serves as both an outline for the holistic evaluation of future merging methods as well as a cookbook for practitioners using model merging.

URL: https://openreview.net/forum?id=j7ye0nXvEm

---

Title: LAW & ORDER: Adaptive Spatial Weighting for Medical Diffusion and Segmentation

Authors: Anugunj Naman, Ayushman Singh, Gaibo Zhang, Yaguang Zhang

Abstract: Medical image analysis depends on accurate segmentation and controllable synthesis, but both tasks face severe spatial imbalance: lesions occupy small regions against large backgrounds. We study adaptive spatial weighting as a task-level design principle and instantiate it in two adapters. LAW learns per-pixel loss weights for mask-conditioned diffusion by modulating a ratio prior with a feature-dependent delta map, with normalization, clamping, and Dice regularization for stability. ORDER improves lightweight segmentation by adding selective bidirectional skip attention with stage-wise confidence gating. On held-out diffusion test sets, LAW lowers FID from 158.13$\pm$0.15 to 108.43$\pm$0.71 on Polyps, from 144.13$\pm$0.31 to 89.51$\pm$0.96 on KiTS19, and from 139.22$\pm$0.38 to 112.58$\pm$0.68 on BRISC, while improving held-out mask-recovery Dice from 0.681$\pm$0.013 to 0.825$\pm$0.003 on Polyps. When the resulting images are added to nnUNet training, downstream Polyps mDice rises from 71.7$\pm$0.4 to 74.1$\pm$0.8. On the cleaned Polyps segmentation protocol, the reported ORDER configuration reaches 76.3$\pm$1.9 mDice and 67.2$\pm$2.0 mIoU at 42K parameters and 0.11 GFLOPs, versus 70.3$\pm$1.5 mDice and 59.9$\pm$1.7 mIoU for matched MK-UNet. On BRISC under the same training recipe, ORDER reaches 77.4$\pm$0.8 mDice and 68.1$\pm$0.7 mIoU. These results position adaptive spatial weighting as a practical design idea for both medical diffusion and efficient segmentation.

URL: https://openreview.net/forum?id=sJXqzr3oLl

---

Title: SafeFix: Targeted Model Repair via Controlled Image Generation

Authors: Ouyang Xu, Baoming Zhang, RUIYU MAO, Yunhui Guo

Abstract: Deep learning models for visual recognition often exhibit systematic errors due to underrepresented semantic subpopulations. While existing debugging frameworks can identify these failure slices, effectively repairing them remains difficult. Current solutions often rely on manually designed prompts to generate synthetic images—an approach that introduces distribution shift and semantic errors, often resulting in new bugs. To address these issues, we introduce SafeFix, a framework for distribution-consistent model repair via controlled generation that employs a diffusion model to generate semantically faithful images that modify only specific failure attributes while preserving the underlying data distribution. To ensure the reliability of the repair data, we implement a verification mechanism using a large vision--language model (LVLM) to enforce semantic consistency and label preservation. By retraining models on the synthetic data, we significantly reduce errors in rare cases and improve overall performance. Our experiments show that SafeFix achieves superior robustness by maintaining high precision in attribute editing without introducing additional bugs.

URL: https://openreview.net/forum?id=TtpW6JiEiW

---

Title: Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches, and Certifiable Robustness

Authors: Róisín Luo, James McDermott, Colm O'Riordan

Abstract: Lipschitz continuity is a fundamental property of neural networks that characterizes their sensitivity to input perturbations. It plays a pivotal role in deep learning, governing robustness, generalization and optimization dynamics. Despite its importance, research on Lipschitz continuity is scattered across various domains, lacking a unified perspective. This paper addresses this gap by providing a systematic review of Lipschitz continuity in deep learning. We explore its theoretical foundations, estimation methods, regularization approaches, and certifiable robustness. By reviewing existing research through the lens of Lipschitz continuity, this survey serves as a comprehensive reference for researchers and practitioners seeking a deeper understanding of Lipschitz continuity and its implications in deep learning.

URL: https://openreview.net/forum?id=pRZ0RKl11f

---

Title: Achieving Adaptivity and Optimality for Multi-armed Bandits using Exponential-Kullback Leibler Maillard Sampling

Authors: Hao Qin, Kwang-Sung Jun, Chicheng Zhang

Abstract: We study the problem of $K$-armed bandits with reward distributions belonging to a one-parameter exponential distribution family. In the literature, several criteria have been proposed to evaluate the performance of such algorithms, including Asymptotic Optimality, Minimax Optimality, Sub-UCB, and variance-adaptive worst-case regret bound. Thompson Sampling-based and Upper Confidence Bound-based algorithms have been employed to achieve some of these criteria. However, none of these algorithms simultaneously satisfy all the aforementioned criteria.

In this paper, we design an algorithm, Exponential Kullback-Leibler Maillard Sampling (abbrev. \expklms), that achieves multiple optimality criteria simultaneously, including Asymptotic Optimality, Minimax Optimality with a $\sqrt{\ln (K)}$ factor, Sub-UCB, and a variance-adaptive worst-case regret bound. Our algorithm design follows the Minimum Empirical Divergence framework~\citep{honda2011asymptotically,maillard2011apprentissage}, with the exploration probability of arm $a$ proportional to $\text{exp}\left(-L(N_{t-1, a}) \text{KL}(\hat{\mu}_{t-1, a}, \max_{a'} \hat{\mu}_{t-1, a'})\right)$, where $L(\cdot)$ is an inverse temperature function, $N_{t-1, a}$ is the number of times arm $a$ that has been pulled before time $t$, $\hat{\mu}_{t-1, a}$ is the empirical mean of arm $a$ before time $t$, and $\text{KL}(\cdot, \cdot)$ is the Kullback-Leibler divergence between two distributions in the one-parameter exponential distribution family. Our analysis allows different choices of inverse temperature function $L(k)$. We also provide numerical simulations demonstrating the effectiveness of our algorithms.

URL: https://openreview.net/forum?id=IuVkRmecVp

---

Title: Not All Structure Is Learned: Disentangling Inherited and Learned Representations in Recurrent Networks

Authors: Mark Alence

Abstract: Structure observed in trained recurrent networks may be inherited from input encodings rather than learned from data. We develop and apply a three-step decomposition to disentangle the two: (1) compare trained representations against untrained baselines to isolate input-driven structure, (2) compare against information-theoretic bounds to quantify what is achievable without learning, and (3) use causal interventions to test whether inherited and learned components are functionally used. Applied to GRUs trained via behavioral cloning on aliased navigation in a 127-node binary tree, the most prominent hidden-state feature, a depth gradient on PC1, is already present before training: an untrained GRU captures 97% of the trained correlation, reflecting input structure rather than learned spatial knowledge. What training adds is within-class node discrimination via sequential memory. Replacing depth-stratified observations with random class assignments eliminates the inherited axis; the GRU compensates with 7x greater learned spatial discrimination while maintaining comparable performance. PCA ablation reveals a double dissociation in exploration pattern, confirming that both inherited and learned components are causally involved in behavior. Applied to a non-hierarchical radial arm maze, the framework recovers an analogous inherited axis but qualitatively different learned structure: visit history tracking rather than spatial disambiguation.

URL: https://openreview.net/forum?id=1RfgHzf5IA

---

Title: CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization

Authors: Min Zhang, YUYIN WANG, Zhongxiang Dai, Zhikang Chen, Jie Zhou, Miao Liu, Sen Cui

Abstract: Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potential in learning out-of-distribution (OOD) representations. Despite showing competitive performance, the prompt-based CLIP methods still suffer from: i) inaccurate text descriptions, which leads to degraded accuracy and robustness, and poses a challenge for zero-shot CLIP methods. ii) limited vision-language embedding alignment, which significantly affects the generalization performance. To tackle the above issues, this paper proposes a novel Conditional Domain prompt Learning (CoDoL) method, which utilizes readily-available domain information to form prompts and improves the vision-language embedding alignment for improving out-of-distribution (OOD) generalization. To capture both instance-specific and domain-specific information, we further propose a lightweight Domain Meta Network (DMN) to generate input-conditional tokens for images in each domain. Extensive experiments on four OOD benchmarks (PACS, VLCS, OfficeHome, and DigitDG) validate the effectiveness of our proposed CoDoL method in terms of improving the vision-language embedding alignment as well as the out-of-distribution generalization performance.

URL: https://openreview.net/forum?id=MDxhbeE21D

---

Title: Do We Really Need to Approach the Entire Pareto Front in Many-Objective Bayesian Optimisation?

Authors: Chao Jiang, Jingyu Huang, Miqing Li

Abstract: Many-objective optimisation, a subset of multi-objective optimisation, involves optimisation problems with more than three objectives. As the number of objectives increases, the number of solutions needed to adequately represent the entire Pareto front typically grows substantially. This makes it challenging, if not infeasible, to design a search algorithm capable of effectively exploring the entire Pareto front. This difficulty is particularly acute in the Bayesian optimisation paradigm, where sample efficiency is critical and only a limited number of solutions (often a few hundred) are evaluated. Moreover, after the optimisation process, the decision-maker eventually selects just one solution for deployment, regardless of how many high-quality, diverse solutions are available. In light of this, we argue an idea that under a very limited evaluation budget, it may be more useful to focus on finding a single solution of the highest possible quality for the decision-maker, rather than aiming to approximate the entire Pareto front as existing many-/multi-objective Bayesian optimisation methods typically do. Bearing this idea in mind, this paper proposes a \underline{s}ingle \underline{p}oint-based \underline{m}ulti-\underline{o}bjective search framework (SPMO) that aims to improve the quality of solutions along a direction that leads to a good tradeoff between objectives. Within SPMO, we present a simple acquisition function, called expected single-point improvement (ESPI), working under both noiseless and noisy scenarios. We show that ESPI can be optimised effectively with gradient-based methods via the sample average approximation (SAA) approach and theoretically prove its convergence guarantees on acquisition optimisation under the SAA. We also empirically demonstrate that the proposed SPMO is computationally tractable and outperforms state-of-the-arts on a wide range of benchmark and real-world problems.

URL: https://openreview.net/forum?id=ZsI4bmqUD8

---

Title: A Survey on Efficient Protein Language Models

Authors: Shouren Wang, Debargha Ganguly, Vinooth Rao Kulkarni, Van Yang, Zhuoran Qiao, Daniel Blankenberg, Vipin Chaudhary, Xiaotian Han

Abstract: Protein language models (pLMs) have become indispensable tools in computational biology, driving advances in variant effect prediction, functional annotation, structure prediction, and engineering. However, their rapid expansion from millions to tens of billions of parameters introduces significant computational, accessibility, and sustainability challenges that limit practical application in environments constrained by GPU memory, hardware availability, and energy budgets. This survey presents the first comprehensive review of efficient pLMs, synthesizing recent advancements across four key dimensions. We first examine (1) dataset efficiency through meta-learning-based few-shot and scaling-law-guided data allocation; and (2) architecture efficiency via lightweight alternatives including quantized transformers, embedding compression, and convolution-based designs. Furthermore, we review (3) training efficiency through scaling-law-informed pretraining, structure-integrated multimodal approaches, and low-rank adaptations with diverse distillation strategies; and (4) inference efficiency via quantization, dense-retrieval, and structure-search methods. By providing a structured taxonomy and practical guidance, this survey enables the development of high-performance, scalable, yet sustainable next-generation pLMs.

URL: https://openreview.net/forum?id=PTReuOwsXz

---

Title: MEMETRON: Memetic Response Optimizer for Reward-Guided Post-Decoding Optimization of Large Language Models

Authors: Son The Nguyen, Theja Tulabandhula

Abstract: Modern large language models (LLMs) are commonly optimized using scalar reward signals defined over completed responses, applied both during training and at inference time. However, most such reward-guided post-decoding methods remain one-shot: they independently sample a set of responses, score each once, and select the best. Staying shallow and narrow leaves higher-reward responses unrealized, while scaling up to shallow and wide sampling exacerbates reward hacking, making downstream selection methods such as Best-of-$N$ and Self-consistency unreliable. We propose MEMETRON, an anytime memetic optimization framework that formulates reward-guided post-decoding optimization (RPDO) as discrete black-box optimization over completed responses. MEMETRON alternates between GENETRON for population-based optimization and ANNETRON for annealing-based local refinement under a black-box scalar reward. On mathematical reasoning, MEMETRON increases Pass@$k$ correctness coverage and improves the selection reliability of Best-of-$N$ and Self-consistency; on instruction following, it improves LLM-judge preference. On AMC 12 2025 with Qwen3-8B (non-thinking) and Skywork-Reward-V2-Llama-3.1-8B-40M, for instance, MEMETRON raises Best-of-$N$ accuracy from 61.9% at its 16-sample initialization to 78.6% after one generation and 85.7% after five, at an average of 464 seconds and 12 batched LLM requests per generation. By contrast, one-shot Best-of-$N$ with 1024 samples reaches only 47.6%. On verifiable tasks, MEMETRON can incorporate ground-truth correctness via reward shaping. Comparing shaped and unshaped runs exposes extreme cases where the reward model and ground-truth correctness disagree, and the resulting contrastive pairs serve as training signal for reward model fine-tuning, rejection-sampling SFT warmups for RL-based training pipelines such as PPO and GRPO, and direct preference learning such as DPO.

URL: https://openreview.net/forum?id=QRW8OGn3vb

---

Title: Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks

Authors: Hamid Dashtbani, Mehdi Dousti Gandomani, AmirMahdi Sadeghzadeh

Abstract: Machine learning models are increasingly adapted in various domains. However, adversarial examples pose a significant threat to the reliable deployment of these models. In recent years, some powerful adversarial example attacks have been proposed for the fast and query-efficient generation of adversarial examples, even in black-box scenarios, highlighting the need for scalable, low-cost, and powerful defenses. In this work, we present two contributions to the domain of black-box adversarial example attacks and defenses. First, we propose Random Logit Scaling (RLS), a randomization-based defense against black-box score-based adversarial example attacks. RLS is a plug-and-play, post-processing defense that can be implemented on top of any existing ML model with minimal effort. The idea behind RLS is to confuse an attacker by outputting falsified scores resulting from randomly scaled logits while maintaining the model accuracy. We show that RLS significantly reduces the success rate of state-of-the-art black-box score-based attacks while preserving the accuracy and minimizing confidence score distortion compared to state-of-the-art randomization-based defenses. Second, we introduce a novel adaptive attack against AAA, a SOTA non-randomized black-box defense against black-box score-based attacks that also modifies output logits to confuse attackers, demonstrating its vulnerability against adaptive attacks.

URL: https://openreview.net/forum?id=CXafPv4aAG

---

Title: White-Box Sensitivity Auditing with Steering Vectors

Authors: Hannah Cyberey, Yangfeng Ji, David Evans

Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of large language models (LLMs) primarily rely on black-box evaluations that assess model behavior only through input–output testing. These methods are limited to tests constructed in the input space, often generated by heuristics. In addition, many socially relevant model properties (e.g., gender bias) are abstract and difficult to measure through text-based inputs alone. To address these limitations, we propose a white-box sensitivity auditing framework for LLMs that leverages activation steering to conduct more rigorous assessments through model internals. Our auditing method conducts internal sensitivity tests by manipulating key concepts relevant to the model's intended function for the task. We demonstrate its application to bias audits in four simulated high-stakes LLM decision tasks. Our method consistently indicates substantial dependence on protected attributes in model predictions, even in settings where standard black-box evaluations suggest little or no bias.

URL: https://openreview.net/forum?id=EfinGGyQRz

---

Title: ReDiTT: Retrieval Augmented Conditional Diffusion Transformers for Asynchronous Time Series

Authors: Saiyue Lyu, Zhitian Zhang, Ruizhi Deng, Thibaut Durand

Abstract: We present a diffusion based model for asynchronous time series prediction, where the goal is to predict the next inter event time and event type. To address the inherent uncertainty of future events, we introduce ReDiTT, a retrieval augmented conditional diffusion transformer that operates in latent space. ReDiTT retrieves structurally similar latent sequences from a memory bank during both training and inference and incorporates them as reference conditions through cross attention. This retrieval based conditioning allows the model to attend to relevant temporal dynamics and provides global structural guidance for generation. As a result, ReDiTT stabilizes long horizon forecasting and improves sample diversity. Experiments on seven real world datasets demonstrate state of the art performance on next event prediction and long horizon forecasting. Our code is available at https://github.com/BorealisAI/ReDiTT.

URL: https://openreview.net/forum?id=sv35KiCipb

---

Title: Task-Aware Model Merging via Fisher-Weighted Median

Authors: Baban Gain, Saswati Dana, Udit Sharma, Arnab Kumar Mondal, Prathosh AP, Dinesh Garg, Amith Singhee, Asif Ekbal

Abstract: Fine-tuning large language models provides strong in-domain performance, but it can limit generalization and requires the storage of many specialized models. Retraining a unified multitask model is often infeasible due to data unavailability or high computational cost. Most model merging approaches perform arithmetic operations directly on model parameters. Although research on model merging has expanded significantly in recent years, two distinct directions have become dominant: (1) techniques that mitigate interference from redundant parameters and sign conflicts, and (2) techniques that account for the varying sensitivity of individual parameters. However, these directions have largely evolved independently, without leveraging their complementary strengths. In this work, we aim to bridge this gap by integrating insights from both. We propose DRIFT-MEDIAN, a sensitivity-aware model merging approach that combines Fisher information and a coordinate-wise importance measure in a weighted median aggregation framework. Comprehensive experiments on several LLMs and CLIP-based models demonstrate that task-vector interference mitigation and parameter sensitivity provide complementary signals for model merging. DRIFT-MEDIAN integrates both principles within a unified framework and improves mean performance retention (PRR) across the evaluated settings, with gains varying across individual tasks. We make the code publicly available at https://github.com/babangain/drift-median.

URL: https://openreview.net/forum?id=tB6bb0ZosX

---

Title: GE-FM: Geometry-aware Energy-based Flow Matching for Non Euclidean Manifolds

Authors: Ayush Roy, Arjun Ramesh Kaushik, Vishnu Suresh Lokhande, Nalini K. Ratha, Venu Govindaraju

Abstract: Flow Matching has emerged as a powerful framework for generative transport and denoising, yet existing formulations are inherently Euclidean, neglecting the curved and time-evolving geometry of diffusion manifolds. Recent higher-order extensions seek to recover curved transport by explicitly modeling higher derivatives, but these approaches introduce instability and accumulate discretization error, particularly in few-step ODE sampling regimes. We propose a strictly first-order-in-time, energy-based flow matching framework that incorporates geometry through Christoffel-adjusted dynamics. Our method defines a total energy as the sum of kinetic energy induced by the predicted velocity field and a learned potential energy, and enforces approximate energy conservation along transport trajectories. Energy conservation encourages optimal low-energy denoising paths and yields a smoother optimization landscape, leading to faster and more stable convergence. Crucially, the energy formulation induces a time-dependent Riemannian metric that captures the evolving diffusion geometry without explicit manifold supervision. Christoffel symbols derived from this induced metric adjust the velocity field to account for curvature, implicitly modeling higher-order effects without introducing additional learnable dynamics. This geometric correction is manifold-agnostic and adapts automatically to the evolving diffusion structure. Empirically, our method outperforms existing few-step baselines, achieving improved performance on both FID and mode coverage ($\sim$ 80\% $\uparrow$ on synthetic spiral datasets). To the best of our knowledge, this is the first geometry-aware flow matching framework that integrates energy conservation and Christoffel dynamics for stable curved generative transport. The code used in this work is available at: \url{https://github.com/AyushRoy2001/GE-FM.git}.

URL: https://openreview.net/forum?id=Wb9wwUF7aV

---

Title: Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out-of-Distribution Generalization

Authors: Mahdi Samiei, Arash Marioriyad, Arman Tahmasebi-Zadeh, Mohamadreza Fereydooni, Mahdi Ghaznavi, Mahdieh Soleymani Baghshah

Abstract: Recent progress has pushed AI frontiers from pattern-recognition tasks toward problems that require step-by-step, System-2-style reasoning, especially with large language models. Yet, unlike learning, where generalization and out‑of‑distribution (OoD) evaluation concepts are well formalized, there is no clear, consistent definition or metric for “reasoning ability.” We propose Complexity Out‑of‑Distribution (Complexity OoD) generalization as a framework and problem setting to measure reasoning. A model exhibits Complexity OoD generalization when it maintains performance on test instances whose minimal required solution complexity, either representational (richer solution structure) or computational (more reasoning steps/program length), exceeds that of all training examples. We formalize complexity via solution description Kolmogorov complexity and operational proxies (e.g., object/relation counts; reasoning‑step counts), clarifying how Complexity OoD differs from length and compositional OoD. This lens unifies learning and reasoning: many cases solvable with System‑1‑like processing at low complexity become System‑2‑like under complexity pressure, while System‑2 can be viewed as generalization over solution structures. We translate this perspective into practice with recommendations for operationalizing Complexity OoD across the stack: incorporating complexity into benchmark and evaluation metric design, rethinking supervision to target solution traces (from final outcomes to process‑level feedback and RL/search), seeking and designing inductive biases for Complexity‑OoD generalization, addressing learning‑to‑reason spillovers such as spurious shortcuts, semantic robustness, catastrophic forgetting, and step‑wise calibration. In light of recent controversies over LLM reasoning, we put the problem on firm footing: treat reasoning as Complexity OoD, enabling rigorous evaluation and more systematic research.

URL: https://openreview.net/forum?id=07fh13gWs0

---

Title: Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Authors: Sarthak Kumar Maharana, Shambhavi Mishra, Yunbei Zhang, Shuaicheng Niu, Taki Hasan Rafi, Jihun Hamm, Marco Pedersoli, Jose Dolz, Yunhui Guo

Abstract: Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-world deployment, where data undergoes continual distributional shifts. Continual Test-Time Adaptation (CTTA) addresses this challenge by adapting pretrained models to non-stationary target distributions on-the-fly, without access to source data or labeled targets, while mitigating two critical failure modes: catastrophic forgetting of source knowledge and error accumulation from noisy pseudo-labels over extended time horizons. In this comprehensive survey, we formally define the CTTA problem, analyze the diverse continual domain shift patterns that characterize different evaluation protocols, and propose a hierarchical taxonomy that categorizes existing methods into three families: optimization-based strategies (entropy minimization, pseudo-labeling, parameter restoration), parameter-efficient methods (normalization layer adaptation, adaptive parameter selection), and architecture-based approaches (teacher-student frameworks, adapters, visual prompting, masked modeling). We systematically review representative methods within each category and present comparative benchmarks and experimental results across standard evaluation settings. Finally, we discuss limitations of current approaches and highlight emerging research directions, including adaptation of foundation models and black-box systems, providing a roadmap for future research in robust continual test-time adaptation. We encourage visiting our repository at [https://github.com/sarthaxxxxx/Awesome-Continual-Test-Time-Adaptation](https://github.com/sarthaxxxxx/Awesome-Continual-Test-Time-Adaptation)

URL: https://openreview.net/forum?id=mM3r03Xw1V

---

Title: Investigating the limits of free-form debate as a scalable oversight strategy

Authors: Gareth Tan, Leonid Tsyplenkov, Edy Nastase, Gabriel Recchia

Abstract: Debate is a scalable oversight method involving two copies of a strong model trained to defend alternative responses to a question, with a judge with less task-relevant information, time, or domain-specific capability evaluating which answer is better supported. We replicate and extend a result from prior work demonstrating that training Llama3-8B-Instruct-262k as a debater led to increased performance of a GPT-4-class judge model on QuALITY, a question-answering task that grants the debaters a capability advantage via information asymmetry. When replicating the original setup as closely as possible, we confirm that training debater models in free-form, multi-round debate increased judge accuracy. However, this finding did not generalize across alternative tasks or models, and did not replicate consistently under our closest approximation to the original setting. These results suggest that the effectiveness of free-text debate as a scalable oversight method is sensitive to task structure, model pairing, and training conditions, and highlight the need for greater understanding of when and why debate improves judge accuracy. We identify several factors that may influence debate's success and outline directions for future work aimed at characterizing the conditions under which debate strengthens oversight reliably.

URL: https://openreview.net/forum?id=vRCGzuAhOM

---

Title: MatchEx: Model-Level GNN Explanations with Multi-Granular Insights

Authors: Sayan Saha, Sanghamitra Bandyopadhyay

Abstract: Graph Neural Networks (GNNs) are increasingly deployed in high-stakes domains where interpretability is crucial. Existing model-level explanation methods largely rely on generative models, which often produce motifs that fail to resemble real instances, cannot account for the diversity of discriminative motifs recognized by the classifier for a target class and lack mechanisms for translating global explanations to instance-level insights. We present MatchEx, a framework that discovers discriminative motifs directly from real instances by optimizing a novel matching objective. Unlike isomorphism, which can only recover identical motifs that rarely occur in real-world graphs, this objective extends beyond exact matches to provably recover semantically similar motifs, allowing generalizable explanations. The matching mechanism also enables projection of class level rationales onto individual graphs for faithful instance-level insights. When a single motif fails to explain all instances, MatchEx adaptively partitions the instances in a class into coherent subgroups with distinct rationales. Extensive experiments across six real and synthetic datasets show that MatchEx consistently outperforms state-of-the-art baselines, delivering coherent, generalizable, and multi-granular explanations.

URL: https://openreview.net/forum?id=YMETLG2WvM

---

Title: Predicting integers from continuous parameters

Authors: Bas Joris Maat, Peter Bloem

Abstract: We study the problem of predicting numeric labels that are constrained to the integers or to a
subrange of the integers. For example, the number of up-votes on social media posts, or the
number of bicycles available at a public rental station. While it is possible to model these as
continuous values, and to apply traditional regression, this approach changes the underlying
distribution on the labels from discrete to continuous. Discrete distributions have certain
benefits, which leads us to the question whether such integer labels can be modeled directly
by a discrete distribution, whose parameters are predicted from the features of a given
instance. Moreover, we focus on the use case of output distributions of neural networks,
which adds the requirement that the parameters of the distribution be continuous so that
backpropagation and gradient descent may be used to learn the weights of the network.
We investigate several options for such distributions, some existing and some novel, and
test them on a range of tasks, including tabular learning, sequential prediction and image
generation. We find that overall the best performance comes from two distributions: Bitwise,
which represents the target integer in bits and places a Bernoulli distribution on each, and
a discrete analogue of the Laplace distribution, which uses a distribution with exponentially
decaying tails around a continuous mean.

URL: https://openreview.net/forum?id=d1WKFlKFEa

---

Title: Extracting Common Components from Partially Observed Views Using Diffusion Geometry

Authors: Bar Weiss, Hau-Tieng Wu, Ronen Talmon

Abstract: Data acquired from multiple sensors or modalities, commonly referred to as multiview data, is prevalent in real-world applications. A core problem in multiview data analysis is finding representations of common components across views while filtering out view-specific nuisance factors. A widely spread assumption in existing methods is that the views are fully aligned, where each sample has measurements from all views. However, in practice, data is often partially aligned, where some samples have missing measurements from one or more views, and only a subset of the samples are fully aligned. In this work, we propose ADM+, a multiview manifold learning algorithm that computes a low-dimensional embedding of common information from partially aligned data. ADM+ extends Alternating Diffusion Maps (ADM), an existing multiview manifold learning method, to the partial alignment setting by using fully aligned samples as anchor points for extracting common components for unaligned samples. Unlike existing methods, ADM+ does not require prior imputation of missing data or interpolation in the embedding space and makes use of all available data. We provide a computationally efficient implementation, improving upon the $O(N^3)$ time complexity of ADM, and a theoretical analysis showing that ADM+ approximates an anisotropic diffusion process that emphasizes common components. Empirical evaluations across three domains -- dynamical systems, synthetic multiview images, and real-world functional magnetic resonance imaging (fMRI) -- demonstrate that ADM+ achieves favorable performance compared to kernel- and manifold-based baselines. In addition, ADM+ shows robustness to distributional discrepancies between aligned and unaligned samples.

URL: https://openreview.net/forum?id=yTHGIV8ToF

---

Title: Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

Authors: Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo, Christina Lioma

Abstract: Pruning Large Language Models (LLMs) reduces memory and inference costs by removing parts of the network, producing smaller models that retain most of their accuracy. As attention layers are the most resource-intensive parts of LLMs, pruning them is a promising compression strategy. Prior work shows that up to 33% of attention layers can be pruned with minimal accuracy loss. Nevertheless, the impact of attention pruning on model interpretability, specifically faithfulness and confidence calibration, remains unstudied. To address this gap, we study how pruning attention layers affects explanation faithfulness and confidence calibration across five LLMs and eight datasets. While the pruned models often maintain high accuracy, we find that their faithfulness and calibration often degrade. Notably, faithfulness and calibration can fluctuate significantly, even when accuracy remains stable, highlighting a misalignment between model confidence, interpretability, and accuracy. Our findings suggest that layer pruning can affect LLMs' interpretability and reliability in ways not captured by accuracy and efficiency measures alone. We recommend including explainability and calibration metrics when evaluating pruned models.

URL: https://openreview.net/forum?id=VxZd6HfMOo

---

Title: Federated Learning with Projected Trajectory Regularization

Authors: Tiejin Chen, Yuanpu Cao, Yujia Wang, Cho-Jui Hsieh, Jinghui Chen

Abstract: Federated learning enables joint training of machine learning models from distributed clients without sharing their local data. One key challenge in federated learning is to handle non-identically distributed data across the clients, which leads to deteriorated model training performance. Prior works in this line of research mainly focus on utilizing last-step global model parameters/gradients or the linear combinations of the past model parameters/gradients, which do not fully exploit the potential of global information from the model training trajectory. In this paper, we propose a novel federated learning framework with projected trajectory regularization (FedPTR) for tackling the data heterogeneity issue, which proposes a unique way to better extract the essential global information from the model training trajectory. Specifically, FedPTR allows local clients or the server to optimize an auxiliary (synthetic) dataset that mimics the learning dynamics of the recent model update and utilizes it to project the next-step model trajectory for local training regularization. We conduct rigorous theoretical analysis for our proposed framework under nonconvex stochastic settings to verify its fast convergence under heterogeneous data distributions. Experiments on various benchmark datasets and non-i.i.d. settings validate the effectiveness of our proposed framework.

URL: https://openreview.net/forum?id=vfCztZvcP3

---

Title: End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI

Authors: Yihong Chen, Jiancheng Yang, Deniz Sayin Mercadier, Hieu Le, Juerg Schwitter, Pascal Fua

Abstract: Reconstructing cardiac motion from CMR sequences is critical for diagnosis, prognosis, and intervention. Existing methods rely on complete CMR stacks to infer full heart motion, limiting their applicability during intervention when only sparse observations are available.
We present TetHeart, the first end-to-end framework for unified 4D heart mesh recovery from both offline full-stack and intra-procedural sparse-slice observations.
Our method leverages deformable tetrahedra to capture shape and motion in a coherent space shared across cardiac structures. Before a procedure, it initializes detailed, patient-specific heart meshes from high-quality full stacks, which can then be updated using whatever slices can be obtained in real time, down to a single slice during the procedure.
TetHeart incorporates several key innovations: (i) an attentive slice-adaptive 2D–3D feature assembly mechanism that integrates information from arbitrary numbers of slices at any position; (ii) a distillation strategy to ensure accurate reconstruction under extreme sparsity; and (iii) a weakly supervised motion learning scheme requiring annotations only at keyframes, such as the end-diastolic and end-systolic phases. Trained and validated on three large public datasets, and evaluated on additional private interventional and public datasets without retraining, TetHeart achieves state-of-the-art accuracy in both pre- and intra-procedural settings. Code and dataset is available at \url{https://github.com/Scalsol/TetHeart}.

URL: https://openreview.net/forum?id=9k00kN5yk2

---

Title: POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

Authors: Zhangheng LI, Jianing Zhu, Junyuan Hong, Sungmin Eum, Shuowen Hu, Suya You, Zhangyang Wang

Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and visual data, where privacy-sensitive examples could be unintentionally encoded, raising concerns about privacy or copyright violation. To this end, Multi-modality Machine Unlearning (MMU) was proposed as a mitigation that can effectively force MLLMs to forget private information. However, the robustness of such unlearning methods is not fully exploited when the model is published and accessible to malicious users. In this paper, we propose a novel adversarial strategy, namely Prompt-Optimized Parameter Shaking (POPS), aiming to recover the supposedly unlearned multi-modality knowledge from the MLLMs. Our method elicits the victim MLLMs to generate potential private examples via prompt-suffix optimization, and then exploits these synthesized outputs to fine-tune the models so they disclose the true private information. The experiments on the different MMU benchmarks reveal substantial weaknesses in the existing MMU algorithms. Our POPS can even achieve a near-complete recovery of supposedly erased sensitive information on the unlearned MLLMs, exposing fundamental vulnerabilities that challenge the foundational robustness of representative MMU-based privacy protections.

URL: https://openreview.net/forum?id=wMiEcH84l9

---

Title: Neural Networks Performance Prediction using Weights and Gradients Analysis

Authors: Michael Bohadana, Alon Schneider, Gilad Katz

Abstract: Neural network performance predictors are widely used to accelerate neural architecture search, but existing methods face a persistent trade-off: learning-based predictors require costly per-dataset initialization, while lightweight proxies are fast yet struggle to exploit prior experience and often degrade under dataset shift. We introduce NAP2, a hybrid performance predictor that models early training dynamics. NAP2 tracks the temporal evolution of layer-wise weight and gradient statistics over a small number of mini-batches, producing accurate rankings from as little as 100 mini-batches per candidate. Crucially, NAP2 supports cross-dataset reuse: a predictor trained on one dataset can be applied to another without fine-tuning, avoiding the re-initialization overhead incurred by many model-based approaches. Experiments on NAS-Bench-201 across CIFAR-10, CIFAR-100, and ImageNet16-120 show that NAP2 is competitive with strong hybrid baselines under limited budgets and delivers cost-effective cross-dataset transfer, outperforming established learning-curve and zero-cost baselines at short query times. We further demonstrate robustness to significant distribution shift, with a predictor trained on CIFAR-10 transferring effectively to SVHN. Our code and trained models are available at https://anonymous.4open.science/r/NAP2-6027/README.md.

URL: https://openreview.net/forum?id=51TWh8tlSy

---


New submissions
===============


Title: More Than Parameter Count: Depth–Width Shape in Controlled Pre-LN Transformer Sweeps

Abstract: Parameter-count scaling models summarize architecture using a single quantity, $N$, even though models with similar parameter counts can allocate capacity very differently between depth and width. We ask whether depth–width shape adds predictive information in controlled Pre-LN Transformer sweeps. We analyze 30 decoder-only runs spanning 27M to 6.9B parameters and up to 143B training tokens. All runs use the same corpus, predetermined training seed, and final-step training-loss endpoint; within the documented 1B, 3B, and 7B tiers, the optimization configuration is held fixed while depth and width vary. The objective is to compare architecture shapes under controlled tier-specific protocols rather than after architecture-specific tuning. We compare a parameter-count baseline, a smooth $D/W$ correction, and two thresholded excess-depth models. AICc favors the parameter-count baseline ($R^2=0.966$), whereas leave-one-run-out cross-validation within the observed architecture grid favors a logarithmic excess-depth model, reducing RMSE from $0.108$ to $0.102$ nats. Descriptively, the deepest configurations have higher final training loss than shallower, wider configurations by $0.157$, $0.162$, and $0.119$ nats at the 1B, 3B, and 7B tiers. The 1B contrast is closely matched; the 3B contrast is approximately compute-matched despite small offsetting parameter and token differences; and the 7B contrast gives the wider run an advantage on all recorded budget dimensions. The fitted logarithmic onset lies at approximately 40–54 layers over $W=512$–4096, but marks activation of a shape penalty rather than an optimal depth. The result is a modest, protocol-conditioned association and does not estimate seed-to-seed variability, validation performance, or a universal law of Transformer depth.

URL: https://openreview.net/forum?id=InOiMlQeNT

---

Title: Best Policy Tracking in Reinforcement Learning

Abstract: Reinforcement learning commonly exhibits non-monotonic performance dynamics in single-task settings. During learning, a sequence of policies is generated whose true performance is unknown and may oscillate significantly across iterations. Despite this well-known phenomenon, standard experimental practice typically reports either the final visited policy or the best among periodically evaluated ones, implicitly assuming that these selection criteria reliably approximate the best-performing policy visited throughout the entire sequence.

In this work, we present an empirical study on the implications of non-monotonicity for policy selection in single-task policy-based reinforcement learning. We introduce metrics both to characterize performance degradation and to assess the accuracy and efficiency of different selection criteria. These metrics allow us to evaluate commonly used selection strategies across standard continuous control benchmarks and actor-critic method. Our results show that widely adopted selection practices can deviate considerably from the best policy encountered during learning, leading to underperforming estimates and suboptimal use of computational resources. As an alternative, we recommend a validation-based policy selection strategy constrained by computational cost, which provides a more consistent and accurate approximation of the best visited policy while improving selection efficiency. Overall, our findings highlight policy selection as an underexplored factor in the evaluation protocols of reinforcement learning and emphasize the need for more explicit and standardized validation-based selection procedures.

URL: https://openreview.net/forum?id=b4BbTUF8Xe

---

Title: Rethinking Negative Prompting for Concept Erasure: Unconditional Anchoring and Localized Guidance

Abstract: Text-to-image diffusion models have rapidly transitioned from research prototypes to widely deployed tools. Along with their impressive capabilities, these models can also generate harmful or undesirable content (e.g., NSFW imagery), motivating practical mechanisms for removing such concepts. Negative prompting, which uses the negative-prompt prediction as the CFG anchor, is among the most widely used mechanisms for steering sampling trajectories away from unwanted concepts. However, for jailbreak prompts, it can fail in a regime we refer to as guidance collapse, where the model's denoising predictions for the prompt and negative prompt become similar, so that the guidance difference driving negative prompting vanishes and the guided prediction degenerates to the negative-prompt prediction, undermining erasure. To fix this issue, we propose TraSCE (Trajectory Steering for Concept Erasure), a training-free inference-time method that makes negative prompting reliable under guidance collapse. TraSCE combines two complementary components: (i) unconditional anchoring, which makes the fallback under guidance collapse a neutral unconditional prediction rather than the negative-prompt prediction, and (ii) an alignment-localized steering step, which deliberately induces this collapse for adversarial prompts and steers the sample to the unconditional fallback, while remaining minimally intrusive on benign prompts. We evaluate TraSCE on adversarial prompt benchmarks targeting nudity and violence, and further demonstrate its applicability to erasing artistic styles and objects. Across these settings, TraSCE substantially improves the robustness of negative prompting to jailbreak prompts without any training, weight updates, or pre-collected concept-specific data (prompts or images), while keeping the impact on general image quality small.

URL: https://openreview.net/forum?id=RgdWbRqviu

---

Title: Partition-Losses Fine-Tuning: Contamination-Robust Backdoor Unlearning

Abstract: The widespread use of large-scale, weakly curated training data and third-party checkpoints makes training convenient, but also exposes models to poisoning-based backdoor attacks. These attacks embed a hidden trigger-to-target association during training: the infected model behaves normally on clean inputs but predicts an attacker-chosen label whenever the trigger appears, posing risks for security-sensitive deployment and model reuse. Post-training fine-tuning has become a practical default defense as it is computationally efficient and does not require control over the original training pipeline. However, most existing post-training fine-tuning defenses rely exclusively on a clean dataset and ignore suspicious inputs, leaving potentially useful signals unexploited. In this paper, we propose Partition-Losses Fine-Tuning (PL), a simple and architecture-agnostic post-training method that leverages a defender-constructed candidate suspicious pool, either synthesized from publicly known trigger templates or flagged by off-the-shelf detectors. PL minimizes the benign loss on clean data while maximizing the loss on these candidate suspicious samples, directly breaking the trigger-to-target association. Importantly, PL remains effective even when the synthesized trigger pattern differs from that of the true attack or the target class is unknown. Comprehensive experiments show that PL matches or surpasses clean-only fine-tuning methods under the same computational budget while substantially reducing the amount of clean data required. It also remains effective under realistic contamination, hyperparameter variation, and cross-attack settings.

URL: https://openreview.net/forum?id=GblbNsIZyn

---

Title: OmegaMGT: Isometric Dynamics and Asymptotic Dimensional Contraction in Single-Highway Transformers

Abstract: Current Transformer architectures dominate natural language processing under a constant dimensionality paradigm, which imposes a computational cost that grows linearly with network depth ($\mathcal{O}(L)$). This dimensional redundancy demands massive hardware capabilities for parameter storage and context maintenance, severely limiting their operational scalability and prohibiting the local deployment of state-of-the-art models in resource-constrained environments. In this work, we introduce the \textbf{\textit{Omega Minimal Gated Transformer (OmegaMGT)}}, a deep neural architecture that challenges this standard through a geometric dimensionality reduction topology. We propose two fundamental innovations: (1) a \textit{per-instance} spectral token compression mechanism based on Principal Component Analysis (PCA), which aligns the sequence length with its intrinsic informational content; and (2) a \textit{single-highway} structure based on the Minimal Gated Unit (MGU), which replaces additive residual connections with a modulated deep recurrence. Through rigorous theoretical analysis, we demonstrate that OmegaMGT decouples the asymptotic \textit{time} complexity of inference from model depth, achieving an $\mathcal{O}(1)$ bound as $L \to \infty$, while activation memory grows only through the compressed attention term ($\mathcal{O}(L\omega^2)$) rather than the full model width. We further prove that, at initialization and under (near-)isometric projections, the gradient norm is preserved across depth, providing a stable pathway for backpropagation in massively deep regimes. This approach offers a theoretical framework for building massively deep yet computationally bounded language models, pointing toward LLM inference under tight compute budgets.

URL: https://openreview.net/forum?id=3ux6hTsXlC

---

Title: Evaluating Ensemble Methods for Algorithmic Stability Under Data Perturbations

Abstract: In this paper, we assess methods that can be utilized to induce selection stability in classification problems, such that predictions remain consistent under perturbations to the data. Standard classification algorithms often exhibit sensitivity near decision boundaries, which can compound instability that already exists from overfitting to a subset of high-influence samples. The recently proposed inflated argmax and subbagging framework --- which takes an ensemble of predictions from models trained on different subsamples of the data combined with a selection rule that allows for multi-class predictions in ambiguous cases --- addresses this and can yield substantial improvements in stability, as we validate in the cause-of-death classification setting. However, this approach still requires careful consideration of the stability of the base algorithm. In our experiments, the subbagging ensemble technique is not universally beneficial, as subbagging an already stable classifier can be detrimental to stability. We further examine the extent to which the Rashomon set --- the collection of distinct models that achieve near-optimal predictive performance --- can serve as a substitute for subbagging by leveraging predictive multiplicity of the Rashomon effect rather than resampling. We find that it is possible for the Rashomon ensemble to stabilize predictions, though this effect is by no means guaranteed. Our results are consistent with an influence equalization interpretation of ensemble techniques, which helps provide a natural explanation for both when subbagging fails and why the Rashomon ensemble, despite its theoretical appeal, can struggle in practice without a significant multiplicity effect.

URL: https://openreview.net/forum?id=7CCr0dgNwm

---

Title: Rule-Dependent Neighbor Routing in Transformers Trained on Life-like Cellular Automata

Abstract: Mechanistic interpretability benefits from learned-model testbeds whose task semantics, local dependencies, and canonical sufficient statistics are exactly specified. We train small transformer encoders to predict one-step updates of the Conway, Seeds, and DayNight cellular automata on $16\times16$ toroidal grids.Across three token-ordering conditions, three seeds, and a fixed non-grid graph control, the models reach near-perfect accuracy and depend causally on true neighbors rather than matched non-neighbor controls. Comparing the trained models under a fixed architecture, data distribution, and analysis pipeline, we find rule-associated differences in the layer distribution of neighbor-attending heads: routing is concentrated early for Seeds, late for DayNight, and seed-sensitive for Conway, and a head's neighbor-attention mass correlates strongly with its single-head ablation damage. In Seeds, the predicate dead & count == 2 is perfectly linearly decodable on dead cells, where it determines the output — and remains so after removing the explicit readout direction — but only partially on alive cells, while a uniformly faithful $0..8$ neighbor count is not recovered under the same probes. Selected-head case studies show that the learned within-neighbor attention weighting is functionally important to the model's output, and DayNight additionally preserves complement-related representational geometry across layers. These results establish Life-like cellular automata as a controlled comparative testbed for learned routing mechanisms, while leaving the complete decision circuit and the causal role of specific rule properties open.

URL: https://openreview.net/forum?id=Yajrs3kwZx

---

Title: Benchmarking Stochastic Interpolants for Modeling Physical Systems

Abstract: Generative models have recently emerged as powerful surrogates for physical systems, demonstrating increased accuracy, stability, and/or statistical fidelity. However, most approaches rely on iteratively denoising a Gaussian, a choice that may not be the most effective for autoregressive prediction tasks, where the current and future states are often closely related. We therefore benchmark generative models based on stochastic interpolants, which can learn maps between arbitrary densities, against standard Gaussian-denoising frameworks across a range of fluid and climate systems. Following prior work, we consider a transport map for dynamical systems that learns an interpolant between a point mass at the current state and a conditional distribution of future states. We find that no single model is universally best: whether a generative model is needed at all, and which one performs best, depends on the physics of the system, the sampling scheme, and the performance metric of interest. Within this landscape, stochastic interpolants can be a strong and competitive baseline. They benefit from the proximity of successive states and are highly flexible during inference, with the capability to tradeoff deterministic accuracy, spectral consistency, and probabilistic calibration. This flexibility makes them well-suited to tailoring a generative surrogate to a given system and objective, and motivates further exploration of transport maps designed specifically for physical emulation.

URL: https://openreview.net/forum?id=OOlMYM5bLr

---

Title: Returning the First Step to the Policy: Repairing the Multistart Bias of Neural Solvers on Selective Routing

Abstract: Multistart decoding is central to state-of-the-art constructive neural routing solvers. On selective routing problems such as the Orienteering Problem (OP) and the Prize-Collecting TSP (PCTSP), a solver must decide not only how to order visits but which nodes are worth visiting at all. We identify a bias of POMO-style multistart decoding in this setting: the first node of every rollout is fixed outside the policy, assigned zero log-probability, and therefore excluded from learning. We repair it with selection-aware starts (SAS), a minimal change: one 128-parameter linear head that selects rollout starts and restores step-zero credit, leaving the POMO backbone and shared-baseline training unchanged. On OP50, SAS cuts the single-seed one-shot gap from 12.4% to 3.9% with full-budget performance unchanged; the repair transfers to PCTSP50 at the mechanism level, although the one-shot gap remains larger than on OP50. We then characterize in advance when the bias matters. Across five OP50 budgets, a model-free start-stakes measure, computed by classical local search as the share of achievable value lost when the start is forced to a uniformly random feasible node, tracks the realized one-shot benefit of the repair for this model family trained on the target distribution ($R^2 = 0.96$). Part of that correlation is structural; the empirical finding is the conversion rate: the realized benefit is a near-constant 0.72 to 0.75 of the available start value at every budget, a ratio nothing in the training objective guarantees, and it transfers unchanged to PCTSP. Within the measured range, the near-zero-benefit end of the spectrum is reached only through problem structure. Under the evaluated shifts the two factors separate: the measured start value persists, but the trained head converts less of it: the advantage halves on OP100 and reverses on OPLib, where every configuration decodes at gaps between 49% and 64%.

URL: https://openreview.net/forum?id=KfzwYcSe7I

---

Title: From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models

Abstract: YOLO-Pose models provide efficient keypoint localization, but do not quantify the associated spatial uncertainty. We introduce a lightweight post-hoc probabilistic extension that augments a trained YOLO-Pose model with calibrated bivariate predictive distributions over keypoint locations, centered at the model’s original predictions. Concretely, we train additional probabilistic heads with an importance-weighted negative log-likelihood to predict an input-dependent $2\times2$ dispersion matrix for each keypoint, followed by Gaussian calibration for broad downstream compatibility or Student-$t$ calibration for distributional fidelity. Complementing this, we propose an evaluation protocol that combines a suite of distributional calibration diagnostics with average keypoint precision (AKP), a keypoint-level extension of the COCO AP protocol for assessing reliability rankings. Experiments on COCO show that the learned uncertainty estimates enable effective keypoint-level reliability ranking, Student-$t$ calibration best captures the empirical residual distribution, and uncertainty-based pruning removes unreliable keypoints. A central application-level demonstration is vision-based aircraft landing, where calibrated covariances for runway keypoints support uncertainty-aware aircraft position estimation and downstream sensor fusion.

URL: https://openreview.net/forum?id=MHG2uIP3hB

---

Title: UNO: Unlearning via Orthogonalization in Generative Models

Abstract: As generative models become increasingly powerful and pervasive, the ability to unlearn specific data, whether due to privacy concerns, legal requirements, or the correction of harmful content, has become increasingly important. Unlike in conventional training, where data are accumulated and knowledge is reinforced, unlearning aims to selectively remove the influence of particular data points without costly retraining from scratch. To be effective and reliable, such algorithms need to achieve (i) forgetting of the undesired data, (ii) preservation of the quality of the generation, (iii) preservation of the influence of the desired training data on the model parameters, and (iv) a small number of training steps. We propose fast unlearning algorithms based on loss gradient orthogonalization for unconditional and conditional generative models. We show that our algorithms are able to forget data while maintaining the fidelity of the original model. On standard image benchmarks, our algorithms achieve orders of magnitude faster unlearning times than their predecessors, such as gradient surgery. We demonstrate our algorithms with datasets of increasing complexity (MNIST, CelebA and ImageNet-1K) and for generative models of increasing complexity using VAEs and diffusion transformers.

URL: https://openreview.net/forum?id=N2iiNBZkVi

---

Title: Self-Improvements in Modern Agentic Systems: A Survey

Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains. We offer a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and control logic. Within this framework, self-improvement is formalized as a self-induced update operator that obtains and commits updates to model parameters or scaffold components. We organize prior work by updating targets and by the signals that drive change, then review applications and discuss evaluation before closing with open problems and future directions.

URL: https://openreview.net/forum?id=JyrcEL77Nq

---

Title: A Unified Framework for Zero-Shot Reinforcement Learning

Abstract: Zero-shot reinforcement learning (RL) has emerged as a setting for developing general agents, capable of solving downstream tasks without additional training or planning at test-time. While conventional RL optimizes policies for fixed rewards, zero-shot RL requires learning representations that enable immediate adaptation to arbitrary reward functions. As the field matures, the growing diversity of approaches demands a foundational framework reconciling different perspectives under a common structure. In this work, we introduce a formal, unified framework for zero-shot RL, allowing for rigorous comparisons across methods. We propose a taxonomy organizing the algorithmic landscape along two levels: representation, distinguishing between compositional and direct methods based on their exploitation of value function decompositions; and learning paradigm, differentiating between reward-free and pseudo reward-free training. Additionally, we propose a unified view of existing error bounds, decomposing the total error into three primary contributing components: inference, reward, and approximation, serving as a foundation for more grounded comparisons of zero-shot methods.

URL: https://openreview.net/forum?id=Nv26zx1Ywa

---

Title: Two-way Cutting-plane Algorithm for Best Subset Selection Considering Multicollinearity

Abstract: When linear dependence exists between some explanatory variables in a regression model, the estimates of regression coefficients become unstable, thereby making the interpretation of the estimation results unreliable. To eliminate such multicollinearity, we propose a high-performance method for selecting the best subset of explanatory variables for linear and logistic regression models. Specifically, we first derive a bilevel reformulation of the optimization problem for best subset selection with a multicollinearity constraint. We then develop a two-way cutting-plane algorithm that uses cutting planes in two ways; one type of cutting planes is used to approximate an upper-level nonlinear objective function, and the other type of cutting planes is used to remove subsets with collinearity. We prove that this algorithm outputs a solution with guaranteed global optimality within a finite number of iterations. Computational results based on synthetic and public datasets demonstrate the effectiveness of our method by comparison with the L1-regularized estimation and the previous cutting-plane algorithm for eliminating multicollinearity. Our method is not only a fast computational framework for best subset selection, but also has the advantage of improving the reliability of regression analysis by eliminating multicollinearity.

URL: https://openreview.net/forum?id=9MjXUmQLZP

---

Title: SPID: Distilled Protein Backbone Generation

Abstract: Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unprecedented capabilities for de novo protein design. However, despite their generation quality, these models are constrained by slow sampling, often requiring hundreds of iterative steps. This computational bottleneck limits their practical utility in large-scale protein discovery, where thousands to millions of candidate structures are needed. To address this challenge, we explore the techniques of score distillation, which has shown great success in reducing the number of sampling steps in the vision domain while maintaining high generation quality. However, a straightforward adaptation of these methods results in unacceptably low designability. We introduce Score Protein identity Distillation (SPID), which resolves this incompatibility by combining few-step generation with inference-time noise scaling. SPID adapts the Score identity Distillation (SiD) framework to both diffusion- and flow-based models without requiring access to pretraining data. Applied to the Proteina flow-matching model, our 16-step generator achieves 94.4% designability — exceeding the 400-step teacher (94.2%) — while delivering more than a 20-fold reduction in effective sampling time and maintaining comparable diversity and novelty. SPID generalizes across unconditional generation, fold-class conditional generation, and motif scaffolding, and extends to equivariant diffusion architectures, achieving significant speedup with improved designability.
The resulting reduction in inference cost facilitates large-scale in silico protein design, thereby advancing diffusion-based models toward real-world protein engineering applications.

URL: https://openreview.net/forum?id=l1fS626LVB

---

Title: Neural Boundary Integral Operators: Solver-Consistent Learning Across Geometry, Resolution, and Kernel Mismatch

Abstract: Black-box neural operators such as FNO and DeepONet learn PDE solution maps from data, but they often struggle on irregular geometry, do not encode known singular operator structure, and do not natively handle unbounded domains without truncation, mapping, or specialized exterior treatment. We introduce the Neural Boundary Integral Operator (NBIO), a hybrid neural operator that embeds a boundary integral equation (BIE) and restricts learning to smooth model discrepancy. NBIO combines: (M1) an analytically split kernel, consisting of a known singular Green’s kernel and a learned smooth correction; (M2) a differentiable second-kind solve for the boundary density; and (M3) discretization-independent continuous evaluation in boundary arc length and query location. This construction supports zero-shot transfer across geometry and resolution and across interior/exterior regimes in 2D; a jointly trained 3D proof of concept covers both regimes on deformed spheres. In the zero-correction limit, NBIO recovers a spectrally accurate boundary element solver on 2D Laplace problems, with relative L2 error of 4.8 × 10^-7, while learned black-box baselines on the same data remain near 10^-2–10^-1 and do not improve under boundary refinement. When the available classical prior is misspecified or incomplete, learning becomes useful: on a modified-Helmholtz task with a Laplace prior, the learned correction reduces error from 1.13 to 8.3 × 10^-3; on a variable-coefficient problem it improves a fixed prior from 9.0 × 10^-2 to 3.4 × 10^-3, although its edge over a strong FNO is modest. NBIO also extends to interface problems with prescribed jumps and volume sources. Its zero-parameter structured limit matches the reference solver and is substantially more accurate than same-protocol black-box operators. A differentiable QBX module resolves near-boundary and near-touching evaluation. On real ABC and Fusion 360 CAD meshes, the zero-parameter BIE structure achieves median errors of 2.1 × 10^-3 and 5.6 × 10^-3, respectively, 100–200× lower than same-protocol learned geometry operators, while a learned correction does not improve the known-kernel CAD solver. Overall, NBIO shows that exact numerical structure can provide solver consistency and geometry-resolution transfer, while learning is most valuable when the classical model is incomplete.

URL: https://openreview.net/forum?id=K8fA2QOlV8

---

Title: MoCo: A One-Stop Shop for Model Collaboration Research

Abstract: Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multiple LMs collaborate, compose, and complement each other. Existing research on this topic has mostly been disparate and disconnected, from different research communities, and lacks rigorous comparison. To consolidate existing research and establish model collaboration as a school of thought, we present MoCo: a one-stop Python library of executing, benchmarking, and comparing model collaboration algorithms at scale. MoCo features 26 model collaboration methods, spanning diverse levels of cross-model information exchange such as routing, text, logit, and model parameters. MoCo integrates 25 evaluation datasets spanning reasoning, QA, code, safety, and more, while users could flexibly bring their own data. Extensive experiments with MoCo demonstrate that most collaboration strategies outperform models without collaboration in 61.0% of (model, data) settings on average, with the most effective methods outperforming by up to 25.8%. We further analyze the scaling of model collaboration strategies, the training/inference efficiency of diverse methods, highlight that the collaborative system solves problems where single LMs struggle, and discuss future work in model collaboration, all made possible by MoCo. We envision MoCo as a valuable toolkit to facilitate and turbocharge the quest for an open, modular, decentralized, and collaborative AI future.

URL: https://openreview.net/forum?id=jBHaZSA4nv

---

Title: Retrieval-Augmented Multimodal Language Model with Heterogeneous Evidence Grounding for Clinical Prognosis

Abstract: Clinical prognosis requires reasoning over heterogeneous evidence: continuous waveforms, dense laboratory measurements, longitudinal patient history, and the empirical outcome patterns of similar patients. Multimodal language models offer a principled substrate for this integration, yet current systems operate over an impoverished version of the clinical record. Structured measurements lose their numerical meaning when serialized as text. Each encounter is treated in isolation, discarding the prognostic signal in serial waveform change. And all predictions are grounded exclusively in knowledge memorised during pre-training, with no mechanism to consult the outcomes of clinically similar patients. We introduce PRISM (Population-Retrieved Inference for Structured Multimodal Prognosis), which recovers each of these missing evidence streams through a matched representational strategy: per-feature numerical tokenization that preserves measurement fidelity, temporal aggregation via a learnable recency decay that distills longitudinal ECG context, and a dual-modality population memory bank that retrieves outcome-concordant historical patients as non-parametric evidence for language model reasoning. On the 1,443-task MDS-ED benchmark (MIMIC-IV), PRISM achieves state-of-the-art performance across every task group: 86.2% Diagnosis, 93.1% Deterioration, 91.8% ICU admission, and 92.8% Mortality (AUROC). Our ablations reveal that the three evidence streams serve complementary roles: numerical encoding drives laboratory-dependent diagnoses, population evidence sharpens diagnosis and ICU admission prediction, and temporal context resolves trajectory-sensitive outcomes, suggesting broader design principles for grounding clinical language models in structured, longitudinal, and population-level evidence.

URL: https://openreview.net/forum?id=5ChJTFLi5r

---

Title: Matching correlated VAR time series

Abstract: We study the problem of matching correlated VAR time series databases, where a multivariate time series is observed along with a perturbed and permuted version, and the goal is to recover the unknown matching between them. To model this, we introduce a probabilistic framework in which two time series $(x_t)_{t\in[T]},(x^\#_t)_{t\in[T]}$ are jointly generated, such that $x^\#_t=x_{\pi^*(t)}+\sigma \tilde{x}_{\pi^*(t)}$, where $(x_t)_{t\in[T]},(\tilde{x}_t)_{t\in[T]}$ are independent and identically distributed vector autoregressive (VAR) time series of order $1$ with Gaussian increments, for a hidden $\pi^*$. The objective is to recover $\pi^*$, from the observation of $(x_t)_{t\in[T]},(x^\#_t)_{t\in[T]}$. This generalizes the classical problem of matching independent point clouds to the time series setting.

We derive the maximum likelihood estimator (MLE), leading to a quadratic optimization over permutations, and theoretically analyze an estimator based on linear assignment. For the latter approach, we establish recovery guarantees, identifying thresholds for $\sigma$ that allow for perfect or partial recovery. Additionally, we propose solving the MLE by considering convex relaxations of the set of permutation matrices (e.g., over the Birkhoff polytope). This allows for efficient estimation of $\pi^*$ and the VAR parameters via alternating minimization.
Empirically, we find that linear assignment often matches or outperforms MLE relaxation based approaches.

URL: https://openreview.net/forum?id=V6vGEZC0It

---

Title: Analysis of Dirichlet Energies as Over-smoothing Measures

Abstract: We analyze the distinctions between two metrics often used to measure over-smoothing: the Dirichlet energies induced by the unnormalized graph Laplacian and the normalized graph Laplacian. We demonstrate that the latter fails to satisfy the axiomatic definition of a node-similarity measure proposed by Rusch \textit{et al.} By formalizing fundamental spectral properties of these two definitions, we highlight critical distinctions necessary to select the metric that is spectrally compatible with the GNN architecture, thereby resolving ambiguities in monitoring the dynamics.

URL: https://openreview.net/forum?id=egSZ5SMvUk

---

Title: Unifying Back-Propagation and Forward-Forward Algorithms through Model Predictive Control

Abstract: We propose a Model-Predictive-Control (MPC) view of neural-network training that unifies back-propagation (BP) and the Forward-Forward (FF) algorithm as two extreme instances of a receding-horizon planner family: at each layer the planner solves a length-$h$ terminal-cost subproblem and stitches the partial gradients into a global update. Three planner backends instantiate the family by varying the window stride: **chunked** and **overlap-all** both collapse to FF at $h=1$ and recover full BP at $h=T$, while **deep-supervision** also collapses to FF at $h=1$ but at $h=T$ instead recovers a full deep-supervision schedule, making the horizon $h$ a single dial that gradually interpolates between the FF and BP regimes the title alludes to.Theoretically, on a deep linear network with $T$ blocks (or nonlinear residual networks with bounded block Jacobians) we prove a finite-$T$ gradient-cosine deviation
$1-\cos^2(\theta_h) = O(\delta^{p}), \qquad \delta = 1-h/T,$
where $p=3$ for **deep-supervision** and $p=2$ for **overlap-all** and **chunked** (with a matching $\Omega(\delta^{2})$ lower bound); combined with a biased-PL analysis this gives per-iteration suboptimality $1-r(h) = O(\delta^{p})$, so the marginal gain of an extra horizon step diminishes polynomially as $h \to T$ while memory grows linearly, ruling out FF and BP as optimal and forcing an interior $h^\star \in (1,T)$ as the unique minimiser of any reasonable accuracy–memory objective.
Empirically, on CIFAR-100 (ViT-b/16+LoRA, ResNet-50 stage- and block-cut) and RoBERTa-base+LoRA on four GLUE tasks, the measured $1-\cos^2\theta_h$ tracks the predicted polynomial decay, and the resulting test-accuracy curve shows the same diminishing returns: an interior horizon $h \approx T/3$ closes $80\text{–}98\%$ of the FF–BP accuracy gap across all tasks and, on the majority of them, matches or slightly surpasses full BP ($\le +1.3$ pp), while peak GPU memory scales only linearly in $h$. This pins an interior $h^\star$ as the optimal accuracy–memory operating point — ruling out both FF and full BP — and delivers substantially better memory efficiency (~30–80% lower peak GPU usage on transformer backbones) at no cost to accuracy.

URL: https://openreview.net/forum?id=5WtC6PiRAs

---

Title: Finite-Step Sinkhorn Certificates for Nonnegative Stream-Mixing Products

Abstract: We develop deterministic finite-step certificates for products of nonnegative stream-mixing matrices in expanded-stream residual architectures, with Manifold-Constrained Hyper-Connections as the motivating example. The certified object is the mixer product $T_L=H_{\rm res}^{(L)}\cdots H_{\rm res}^{(1)}$, which controls the backward streamwise $A_{\max}$ gain of the mixing pathway. We show that marginal errors of nonnegative mixers give depth-wise induced-norm certificates, including a structural ceiling $\|T_L\|_1\le n$, a telescoping refinement, and a depth-uniform Dobrushin certificate under scrambling. We then analyze finite column-then-row Sinkhorn normalization: Hilbert-projective contraction gives an explicit output column-defect bound after finitely many steps, including the half-step loss from ending each full step with row normalization. Combining these estimates gives calibration rules mapping logit clipping, temperature, depth, and Sinkhorn step count to a certified mixer-path backward budget. Matrix-specific local analysis identifies the sharp asymptotic full-step rate $\sigma_2(P)^2$ near the Sinkhorn limit while separating certificates from local predictors. CIFAR-10 diagnostics and a matched short-horizon transformer run report marginal defects, product certificates, exact propagated gains, Dobrushin audits, adaptive stopping behavior, and local-predictor curves.

URL: https://openreview.net/forum?id=NyEeOl6Gn7

---

Title: Design Principles for Sequence Models via Coefficient Dynamics

Abstract: Deep sequence models, ranging from Transformers and State Space Models (SSMs) to more recent approaches such as gated linear RNNs, fundamentally compute outputs as linear combinations of past value vectors. To draw insights and systematically compare such architectures, we develop a unified framework that makes this output operation explicit, by casting the linear combination coefficients as the outputs of autonomous linear dynamical systems driven by impulse inputs. This viewpoint, in spirit substantially different from approaches focusing on connecting linear RNNs with linear attention, reveals a common mathematical theme across diverse architectures and crucially captures softmax attention, on top of RNNs, SSMs, and related models. In contrast to new model proposals that are commonly evaluated on benchmarks, we derive design principles linking architectural choices to model properties. Thereby identifying tradeoffs between expressivity and efficient implementation, geometric constraints on input selectivity, and stability conditions for numerically stable training and information retention. By connecting several insights and observations from recent literature, the framework both explains empirical successes of recent designs and provides guiding principles for systematically designing new sequence model architectures.

URL: https://openreview.net/forum?id=DRR5ZCN7Qe

---

Title: DRIFT: Constrained Dynamic Topic Models for Weak-Signal Trend Discovery

Abstract: Standard topic models recover a corpus's dominant topics, and existing extensions add either temporal dynamics or domain-guided discovery, but rarely both. This gap matters in settings such as mental-health discourse in social media, where the trends of real interest are rare, evolve in both prevalence and content over time, and are easily crowded out by a large, shifting background of everyday content. We introduce DRIFT (Dynamic Recovery of Infrequent Fine-grained Trends), a constrained dynamic topic model that jointly represents time-varying topic prevalence through covariate-driven document-level priors, time-varying topic content through a random-walk prior over topic-word logits, and mild guidance for a domain of interest, without requiring seed guidance to be pre-assigned to individual trends. Two deterministic reparameterizations enforce this guidance by construction rather than by penalty: a seed-mass floor on the content of weak-signal trends, and a downweighting of their prevalence in documents without seed words, so that every topic-proportion and topic-word distribution remains a valid probability simplex throughout inference. We derive a variational EM algorithm that handles the resulting nonconjugate terms with reparameterized Monte Carlo estimation. On a synthetic benchmark built to test recovery of rare, and drifting trends, DRIFT outperforms LDA, DTM, STM, and a constrained matrix-factorization baseline on rare-trend identification, topic-content quality, and temporal recovery, at a modest cost in held-out perplexity relative to the strongest baselines. On a real case study of mental-health discourse in youth YouTube vlogs, DRIFT recovers interpretable, temporally coherent weak-signal trends without ground-truth labels. These results suggest that explicit, feasibility-preserving constraints are an effective way to keep dynamic topic models sensitive to rare but meaningful signals.

URL: https://openreview.net/forum?id=n9nktcl5Eh

---

Title: The Kernel Inner Product Space: Dimensionality Reduction as Kernel Alignment

Abstract: Dimensionality reduction has splintered into families (spectral methods such as PCA, kernel PCA, Isomap, and Laplacian eigenmaps on one side, neighbor embeddings such as t-SNE and UMAP on the other), each carrying its own objective and its own solver. We show that these methods are often a single operation seen from different angles. Representing both the data and its embedding by centered \emph{kernels} and scoring their agreement with the \emph{RV coefficient}, a cosine in matrix space, casts dimensionality reduction as the projection of a fixed input kernel onto the set of output kernels that a $q$-dimensional configuration can realize. When the output kernel is linear this achievable set is a convex cone, the projection is a closed-form spectral truncation with an exact alignment ceiling, and the classical spectral methods are its optima for different input kernels; a class-label target extends the same projection to a continuous soft-LDA. Beyond the linear output, an invariance argument leaves exactly two canonical readout families, dot-product and distance, and each turns the cone into a smooth manifold whose dimension we compute, set by the motions its readout cannot see; on the distance manifold of the Student-t readout, the RV gradient becomes force-directed and its attraction a precise relative of t-SNE's. We further show that the kernel diagonal is a degree metric that tethers the embedding's spread, and that repulsion is not part of the objective but the gradient of the one mode, the embedding's global volume, that the centering discards, which places t-SNE, but not UMAP, within the framework's reach. A sequence of experiments verifies each of these predictions in turn.

URL: https://openreview.net/forum?id=UxUdmXFV0Z

---

Title: BayPrAnoMeta: Bayesian Proto-MAML for Few-Shot Industrial Image Anomaly Detection

Abstract: Industrial image anomaly detection is a challenging problem owing to extreme class imbalance and the scarcity of labeled defective samples, particularly in few-shot settings. We propose BayPrAnoMeta, a Bayesian generalization of Proto-MAML for few-shot industrial image anomaly detection. Unlike existing Proto-MAML approaches that rely on deterministic class prototypes and distance based adaptation, BayPrAnoMeta replaces prototypes with task-specific probabilistic normality models and performs inner loop adaptation via a Bayesian posterior predictive likelihood. We model normal support embeddings with a Normal–Inverse–Wishart (NIW) prior, producing a Student-$t$ predictive distribution that enables uncertainty-aware, heavy-tailed anomaly scoring and is essential for robustness in extreme few-shot settings. We further extend BayPrAnoMeta to a federated meta-learning framework with supervised contrastive regularization for heterogeneous industrial clients and prove convergence to stationary points of the resulting nonconvex objective. Experiments on the MVTec AD and VisA datasets show that the proposed method consistently achieves significant AUROC improvements over MAML, Proto-MAML, and PatchCore based methods in few-shot anomaly detection settings.

URL: https://openreview.net/forum?id=sfel3uqWCE

---

Title: TNODEV: Toolbox for Neural ODE Verification

Abstract: Neural ordinary differential equations (neural ODE) gained attention in safety critical settings such as continuous-time controllers for cyber-physical systems and classifiers integrated into automated decision pipelines, raising the question whether their behavior can be formally verified. Existing tools dedicated to neural ODE provide only a single reachability call without iterative input-set refinement, limiting the precision of their verdicts to whatever one reachability call can deliver. We present TNODEV, the first formal verifier for neural ODE that integrates a falsification checker, a fast interval-based reachability backend based on continuous-time mixed monotonicity, a verification and refinement loop with three input-set splitting heuristics, and a parallel scheduler in a single end-to-end pipeline. TNODEV supports safe-set inclusion verification on pure neural ODE, neural ODE in closed loop with a neural network controller and general neural ODE (GNODE), with the safe set specified either as an interval or as the half-space intersection induced by a target classification label. We evaluate TNODEV on a range of benchmarks across safe-set inclusion and classification-robustness properties, including a direct reachability comparison against NNV 2.0 and CORA and a verification comparison against NNV 2.0 on MNIST general neural ODE classifiers.

URL: https://openreview.net/forum?id=3lbdS0A1N5

---

Title: Rollback-Safe Long-Horizon Generation: Auditing and Transactional State Commit for Narrative Agents

Abstract: Long-horizon language-model systems do not only fail by forgetting context; they also fail when fluent but unsafe continuations are written into persistent memory as state. We study a rollback-safe runtime for narrative agents that treats each continuation as a candidate state transition: text, extracted events, and state deltas are staged, audited, gated, and then either committed atomically or preserved as a rejected trace. The paper’s organizing claim is that runtime integrity depends on coupling detection to the write boundary. Three findings support this claim. First, post-hoc audit over 48 internal outputs finds human-confirmed corruption in every condition family, showing that context and memory alone do not eliminate unsafe drafts. Second, auditor calibration exposes detection as the bottleneck: the current auditor is high-recall but poorly calibrated, with overall precision/recall/F1 of 0.113/0.727/0.196 and a failed motivation-drift channel. Third, write-control replay shows that merely critiquing or recording a failure is architecturally insufficient; the empirical question is how much corruption the deployed gate actually intercepts. In the present replay, the current gate misses most human-confirmed corruptions, while a human-calibrated upper bound eliminates silent commits by construction. Public sanity checks on ToolSandbox, ConStory-Bench, and LongMemEval test the same runtime, narrative-consistency, and memory/KFI mechanisms outside the internal story scaffold. The contribution is diagnostic: rollback-safe commit is a useful control boundary, but its value is gated by calibrated audit signals.

URL: https://openreview.net/forum?id=xVD5lur9ZA

---

Title: Proxy-to-Preference Transfer Under Structured Label Mismatch: Tail-Class Collapse in Few-Shot Adaptation from Physics-Derived Proxies to Human Preference Labels

Abstract: Many learning systems must predict a scarce, subjective human target using supervision that is abundant but only a structured proxy for that target: a simulator output, a heuristic score, or a physical model. We study the resulting proxy-to-preference transfer problem — how to exploit abundant proxy-labeled data without letting the proxy silently redefine the target it approximates — and identify a specific, previously undocumented failure mode of naive adaptation in this setting. We instantiate the problem in multimodal thermal sensation inference: a source domain of 141,330 synthetic state-text pairs labeled with a physics-derived comfort index (Predicted Mean Vote, PMV) and a target domain of 279 real, chronologically ordered human comfort judgments (Thermal Sensation Vote, TSV). The proxy-target relationship here is not merely noisy: PMV is directionally biased and compresses the dynamic range of TSV, and the bias itself is regime-dependent. We pretrain a multimodal encoder on proxy supervision and evaluate zero-shot and few-shot transfer to the human target. Aggregate metrics reproduce a familiar-looking tension: zero-shot transfer is already strong, while few-shot adaptation raises exact agreement but degrades mean absolute error (MAE). Disaggregating by class reveals what the aggregate numbers hide: adaptation does not fail uniformly. On our main split, the adapted model's predictions never fall below the majority label region of its small adaptation set — it stops covering 38% of the true label distribution entirely, in the direction opposite the proxy's own bias. We argue this is not an artifact of one dataset but a structural risk whenever (i) the adaptation set is small, (ii) its label distribution is unrepresentative of the deployment distribution — which chronological, non-i.i.d. real-world adaptation windows often are by construction, and (iii) the target label space is discrete and imbalanced. We call this tail-class collapse under adaptation, give empirical evidence for the mechanism (not just the symptom), and argue that structured proxy supervision is best used as an inductive bias for representation learning rather than a target substitute, with adaptation procedures evaluated for label-tail coverage rather than aggregate error alone.

URL: https://openreview.net/forum?id=mnahpPEIqM

---

Title: VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Abstract: Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer's focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.

URL: https://openreview.net/forum?id=zfM3xSZMQi

---

Title: Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams

Abstract: Team science holds that leadership is \emph{contingent}: it helps only under specific conditions, and capable, autonomous teams may need none at all. We ask the analogous question for multi-agent LLM teams: under what \emph{measurable} conditions does process-level coordination control add value, and do those conditions match what team science predicts? Answering it takes measurement, not just accuracy: we use \emph{behavioral signatures} (majority lock-in, exploration, recovery from an incorrect round-0 consensus) and \emph{per-action ablations}, clean because each controller is an explicit action set rather than a monolithic prompt.

We operationalize three classical leadership styles (transactional, transformational, situational) as controllers over a shared action vocabulary (\texttt{explore}, \texttt{revise}, \texttt{accept}, \texttt{synthesize}), holding the agent set and final aggregation fixed. Team science \citep{hackman2002leading,bass1985leadership,hersey1969management} supplies the substrate: Bass's two-component structure maps to the \texttt{accept} and \texttt{revise} actions, giving the transactional controller its decomposition. A matched controller with the same actions but an \emph{arbitrary} rule recovers no better than majority voting, so it is the theory-derived rule, not the vocabulary, that does the work.

Across four task regimes and three open-weight model families on a single backend, no controller dominates by accuracy, as the contingency view predicts. Against a \emph{shared} round-0 vote, generated once and reused across conditions, transactional control matches the vote on all 12 (model, regime) combinations to within $1.3$pp, and accuracy gains appear in only two of the 36 leadership entries, situational and transformational, both on the single \texttt{llama-4-scout} social combination, where the round-0 majority is unreliable. Against the stronger flat baseline, only situational still gains ($+8$pp). A \emph{recovery-advantage} account, tested with four boundary probes, says when a controller beats plain interaction: only where the round-0 majority is unreliable, the task is recoverable, and undirected interaction does not already repair it. These conditions map onto contingency theory (leadership substitutes, path-goal redundancy, and the situational readiness gap), so a largely null accuracy result is what the theory predicts, not a failure of the controllers. We read process-level coordination control as a contingency to be measured and theory-mapped, not a leaderboard to be topped.

URL: https://openreview.net/forum?id=gp77UrXyjT

---

Title: Looking beyond the next token

Abstract: The most natural way to model language is rarely autoregressive. The structure of causal language model training assumes that each token can be predicted from prior context, a process that contrasts with humans’ natural writing and reasoning process, which is often non-linear and hierarchical. While this mismatch is well-documented, the working assumption has been that architectural changes are needed to address it. We argue that by simply rearranging and modifying the training data, autoregressive modeling can more accurately imitate some aspects of the true data-generating process without any changes to the architecture or training infrastructure. We introduce Trelawney, a purely data-centric method that modifies the training data by interleaving sequences with special lookahead tokens that contain future information. This simple data augmentation, requiring no changes to model architecture or training infrastructure, equips models to both condition on future goals and generate them. We present representative results on high-entropy tasks like path planning, algorithmic reasoning, zebra puzzles, and controllable generation, demonstrating improved performance on tasks with branching paths or long-horizon planning. Finally, our method enables the generation of plausible long-term goals at no additional cost, potentially opening doors to new capabilities beyond the current language modeling paradigm.

URL: https://openreview.net/forum?id=x4Bzz1rY43

---

Title: Efficient Logical Reasoning with Hyperdimensional Computing Through Interference-Canceled Hybrid Encoding

Abstract: Hyperdimensional Computing or Vector Symbolic Architectures (HDC/VSA) offer a middle ground between symbolic and neural reasoning, combining compositional structure with efficient vector operations, yet their application to relational reasoning has hit a wall: interference in role-filler bindings degrades the signal-to-noise ratio as predicate arity increases: an unavoidable $1/\sqrt{k}$ decay that makes reliable inference difficult beyond binary or ternary relations. We address this through query-aware decoding rather than a better encoding alone. Our hybrid encoding combines predicate binding with positional permutations, preserving distinguishability while enabling invertible argument recovery, and successive interference cancellation (SIC) exploits known query constraints to remove interference before cleanup. Retrieval similarity improves from 0.31 to 0.92 on binary predicates, and the approach scales to 100K facts and to real knowledge graphs. Reliable decoding in turn unlocks capabilities that trained embeddings and exact reasoners do not combine: open-vocabulary reasoning over entities unseen at training, graceful robustness to noise, and multi-constraint queries. The framework trades exactness for efficiency: it suits approximate reasoning over large knowledge bases with graceful degradation, but not domains requiring formal guarantees. More broadly, the results suggest that effective reasoning in lossy vector representations comes from algorithms that exploit task structure, not from better encodings alone.

URL: https://openreview.net/forum?id=5t9BjXNGJs

---

Title: Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

Abstract: The local sharpness of the loss, the top Hessian eigenvalue $\lambda_1$, determines the largest stable gradient step, but measuring it normally requires Lanczos or Hessian-vector iterations. We observe that a single Armijo backtracking line search already carries this information at the cost of a few forward passes: the accepted step $\alpha$ brackets the \emph{directional} curvature $q = g^\top H g/\|g\|^2$ within the multiplicative band set by the backtracking factor. Across CIFAR-10, Fashion-MNIST and Imagenette, $\log\alpha$ tracks $\log\lambda_1$ at Pearson $-0.91$ to $-0.95$, giving a low-cost online Edge-of-Stability reading. Used once at initialisation, this measurement yields a learning-rate cap (a safeguard, not a faster optimiser) that makes Adam robust to a too-large initial learning rate across more than three orders of magnitude ($10^{-3}$ to $3.0$), at about one percent overhead, and it is a no-op when the chosen rate is already safe. One probe is enough: periodic in-training probing adds no robust benefit. The raw-gradient probe exposes the mechanism but needs a safety factor calibrated to the architecture by a one-minute divergence sweep. Probing along Adam's own update direction removes this calibration: a single fixed safety factor $\kappa = 2$ avoids divergence on all nine architectures we test and across the full learning-rate grids of all four benchmarks, and the recipe transfers to AdamW unchanged.

URL: https://openreview.net/forum?id=foLGwG3ne4

---

Title: Separation-Utility Pareto Frontier: An Information-Theoretic Characterization

Abstract: We study the Pareto frontier between predictive utility and separation, a fairness criterion requiring predictive independence from sensitive attributes conditional on the true outcome. Through an information-theoretic lens, we characterize the achievable separation–utility region, prove that the revealed randomized frontier is the concave closure of the deterministic frontier, and clarify how the marginal cost of separation varies along the frontier. We further identify sufficient conditions under which the trade-off is strict, providing theoretical guidance for interpreting empirical frontiers and choosing operating points. Motivated by this characterization, we develop a direct empirical regularizer based on conditional mutual information (CMI) for discrete target and sensitive variables. By estimating CMI directly from sample statistics, the resulting plug-in regularizer avoids reliance on adversarial or variational proxy losses, is compatible with deep models trained by gradient-based optimization, and provides a scalar monitor of residual separation violation with finite-sample guarantees. Experiments on COMPAS, UCI Adult, UCI Bank, CelebA, and ACS show that the proposed method traces stable separation-utility frontiers, substantially reduces separation violations, and achieves competitive trade-offs relative to established baselines, often improving the low-violation region. This study offers a principled, stable, and flexible framework for navigating separation–utility trade-offs in deep learning.

URL: https://openreview.net/forum?id=Vtulf3SPT3

---

Title: DecomposeR: Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward

Abstract: Deep research tasks require LLMs to plan what to investigate, retrieve evidence, and synthesize long-form answers across multiple branches of inquiry. Existing training paradigms either rely on short-form verifiable QA as a proxy or optimize monolithic long trajectories, which makes planning and execution difficult to disentangle and yields weak credit assignment for the planning process. We propose DecomposeR, a planner-centric deep research framework that represents research plans as typed directed acyclic graphs (DAGs), allowing planning to be made explicit, structured, and rewardable. We train a Qwen3-8B model in two stages: planner reinforcement learning (RL) first learns graph structure and query decomposition to improve research planning, and answerer reinforcement learning (RL) then learns branch-level execution and final synthesis conditioned on the learned plan. By assigning rewards to explicit planner tokens and structured components rather than to a flat trajectory, DecomposeR enables finer-grained optimization of planning while reducing the ambiguity of end-to-end training. Experiments show that DecomposeR-8B outperforms the strongest comparable open
baselines of similar size by 5.1–8.0 points across all three benchmarks (DeepResearchBench, HealthBench, ResearchQA-Mini). It also rivals open models roughly 4× larger in size (51.7 average, vs. 51.4 and 51.1 for the two strongest 30–32B open systems), due to improved planning and answering.

URL: https://openreview.net/forum?id=qG4deYtMxL

---

Title: Action-Free Learning for Data-Efficient Offline Goal-Conditioned RL

Abstract: We study the data efficiency of offline reinforcement learning (RL) algorithms that leverage action-free data. A truly data-efficient algorithm should perform comparably whether trained on mostly action-free data with limited action-labeled transitions, or on an equivalent amount of fully action-labeled data. To evaluate how well current methods live up to this promise, we conduct controlled experiments using varying ratios of action-free and action-labeled data. Our empirical analysis yields two key findings. First, all evaluated methods struggle to match the performance achieved with a fully action-labeled dataset, revealing a bottleneck between the pretrained model's capacity and the fine-tuned policy's performance that widens as action labels become increasingly scarce. Second, we identify a counterintuitive trend: representation learning methods that provide stronger low-level policy initialization do not necessarily achieve higher success rates, suggesting effective planning, rather than action quality, is the primary driver of goal-reaching performance. Building on these insights, we develop an empirical recipe for data-efficient pretraining: (a) effective subgoal planning with value-based policy extraction is essential for strong goal-reaching performance; (b) representation learning methods should be used solely to initialize the low-level policy, not for planning; and (c) incorporating value functions consistently improves performance under action-limited settings. We instantiate this recipe in QSIA, a hierarchical method that substantially outperforms all evaluated baselines. Notably, fine-tuning QSIA with as little as $10\%$ action-labeled data is generally sufficient to match the performance of training on a fully action-labeled dataset.

URL: https://openreview.net/forum?id=HIYwI8gykw

---

Title: When the Scan and Step Paths Diverge: A Reproducibility and Correctness Audit of a Full-History Mamba VLA (MTIL)

Abstract: MTIL is a full-history Mamba-2 vision-language-action (VLA) policy that advances two premises for state-space backbones in robot imitation: that encoding the entire interaction history improves closed-loop success, and that a recurrent state-space update deploys at constant per-step cost. We reproduce it on the ALOHA benchmark under a fixed, seed-controlled protocol and audit both premises, contributing three corrections. The first (C1) is a correctness defect: the recurrent step() controller shipped in the released code is broken. It collapses the canonical two-stream residual of its blocks and diverges from the sequence-parallel training path by ≈15 in action units (episode-dependent, 15–60), driving closed-loop success to 0.7% (1/150) when it is used. (MTIL's reported results and our 97.3% reproduction both run on the still-correct sequence-parallel path; what is broken is the advertised constant-cost recurrent step(), the efficiency claim's deliverable.) A ~15-line fix restores a controller that matches the parallel path to 8×10⁻⁴ (in fp32 and fp16, below the model's 0.012 ground-truth floor) and recovers rollout success to 96.7% (3×50 seeded, Wilson-95% [92.4, 98.6]), statistically indistinguishable from the 97.3% sequence-parallel reproduction. We verify the defect on a pristine clone of the published repository. The verification gap is not unique to MTIL. By source inspection plus one empirical second-repo kill-test, among the four released SSM robot policies we examined, none ships a recurrent rollout path that is verified equivalent to its parallel scan (only two ship a recurrent path at all; RoboSSM's is unimplemented dead code that raises AttributeError on first use). Across this line, the advertised linear-time recurrent inference is either unrealized or unverified. The second correction (C2) concerns efficiency: with a faithful controller in hand, the constant-per-step advantage is real at the block level (~1.4× on the Mamba blocks) but invisible end-to-end (~1.014×, flat in T), because MTIL's per-frame, uncached DINOv2 vision frontend dominates the SSM scan by more than an order of magnitude (≈27× the O(T) step block, ≈19× the O(T²) re-encode). This is a property of this encoder, not of VLAs in general. The third correction (C3) is narrower in scope: on MTIL's near-Markovian headline benchmark, full history is not necessary for high success. A 16-frame-context-trained policy solves transfer_cube at 90% (a Markovian control, 3×50 seeded, Wilson-95% [84.2, 93.8]). Two further probes target genuinely memory-dependent settings but remain inconclusive rather than extending this claim: an exploratory LIBERO-Mem comparison (n=20) shows no detectable memory gain, though both policies floor on basic manipulation before memory becomes relevant; and a delayed-cue probe we could not make conclusive (the cued task floors at 25%). We do not conclude that full history is generally unimportant for memory-dependent manipulation, only that it is not required on the benchmark MTIL's headline number comes from (cf. DSSP, which shows history helping on tasks designed to require it). A same-protocol ACT baseline reproduces at only ~25%, which we report as an under-tuned outlier and from which we draw no vision-attribution. None of these findings refute MTIL's utility; they correct why it works, and point SSM-VLA efficiency effort, for heavy-encoder designs, at the visual frontend rather than the sequence model. We release a one-command artifact for the central C1 correctness finding (C2/C3 are backed by committed result-dumps and scripts) on a pristine clone of the published code.

URL: https://openreview.net/forum?id=y2fIIRh7Jd

---

Title: Sample-Efficient Deep Cellular Deconvolution with Bulk-Anchored Pseudo-Bulk Augmentation

Abstract: Cellular deconvolution estimates cell-type proportions from bulk RNA-seq and enables cell-population analysis across large transcriptomic cohorts. Most methods depend on a cell-type-resolved reference, typically from scRNA-seq; when that reference is external to the target bulk cohort or differs in processing, protocol, or disease context, the mismatch reduces accuracy, especially for closely related cell states. Paired bulk and scRNA-seq data avoid this by directly connecting the target bulk assay to matched single-cell-derived composition, but large-scale paired profiling is impractical, so in realistic studies paired data are available only for a limited pilot subset of donors. This creates a challenge for deep deconvolution, which learns flexible bulk-to-composition mappings but needs enough labeled examples: the usual fix, training on pseudo-bulks generated from the reference, does not exploit the real paired bulk profiles a pilot design provides. We address this with an anchored data-augmentation strategy: from the matched in-study scRNA-seq we generate pseudo-bulks to expand the number and diversity of labeled examples, and we retain the real paired bulks in training with their matched composition labels, anchoring the model to the target bulk RNA-seq distribution. This converts a small paired pilot into an expanded supervised training resource. On a paired bronchoalveolar-lavage benchmark, the strategy makes a standard deep deconvolver the best method at every cell-type resolution, ahead of the strongest classical method and standard pseudo-bulk training, and controlled experiments show the real-bulk anchors drive the gain. Code is available at https://anonymous.4open.science/r/2dRNA-512F.

URL: https://openreview.net/forum?id=FG7PlICeut

---

Title: Structure Learning in Graphical Models from Indirect Observations

Abstract: This paper considers learning of the graphical structure of a $p$-dimensional random vector $\mathbf{X} \in \mathbb{R}^p$ using both parametric and non-parametric methods. Unlike the previous works which observe $\boldsymbol{x}$ directly, we consider the indirect observation scenario in which samples $\boldsymbol{y}$ are collected via a sensing matrix $\mathbf{A} \in \mathbb{R}^{d\times p}$, and corrupted with some additive noise $\mathbf{w}$, i.e, $\mathbf{Y} = \mathbf{A}\mathbf{X} + \mathbf{W}$. For the parametric method, we assume $\mathbf{X}$ to be Gaussian, i.e., $\boldsymbol{x} \in \mathbb{R}^p\sim \mathcal{N}\left(\mathbf{\mu}, \mathbf{\Sigma}\right)$, $\mathbf{\mu} \in \mathbb{R}^p$, and $\mathbf{\Sigma} \in \mathbb{R}^{p\times p}$. For the first time, we show that the correct graphical structure can be correctly recovered under the indefinite sensing system ($d < p$) using insufficient samples ($n < p$). In particular, we show that for the exact recovery, we require dimension $d = \Omega(p^{0.8})$ and sample number $n = \Omega(p^{0.8}\log^3 p)$. For the nonparametric method, we assume a nonparanormal distribution for $\mathbf{X}$ rather than Gaussian. Under mild conditions, we show that our graph-structure estimator can obtain the correct structure. We derive the minimum sample number $n$ and dimension $d$ as $n\gtrsim (\textup{deg})^4 \log^4 n$ and $d \gtrsim p + (\text{deg}\cdot\log(d-p))^{\beta/4})$, respectively, where $\textup{deg}$ is the maximum Markov blanket in the graphical model and $\beta > 0$ is some fixed positive constant. Additionally, we obtain a non-asymptotic uniform bound on the estimation error of the CDF of $\mathbf{X}$ from indirect
observations with inexact knowledge of the noise distribution. To the best of our knowledge, this bound is derived for the first time and may serve as an independent interest. Numerical experiments on both real-world and synthetic data are provided confirm the theoretical results.

URL: https://openreview.net/forum?id=F6q5LCCeVZ

---

Title: Differentially Private Data-Driven Markov Chain Modeling

Abstract: Markov chains model a wide range of user behaviors. However, generating accurate Markov chain models requires substantial user data, and sharing these models without privacy protections may reveal sensitive information about the underlying user data. We introduce a method for protecting user data used to formulate a Markov chain model. First, we develop a method for privatizing database queries whose outputs are elements of the unit simplex, and we prove that this method is differentially private. We quantify its accuracy by bounding the expected KL divergence between private and non-private queries. We extend this method to privatize stochastic matrices whose rows are each a simplex-valued query of a database, which includes data-driven Markov chain models. To assess their accuracy, we analytically bound the change in the stationary distribution and the change in the convergence rate between a non-private Markov chain model and its private form. Simulations show that under a typical privacy implementation, our method yields less than $2\%$ error in the stationary distribution, indicating that our approach to private modeling faithfully captures the behavior of the systems we study.

URL: https://openreview.net/forum?id=FRLrjh0Bzf

---

Title: When Does Data Value Reduce to Class Balance? A Coverage View of Per-Point Data Valuation

Abstract: Per-point data-valuation scores — Data Shapley, influence, leverage — are routinely used to select training data by taking the top k. We show that fixed, mode-blind per-point rankings are provably suboptimal as selectors in an identifiable latent-mode regime, and characterize that regime exactly. A point's value is not a scalar but a context function g_i(S) = R(S) − R(S ∪ {i}); for a kernel learner it is the point's feature novelty after orthogonalizing against the selected set, so the interaction lives in the off-diagonal (redundant) Gram. In an equal-mode model we prove that any symmetric valuation, used as top-k, has expected excess risk exactly ρ·δ(1,k) over greedy coverage at a binding budget — a fixed scalar cannot place one point per mode, so it covers only an occupancy fraction. The decisive question is whether the learner's modes coincide with the labels available at selection. If yes (standard classification on strong features), a one-line class-balance of the same scores recovers and beats coverage — the gap was just class imbalance. If no, the correction is coverage, not balancing: with no labels, label-free coverage wins by up to +0.20 accuracy across datasets (by several standard errors) while a raw top-k falls below random; with labels coarser than the modes, scalar balancing is insufficient and the best practical selector uses the coarse labels as strata and applies learner-geometry coverage within them. The resulting diagnostic is simple: when the labels available at selection align with the learner's modes, balance within labels; otherwise add an explicit learner-geometry coverage or diversity correction. The advantage is representation-relative — when the modes are not expressed in the learner geometry (spurious subgroups), coverage too fails, exactly as the theory predicts.

URL: https://openreview.net/forum?id=AZExWUZRPd

---

Title: Equal Accuracy, Unequal Cost: Channel Dependence in PatchTST under Controlled Coupling

Abstract: Channel-independent (CI) and channel-dependent (CD) variants of PatchTST are typically compared on observational benchmarks where correlation strength, dimensionality, and temporal dynamics all vary at once, leaving no clean way to attribute either variant win to any single cause. We isolate one factor at a time. A factorial AR(1) experiment varies contemporaneous correlation $\rho$ and variate count $C$ while Granger non-causality holds across channels by construction, and a leader-follower VAR(1) experiment introduces lag-1 coupling of controllable strength $\gamma$. On the synthetic grid we detect no consistent CD advantage at the 1% level: the grand-mean CD$-$CI difference is $+0.0013$ MSE (95% CI $[-0.0002, +0.0028]$) over five seeds, with a pooled within-cell detection half-width of 0.0063 MSE, so any global CD gain (averaged across the nine AR(1) cells) above roughly 0.63% of the CI mean would have moved the interval off zero, and none did. Under the original early-stopping protocol an apparent CD deficit emerges under lagged coupling and grows with $\gamma$, but an instrumented single-trajectory diagnostic at the leader-follower $\gamma=0.6$ cell shows this deficit to be largely a selection artefact rather than an architectural failure: CD reaches a sharper, earlier validation minimum than CI and degrades past it, and selecting each model at its own validation minimum erases the gap (CD$-$CI of $-0.0031$ MSE, 95% CI $[-0.0109, +0.0046]$). A cross-variate prediction head ties CI at every $\gamma$ and removes the coupling-dependent widening, locating that widening in the shared head rather than the encoder. Matching the gradient-update budget across modes removes the deficit, which rules out the per-epoch step-count asymmetry as its source, and a block-covariance family carries the same no-advantage finding beyond compound-symmetry correlation. Across every tested synthetic cell (AR(1), leader-follower, and block-covariance) the absolute CD$-$CI difference stays inside a pre-registered 1% relative-MSE band in both directions. What does not equalise is cost: CD requires smaller feasible batches and far more gradient steps per epoch, its batch size collapses to one at $C=84$, and at $C=321$ (ECL) the flattened-token formulation does not finish under a single-T4 budget while CI trains stably. On the observational ETTh1 ($C=7$), a matched-budget comparison still gives CI a modest but reliable edge (pooled CD$-$CI of $+0.0147$ MSE, 95% CI $[+0.0047, +0.0246]$, significant at horizons 96 and 720). On the constructed AR(1), leader-follower, and block-covariance processes, neither mode shows an accuracy gap that clears our 1% bar once each is selected fairly, and CD stays substantially more expensive to train. CI is the better default on accuracy per unit compute. All findings are scoped to PatchTST's flattened-token CD formulation, the generative families studied, and a single-T4 training budget.

URL: https://openreview.net/forum?id=aiUZ2y8UNl

---

Title: Backdoor Removal by Task Negation in Weight Space

Abstract: Foundation models have revolutionized computer vision by enabling broad generalization across diverse tasks. Yet, they remain highly susceptible to targeted backdoor attacks. Mitigating such vulnerabilities remains an open challenge, especially given that the large-scale nature of these models prohibits retraining to ensure safety. Existing backdoor removal approaches rely on costly fine-tuning to override the harmful behavior, and can often degrade performance on other unrelated tasks. This raises the question of whether backdoors can be removed without compromising the general capabilities of the models. In this work, we address this question and study how backdoors are encoded in the model weight space, and show that they are approximately disentangled from other benign tasks. Specifically, this separation enables the isolation and erasure of the backdoor's influence on the model with minimal impact on clean performance. Building on this insight, we introduce a simple post-hoc removal method that leverages such disentanglement. Through extensive experiments with CLIP-based models and common adversarial triggers, we show that, given a small triggered set, our method reduces attack success rate by over 99% while retaining, on average, 96% of the model's clean accuracy, using less than 2% of the data required by clean-data fine-tuning defenses. We further find that trigger vectors estimated on one dataset transfer to backdoored models trained on other datasets and even other label spaces, indicating that standard attacks induce a consistent direction in weight space. Additionally, we demonstrate that even when the attack and its presence are unknown, our method can be paired with reverse-engineered triggers to mitigate the effects of the backdoor. Overall, our method offers favorable removal and clean-accuracy tradeoffs compared to state-of-the-art clean data defenses, at a fraction of their data cost.

URL: https://openreview.net/forum?id=6apd7HU3y0

---

Title: Density-Weighted Minimax Q-Learning with ANN Retrieval: Aligning Retrieval Geometry to Policy-Mass Recall

Abstract: Approximate nearest-neighbor (ANN) retrieval is typically treated as a system-level detail in large-action offline reinforcement learning. This study treats it as a statistical problem: retrieval truncates the support of the target policy, biasing Bellman updates in proportion to the missing policy mass. It formalizes this as policy-mass recall (PMR@k), proves a value-error bound linear in missing mass, and introduces a retrieval-projection firewall (B2) that aligns ANN index geometry to the Q-ranking without allowing the alignment loss to perturb the critic. Experiments on MovieLens-25M (59,047 items, 25M ratings) show that three baselines (Naive ANN-DQN, CQL, BCQ) collapse to near zero ANN recall (0.024– 0.034, majority of seeds at exactly 0.000), while DW-MQL with the B2 firewall achieves 0.960 ± 0.014 across 10 seeds. A component ablation isolates the mechanism: retaining B2 while removing density correction preserves high recall (0.951 ± 0.038), whereas removing B2 collapses recall to 0.033 and Wolpertinger behavior-cloning retrieval gives 0.025. Thus the listwise-KL projection, not density correction alone, drives the retrieval advantage. A second-dataset experiment on KuaiRand-Pure (7,583 short-video items, binary click feed-back) replicates the pattern: DW-MQL achieves 0.885 ± 0.059 while all baselines collapse below 0.003. These experiments report diagnostic retrieval-support metrics rather than recommender benchmark performance.

URL: https://openreview.net/forum?id=ic14CBkXzK

---

Title: Diversity, Veracity, Independence: Diagnosing and Reversing Model Collapse

Abstract: Training a generative model on data produced by earlier models degrades it---model collapse. The
field frames this as telling synthetic data apart from real data, but that distinction is ill-defined and
undetectable at scale. We instead treat collapse as a loss of distributional integrity, measured on
three provenance-free axes: Diversity (are the outputs still varied?), Veracity (are they
still valid?), and Independence (does each new sample carry fresh information, or merely echo earlier
ones?). A single error law ties collapse to all three axes at once, and the data-processing inequality forces
the Independence term to shrink under repeated self-training. We prove this in tractable
(discrete and Gaussian) settings and show it offers a unifying lens on several prior results---casting
Strong Model Collapse's synthetic-data residual as a Veracity (distribution-mismatch) effect,
complementary to the distinct Independence residual our own decomposition isolates. Building on this, we give a
theory of recovering a collapsed model, by analogy with database recovery: a model can be restored only
against a durable, append-only log of real data, and there is a sharp point of no return once lost
content leaves that log. We provide a model-agnostic metrics suite and validate the framework on recursive
loops in text, image, and audio, as well as on the original model-collapse authors' own Gaussian-mixture and
OPT-125M setups. Two predictions from prior work that our diversity sweep first appeared to miss---the
golden-ratio optimum (recovered as a self-consistency check of the corrected weighted-estimator statement) and
recovery hysteresis---are borne out once tested against the correct knob and
objective: the optimal real-data weight}is $1/\varphi$ at matched sample sizes, and a continuous
recursion exhibits a recovery-hysteresis loop that persists as the ramp slows. The provable core is linear/discrete;
deep-network behavior is stated as falsifiable conjectures and measured.

URL: https://openreview.net/forum?id=MuehunvTEe

---

Title: The Landscape of Agentic Time Series Systems: Architectures, Reliability, and Frontiers

Abstract: Time series analysis is moving beyond the classical goal of accurate forecasting toward interpretation, verification, and decision-making in dynamic environments. Time series foundation models improve cross-domain generalization through reusable temporal representations, while Large Language Models for Time Series (LLM4TS) introduce language interfaces, multimodal alignment, and explicit reasoning. However, most systems remain model-centric, mapping observations to outputs without a closed loop over evidence acquisition, tool or action selection, feedback, and state updates. This survey presents a systematic landscape of agentic time series systems, which treat time series analysis as closed-loop interaction with temporal environments. We trace the progression from foundation models to LLM-based translators, temporal reasoners, and closed-loop agents, and define a time series agent as an operational system that observes temporal evidence, selects tools or actions over evolving states, receives feedback, and adapts subsequent behavior. We organize the literature through five capability layers: perception, reasoning, planning and action, memory and knowledge, and temporal world models. This framework serves as a capability map rather than a mandatory checklist. We further treat benchmarks as evolving evaluation infrastructure, considering reliability and trustworthiness as cross-cutting requirements that span predictive quality, reasoning faithfulness, tool use, grounding, robustness, safety, auditability, and reproducibility. Building on an extensive literature corpus, we synthesize representative methods, benchmarks, applications, and research frontiers, and argue for reliable, auditable, and adaptive agents that can analyze, decide, learn, and act under dynamic uncertainty.

URL: https://openreview.net/forum?id=oS3A7GZg2i

---

Title: Teaching Diffusion to Speculate Left-to-Right

Abstract: Speculative decoding accelerates large language model inference by having a lightweight drafter propose blocks of candidate tokens that a larger target verifies in a parallel pass under a rejection-sampling contract that preserves the target's output distribution exactly. Block-diffusion drafters are attractive in this setting because they emit an entire B-token block in one non-autoregressive pass, but they are trained with fully bidirectional attention within the block, whereas verification proceeds strictly left-to-right and truncates the block at the first rejection. We show that this training–verification mismatch is empirically consequential: on a position-uniform block-diffusion baseline, an average of 46.9% of draft tokens that match the target's greedy output are discarded as a downstream consequence of an earlier within-block rejection. We then analyze three complementary training-time interventions that reshape the drafter's objective along orthogonal axes: position-wise loss decay, a first-error focal term targeting the block's chain-breaking position, and a chain reward that substitutes a differentiable surrogate for the expected accepted length. We show that they compose additively, add negligible compute, and require no change to the drafter architecture, the inference pipeline, or the exactness contract. Across four instruction-tuned targets (Llama-3.2-3B, Llama-3-8B, Qwen3-4B, Qwen3-8B) and six reasoning, code, and dialogue benchmarks, the fully stacked configuration raises average accepted draft length by up to +43.9% over the position-uniform baseline and by up to +75.9% on individual benchmarks. The interventions further compound when composed with target-aligned training data, varied speculation horizons, tree-based verification, and the pathwise streak-distillation objective of SpecDiff-2, indicating that they operate along an axis distinct from data, horizon, and verifier-side alignment. Source code, training scripts, and trained drafter checkpoints will be released upon acceptance.

URL: https://openreview.net/forum?id=49AZSjhHnD

---

Title: Continual Learning Algorithms

Abstract: Continual learning studies how a model can keep acquiring knowledge from a non-stationary stream of data without retraining from scratch. The field has produced a large and fast-growing catalogue of algorithms, yet it is often unclear what they share, what problem they solve, and how much of it has been solved—a confusion that the new vocabulary introduced by recent language-model methods has only deepened. This review argues that most continual learning algorithms pursue the same objective: approximately minimizing the historical risk, the average loss over all data observed so far, under a limited computational budget. They differ mainly in how they approximate the past—through a surrogate term added to the objective, or a trust region that constrains the update—and we use this distinction to organize the literature, classical and recent alike. We complement this taxonomy with an accounting of computational cost that separates the cost of each update from the cost of consolidating the algorithm’s memory. Under this accounting, no surveyed method keeps its excess risk bounded while paying a cost that grows sub-linearly with the history: a closer approximation of the past is always paid for with more computation, memory, or model capacity. Reading recent foundation-model adaptation methods through the same lens shows that they largely reinstantiate classical mechanisms—replay, distillation, parameter isolation, low-rank constraints—under new names: the toolkit has been modernized, but what is computationally achievable has, so far, not changed. We distill from this analysis the open problems whose solution, we argue, would constitute real progress in the field.

URL: https://openreview.net/forum?id=BpotEzxSsJ

---

Title: Naïve Difficulty Conditioning Fails on ShapeNet-55: A Matched-Control Audit

Abstract: We audit a naïve difficulty-conditioned extension of AdaPoinTr and show that the apparent +14%–21% per-bin CD-$\ell_2$ gains we observed for it on ShapeNet-55 are evaluation-discipline artefacts rather than method effects. Warm-starting from a fully-trained checkpoint inflates apparent per-bin gains across Simple/Moderate/Hard irrespective of method, across two independent training runs (a pilot recorded per-bin at the time and a seed-matched rerun whose aggregate is released in the bundle)—an inflation visible only under a matched warm-start control. We propose a stratified matched-control protocol that pins warm-start, schedule, crop path, and pre-specified decision thresholds. Applied as a bounded case study to a natural DiffCond extension, the protocol uncovers a 5.48% Hard and 4.56% Mean CD-$\ell_2$ regression across three seeds, missing the pre-specified acceptance threshold on every bin. Four diagnostic probes localise the failure to a dead ranking-delta branch (architecturally non-differentiable) and to an auxiliary cross-entropy loss as the largest identified contributor. A data-side probe (Hard-Specialist) gives an exploratory point estimate of +7.38% ± 3.82pp on Hard across three seeds, with a 95% CI that crosses zero; a broader Easy-side probe gives no Simple-bin signal. We report this contrast descriptively, not as a tested general principle. A cross-backbone replication on UpTrans (4 seeds) yields no (mode, metric) cell clearing the pre-specified +3% method-bar (largest paired mean +1.79%, 95% CI [-2.72%, +6.29%]); a 2-seed PCN sanity check (underpowered) holds all deltas within ~1.1%. Under matched control, neither backbone reproduces the apparent double-digit signature: warm-start mismatch, not difficulty conditioning, drove the original gain.

URL: https://openreview.net/forum?id=cpe5PHAgoQ

---

Title: A Theoretical Analysis of Shallow Transformers on Noisy In-Context Recall Tasks: Optimality, Training Dynamics and Generalization

Abstract: We study the approximation capabilities, convergence speeds and on-convergence behaviors of $\textit{re-parameterized}$ one-layer decoder-only transformers trained on in-sentence-context recall tasks, which requires to recognize the association between a pair of tokens from in-context examples.
Existing theoretical results have not sufficiently addressed the on-convergence behavior of transformers being trained by gradient descent and how fast the convergence rate is. In addition, the generalization of transformers for in-context recall has not been formally investigated. This work addresses these gaps by first showing that a class of parameterized transformers with either linear, ReLU or softmax attentions, is provably Bayes-optimal for an in-context recall task. When being trained with gradient descent, we show via a finite-sample analysis that the expected loss converges at linear rate to the Bayes risks. Moreover, we show that the trained transformers exhibit out-of-distribution (OOD) generalization, i.e., generalizing to samples outside of the population distribution. Our theoretical findings are further supported by extensive empirical validations, showing that $\textit{without}$ proper re-parameterization, standard parameterized one-layer transformers surprisingly $\textit{fail}$ to generalize OOD after being trained by gradient descent.

URL: https://openreview.net/forum?id=wobDwqf2JB

---

Title: The Open-set Push-Pull Loss: A Revisit of Contrastive Learning for Open-set Supervised Anomaly Detection

Abstract: Open-set Anomaly detection (OSAD) can be formulated as a binary classification problem under severe class imbalance and limited knowledge about anomalies. Despite being increasingly used in classification problems, the current formulations of contrastive learning are not well-suited for OSAD. In this work, we address OSAD through a margin-based contrastive loss, the Open-set Push-Pull (OPP) loss. It considers only normal samples as anchors, aggregates them with positives via an attraction term, and pushes anomalies away via a repulsion term. We show that by minimizing OPP, we are equivalently maximizing a lower bound on the anomaly-normal margin on the hypersphere. We also quantify the bounds of this approximation and derive a gradient analysis, which clarifies the role of the temperature $\tau$. Building on this formulation, we propose two variants: (i) the $\tau$-balanced OPP loss that uses two distinct temperatures for the normals and anomalies, directly calibrated through the bound to account for class imbalance, (ii) a regularized OPP loss that saturates attraction and repulsion term to mitigate overfitting to training anomalies. We evaluate different contrastive objectives and our formulation across multiple open-set regimes on the downstream task of AD, on two industrial benchmarks MVTecAD, RealIAD, and finally Mastcam. We further compare against other state-of-the-art supervised AD architectures. Our results show that the OPP loss yields better robustness to increasing openness than standard supervised contrastive objectives and shows promising results when compared to other OSAD methods.

URL: https://openreview.net/forum?id=rz3RNyrzlQ

---

Title: The Cost of Fewer CI Tests in PC-Style Learning

Abstract: PC-style structure learning is often accelerated by reducing conditional-independence (CI) tests. This objective has a hidden cost: the graph information used to skip tests must itself be computed, and aggressive pruning can force larger conditioning sets. The Path-Driven Independence Testing (PIT) family exposes this trade-off. PIT restricts separator search by reachability in the partially learned skeleton; BPIT and Opt-BPIT add a blind-blocking rule that fixes variables on long paths. Under Markovness, faithfulness, and an exact CI oracle, all three algorithms recover the correct Markov-equivalence class and use no more CI tests than PC. Simulations show broad runtime gains for PIT with competitive accuracy, while further pruning traces the cost curve.

URL: https://openreview.net/forum?id=gVgDh59EDS

---

Title: Controllable EEG Editing Via Guided Diffusion Sampling For Affect Transposition

Abstract: In this paper, we study single-trial EEG editing with diffusion models, using affect transposition as a case study: given a real seed signal and a target valence label, we generate an edited signal that shifts the recognized affect while preserving affect-unrelated brain activity patterns. For this purpose, we propose a controllable EEG-to-EEG diffusion framework that performs seed-conditioned generation and supports various control mechanisms, to balance the tradeoff between affect transposition strength and seed signal fidelity. In our evaluation protocol, we test for editing effectiveness, seed fidelity, and physiological plausibility; on affective EEG (DEAP dataset), we quantify successful affect transposition with a valence classifier while measuring spectral-temporal deviation (e.g., STFT-based distances) from the seed. We probe edit plausibility with Frontal Alpha Asymmetry (FAA) measurements, which is a known neurophysiological marker that can indicate high valence. Furthermore, we test the robustness of affect-unrelated signal preservation on out-of-distribution data on an SSVEP (steady-state visually evoked potential) task using FBCCA (filter-bank canonical correlation analysis)-based frequency recognition. Our experiments indicate that classifier guidance towards the desired affective state combined with minimal noising of the input samples enables targeted affect edits with measurable limits regarding seed-fidelity and affect-unrelated feature degradation. These results suggest that single-trial EEG edits can be viably achieved utilizing our diffusion-based editing framework.

URL: https://openreview.net/forum?id=Cz1vhZp8Y6

---

Title: Selection Bias Inflates Per-Query Rerank-Augmentation Headroom: A Shuffle-Axis Null Calibration with Cross-Platform Evidence

Abstract: A common way to motivate adding a per-query routing or re-weighting signal on top of a strong reranker is to measure the oracle-direction headroom: for each query, pick the auxiliary axis (and its sign) that most improves the ranking, and report the average gain. We show that this quantity is systematically inflated by selection bias: the "pick-the-best-axis-per-query" operator returns a positive gain even when the axes carry no query-conditional information, because it is an order statistic over noise. We calibrate the diagnostic with a shuffle-axis null that permutes each query's axis values to destroy any axis-gold relation while preserving the selection operator, and report the net (observed minus null) with paired bootstrap confidence intervals. On the platforms we test, the surface headroom that would license a per-query router collapses to zero or below after calibration: the direction-level net is -0.0032 (95% CI [-0.017, +0.011]). On the primary Stage-B platform no single axis clears the pre-specified bar of 0.02; across the BGE/MonoT5 generalization cells (two rerankers, two benchmarks: LoCoMo, LongMemEval-S) the direction-level net remains at or below 0 or statistically indistinguishable from zero. A bounded fixed-axis exception appears for entity overlap on LoCoMo, but the direction-level router remains null there, so this does not support per-query routing. A controlled lexical ablation refutes a natural mechanism explanation for the bounded fixed-axis exception (entity overlap): it is not that the base scorer's lexical features internalize entity overlap. The same ablation reproduces a known effect from the reranking literature, where a weaker base leaves more apparent headroom on every axis; we cite it and use it to delimit the scope of any per-axis result. This is a methodological reporting requirement with a precise scope: it concerns per-query axis augmentation on top of a strong shortlist, and it makes the classical winner's-curse concern operational for reranking feasibility checks. The contribution is a concrete null that turns a misleading "oracle headroom exists" into a testable "net headroom after selection correction."

URL: https://openreview.net/forum?id=6yR2QUf51I

---

Title: Diffusion-assisted Matching for 6D Pose Estimation and Registration

Abstract: This paper presents DiffMatcher, a new diffusion-based framework for object-centric point cloud registration and pose estimation. The proposed pipeline estimates and refines a local and global geometric consistency-aware soft assignment matrix between source and target point clouds. The iterative refinement is driven by timestep-dependent spatial coordinate encodings and noise-level embeddings that enrich the backbone features and improve their discriminability across noise scales. As a result, the model first reconstructs coarse geometric structures in high-noise regimes and subsequently refines them with increasing precision over iterations. The framework includes a forward process during training, where noise is added to the ground-truth score matrix to guide denoising, and a reverse process during inference, where the model starts from a noisy score matrix and gradually recovers the target matrix. Experiments on real-world 6D object pose estimation, synthetic object-centric point cloud registration, and scene-level benchmarks show significant improvements in point matching and pose estimation, especially for low-overlap and noisy point clouds.

URL: https://openreview.net/forum?id=Y30q0tKBWT

---

Title: UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation

Abstract: Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks in evaluating how accurately generated images reflect the semantics of their textual prompt. However, (1) existing benchmarks lack the diversity of prompt scenarios and multilingual support, both essential for real-world applicability; (2) they offer only coarse evaluations across primary dimensions, covering a narrow range of sub-dimensions, and fall short in fine-grained sub-dimension assessment. To address these limitations, we introduce UniGenBench++, a unified semantic assessment benchmark for T2I generation. Specifically, it comprises 600 prompts organized hierarchically to ensure both coverage and efficiency: (1) it spans across diverse real-world scenarios, i.e., 5 main prompt themes and 20 subthemes; (2) comprehensively probes T2I models' semantic consistency over 10 primary and 27 sub evaluation criteria, with each prompt assessing multiple test points. To rigorously assess model robustness to variations in language and prompt length, we provide both English and Chinese versions of each prompt in short and long forms.
Leveraging general world knowledge and fine-grained image understanding capabilities of a closed-source Multi-modal Large Language Model (MLLM), i.e., Gemini-2.5-Pro, we develop an effective pipeline for reliable benchmark construction and streamlined model assessment. Moreover, to further facilitate community use, we train a robust evaluation model that enables offline assessment of T2I model outputs. Through comprehensive benchmarking of both open- and closed-source models, we systematically reveal their strengths and weaknesses across various aspects.

URL: https://openreview.net/forum?id=Id3q3eJw4M

---

Title: From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

Abstract: Large language model (LLM)-based agents are rapidly evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration. While these capabilities expand agent autonomy, they also make agent behavior increasingly difficult to verify, debug, and audit. The difficulty is that agents are evaluated almost entirely by their final answers, even though a correct answer reveals nothing about how an output was produced, what evidence supported each claim, whether tool calls were justified, how memory shaped later decisions, or where execution failures originated. To answer such questions, we look beyond outputs to the execution process itself, and structure it through evidence tracing and execution provenance. Provenance offers a process-level accountability layer for trustworthy LLM agents by modeling the connections among retrieved evidence, tool outputs, memory items, observations, intermediate claims, actions, and final answers. In this survey, we treat execution provenance as the full typed graph of an agent execution and evidence tracing as its projection onto evidence-support relations, giving a single framework that spans retrieval grounding through audit and recovery. We introduce a taxonomy that characterizes agent systems along six dimensions: trace sources, evidence and execution units, provenance relations, tracing granularity and timing, representation forms, and trust functions. We then review existing methods, categorized into seven main research threads: provenance representation, evidence attribution, tool-use provenance, runtime guardrails, provenance-bearing memory, observability, and failure diagnosis. Finally, we connect existing benchmarks, datasets, and metrics to provenance-related capabilities, discuss how evaluation can move beyond final-answer correctness toward process-level accountability, and outline open challenges for unified trace schemas, semantic provenance, provenance-aware safety, realistic execution-trace benchmarks, recovery-oriented evaluation, and privacy-aware audit infrastructure.

URL: https://openreview.net/forum?id=iwYk6keMAM

---

Title: How Much Quality Survives LLM Downshifting? Judge Disagreement and the Benchmark-Reality Gap in Routing Evaluation

Abstract: Teams cut large language model (LLM) inference cost by routing each request to the smallest model that is good enough, instead of always calling a single frontier model. To justify the switch they quote one headline number: "the small model preserves about X% of quality." We show that this headline cannot be trusted. We run the most aggressive policy possible, a static 70B->8B downshift within one open-weight family, over a 10,000-prompt curated benchmark spanning nine task types, and make four findings. (1) The cost reduction is large and judge-independent but price-assumption-dependent: token counts are essentially unchanged (total-token ratio 0.99), so projected cost falls by 98.35% ($162.79->$2.68) under frontier-tier list pricing and by 77-80% under commodity open-weight pricing; neither figure depends on any quality judge. (2) The quality "preservation" headline is strongly judge-dependent: re-scoring the identical response pairs with a ten-judge cross-family panel (all open-weight, 7B-67B, and out-of-family relative to the model under test) yields per-judge rates from 61% to 97% under the fixed-order protocol (median 89.9%, mean 85%, with 95% bootstrap CIs), with only fair panel agreement (Fleiss' kappa = 0.37). The single open judge behind the original headline, at 97%, sits at the top of this range, 15-23 points above the larger cross-family judges. (3) The headline is also strongly protocol-dependent: a blinded, order-counterbalanced re-judge of the same pairs with all ten judges finds that every judge favors whichever response is shown second, by 60-94 points, and that the two presentation orders agree on only 5-37% of verdicts. The midpoint of the two order-specific rates is 48-65% (median 55%), which lies 18-42 points below nine of the ten fixed-order rates; because most pairs are "preserved" in only one order, we read this as a midpoint, not an estimate of preservation. Under the blinded protocol the panel spread is 17 points versus 36 fixed-order, and the fixed-order strictness ranking is not reliably reproduced (Spearman rho = 0.29, n = 10), so both the original headline and much of the apparent judge disagreement were position artifacts. Absent a human anchor we still cannot say which level is correct. We can say that a genuine judge disagreement of about 16 points survives de-confounding, that fixed-order judge choice moves the headline by 15-23 points across the larger cross-family judges (up to 36 if the least-reliable panelist is included), and that presentation order moves it by up to 94 points on identical outputs. (4) The headline also fails on real chat traffic. On a 2,000-prompt sample of WildChat, two strict cross-family judges, Qwen-32B and Gemma-27B, rate the small model "equivalent-or-better" only 21-32% of the time, although it stays "usable" 92-95% of the time (91.8-95.2%), a fixed-order rate whose order-stability we did not measure. The curated benchmark over-represents multi-step reasoning by roughly 20x relative to real traffic, and the same judges scored the same downshift far lower on real traffic than on the benchmark. Aggregate single-number preservation claims for downshifting and routing should not be trusted, including an earlier single-judge claim of our own that this paper corrects. We do not claim any protocol here recovers the true rate; without a human anchor, none is validated. What practitioners should do instead is report a judge distribution, with intervals, under a blinded, order-counterbalanced protocol, and validate on real traffic. These quantities are diagnostic. Judge spread and cross-order agreement measure whether the evaluation is reliable enough to support an adoption decision at all, and in our data they show it is not. The data point to confidence-aware escalation, rather than blanket downshifting, as the policy to test next, though our single-judge frontier cannot yet show how much quality it recovers. We release the evaluation harness, prompt set, judge, cross-judge panel scripts, the blinded re-judge harness, and run logs for independent re-scoring.

URL: https://openreview.net/forum?id=6ccINAFoJi

---

Title: Generative Refinement Learning for Temporal Interpolation in Scientific Machine Learning

Abstract: Scientific temporal interpolation aims to recover intermediate physical fields from sparsely sampled simulations or observations. Existing video frame interpolation methods often rely on transport-based or deterministic-regression biases, which can be misaligned with complex systems exhibiting nonlinear, multi-scale, and uncertain dynamics. We propose a novel interpolative FLow EXpert (iFLEX), a generative framework that decomposes interpolation into numerical coarse interpolation and diffusion-based stochastic refinement learning. The coarse estimate captures endpoint-consistent low-frequency evolution, while the diffusion module learns nonlinear, high-frequency, and uncertain corrections. To condition the stochastic refinement process, iFLEX separately encodes high-dimensional physical fields through a shared spatial encoder and low-dimensional scalar variables through FiLM-based modulation. We evaluate iFLEX on simulated Navier-Stokes Kraichnan turbulence, real-world Shanghai radar precipitation, and sea surface temperature data, including various challenging out-of-distribution tests. Across these settings, iFLEX improves reconstruction accuracy and structural fidelity over deterministic, optical-flow-based, and diffusion-based video interpolation baselines, with especially strong gains for large-step interpolation and turbulent or intermittent dynamics.

URL: https://openreview.net/forum?id=esqLKv6O2P

---

Title: Calibrated, Falsifiable Detection of Sparse Mechanism Shift

Abstract: Many methods for distribution shift under causal structure assume the sparse mechanism shift (SMS) hypothesis: that across environments only a few causal conditionals change. This assumption drives mechanism-shift scoring, causal discovery in heterogeneous data, and transportable prediction, yet it is almost never tested on the data at hand. This paper makes SMS testable. We first ask why it is hard to tell which mechanisms changed once the causal graph must be estimated rather than assumed known. A controlled ablation locates the cause: the false positives that limit precision arise at the truly invariant nodes, because their parent sets are mis-estimated; a better global skeleton, repairing the changed nodes, and conditioning-set voting do not remove them. We then give a graph-free, label-free detector that flags a node only when no conditioning subset makes its conditional invariant across environments (an inverse use of invariant causal prediction); this is robust to the same failure mode and matches an oracle that knows the true graph. Building on it, we define a calibrated SMS hypothesis test: a sparsity statistic (the fraction of mechanisms flagged as changed) with a data-driven null floor obtained by splitting one environment in half, and a bootstrap three-way verdict (no shift / sparse / dense). On controlled synthetic data the verdict tracks the true sparsity; on real protein-signalling interventions it rejects SMS, and a paired atomic-versus-fat-hand study explains the rejection and predicts when SMS should hold.

URL: https://openreview.net/forum?id=6AGOgvdze1

---

Title: Comprehensive Evaluation of Similarity-Based Positional Encoding for Medical Vision Transformers

Abstract: Similarity-based positional encoding (SimPE) for Vision Transformers has been shown to improve classification accuracy on medical imaging benchmarks. However, accuracy alone is insufficient to characterise the practical value of a positional encoding strategy for clinical applications. This paper provides a comprehensive evaluation of SimPE addressing three open questions: (i)~\emph{how SimPE affects the structure of the learned feature representations}; (ii)~\emph{whether the accuracy advantage extends to other classification metrics}; (iii)~\emph{what is the computational cost of SimPE}. We show through UMAP dimensionality reduction that SimPE induces more discriminative class clusters than standard learned positional encoding (SLPE) and Rotary Position Embedding (RoPE) across all six tested modalities. Ten-fold cross-validation with 95\% confidence intervals on accuracy, precision, recall, and F1 score confirms that SimPE outperforms alternative approaches consistently, with narrow confidence intervals indicating stable training dynamics. Per-epoch training time and total-convergence analysis show that SimPE is only marginally slower than SLPE per epoch but converges in fewer epochs, usually yielding shorter wall-clock training time; RoPE is substantially more expensive in both dimensions. Attention map analysis further confirms that SimPE focuses on clinically relevant image structures more reliably than the alternatives. Finally, SimPE extends naturally to volumetric (3D) medical images with consistent improvements. These results consolidate SimPE as a leading positional encoding strategy for geometry-structured medical image classification.

URL: https://openreview.net/forum?id=qg87bPv2Wt

---

Title: Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization

Abstract: Deep learning problems rarely involve objectives that are equal in importance. A primary objective defines the goal, whilst secondary objectives, such as sparsity, compression, or robustness constrain the solution. While existing multi-objective methods have proven effective in practice, they have a clear symmetry problem and neglect the inherent objective hierarchy built into these objective spaces. We introduce Priority-Constrained Descent (PCD), a gradient-based optimization framework designed to explicitly exploit hierarchical objective structures. PCD preserves the direction of primary descent whilst allowing for the minimal distortion necessary to guarantee progress on secondary objectives, controlled by a single $\tau \in [0,1]$ that dictates the strength of the distortion. The resulting formulation is invariant to objective scaling and admits exact closed-form solutions for problems with two and three objectives. We evaluate PCD within structured network compression settings, unstructured sparsity and low-rankness, and across a variety of synthetic experiments, showing Pareto dominance and better per-objective performance with secondary progress guarantees over existing methods, further exhibiting the interpretable trade-off that $\tau$ provides.

URL: https://openreview.net/forum?id=HT01yGHLEt

---

Title: Calibrating LLMs for Selective Prediction: Balancing Coverage and Risk

Abstract: Despite the impressive capabilities of large language models (LLMs), their outputs often exhibit inconsistent correctness and unreliable factual accuracy. In high-stakes domains, overconfident yet incorrect predictions can lead to serious consequences, highlighting the need for robust uncertainty estimation. To address this, we introduce SelectLLM, an end-to-end method designed to enhance the ability of LLMs to recognize and express uncertainty effectively. By integrating selective prediction into finetuning, SelectLLM optimizes model performance over the covered domain, achieving a more balanced trade-off between predictive coverage and utility. Experimental results on TriviaQA, CommonsenseQA and MedConceptsQA show that SelectLLM significantly outperforms standard baselines, improving abstention behaviour while maintaining high accuracy.

URL: https://openreview.net/forum?id=WkE6Fw62Pj

---

Title: Majority-of-Three is Optimal

Abstract: We give a short proof that the majority vote of three independent consistent classifiers is an optimal learner in the realizable PAC setting. This proves optimality for the simplest voting scheme, while simplifying both the algorithmic structure and the probabilistic analysis of previous voting learners, including the optimal PAC learner of Hanneke and Green Larsen's analysis of bagging. In particular, this proves the conjecture of Aden-Ali et al.

URL: https://openreview.net/forum?id=bCj8madmyb

---

Title: InkFormer: From Characters to Signatures via Progressive Style-Content Decoupled Online Handwriting Generation

Abstract: Online handwritten signatures provide dynamic signing trajectories and serve as an important biometric trait for identity authentication. However, online signature verification remains constrained by the limited scale and diversity of training data, particularly when forgery samples are unavailable or only a single genuine signature is accessible. Although generative modeling can alleviate data scarcity, existing online handwriting generators remain difficult to adapt to signature synthesis due to limited style controllability and insufficient multi-character structure modeling.
Since signatures can be regarded as writer-specific online handwriting trajectories, we rethink signature synthesis as controllable online handwriting generation and propose $\textbf{InkFormer}$, an auto-regressive framework for style-controllable online handwriting generation. InkFormer includes three key designs: self-constructed online standard characters for modality-consistent content guidance, feature-level content--style residuals for explicit writer-style modeling, and a style--content-aware learnable starting token for auto-regressive decoding. With content--style cross-attention during generation, InkFormer produces high-fidelity handwriting trajectories with accurate structures and consistent styles. Building upon InkFormer, we develop a progressive signature generation framework that transfers model capability from character-level to word-level generation, and finally to online signature synthesis. The model first learns character structures and writer styles, then strengthens multi-character and inter-character modeling through word-level fine-tuning. Finally, real signatures are used as style references and signature contents as prompts to generate online signature trajectories for verification training. Experiments show that InkFormer achieves superior online handwriting generation quality over existing methods. Moreover, as a data augmentation engine for signature verification, the generated signatures significantly improve verification performance under forgery-free and one-shot genuine low-resource scenarios. The source code will be publicly available.

URL: https://openreview.net/forum?id=Nbhw2Y5Nmh

---

Title: Revisiting Class-Incremental Learning in the Era of Foundation Models: A Gaussian Density Estimation Perspective

Abstract: Class-Incremental Learning (CIL) is commonly addressed by preventing catastrophic forgetting through regularization, replay buffers, or architectural adaptations. In this work, we revisit this formulation in the context of modern visual foundation models. We argue that once representations are frozen and sufficiently discriminative, the main challenge of CIL shifts from protecting model parameters to modeling class distributions in a stationary feature space. Building on this observation, we formulate incremental learning as recursive estimation of class-conditional Gaussian densities in foundation embeddings. To handle data scarcity, we introduce the Hybrid-Cov estimator, which decouples the Mahalanobis shape term from the log-normalization constant, providing a numerically stable alternative when the sample budget is small. We validate Gaussianity rigorously using formal statistical tests (D'Agostino-Pearson and Henze-Zirkler) and quantify inter-class separability via the Bhattacharyya distance, which provides a direct connection to the Bayes error rate. Despite its simplicity, the proposed approach achieves competitive performance on CIFAR-100, ImageNet-R, and Caltech-256, while requiring significantly lower computational cost and providing a principled probabilistic interpretation.

URL: https://openreview.net/forum?id=k3uhZV6lfU

---

Title: POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration

Abstract: Reinforcement learning (RL) has improved the reasoning abilities of large language models (LLMs), yet state-of-the-art methods still fail to learn on many training problems. On hard problems, on-policy RL rarely explores even a single correct rollout, yielding zero reward and no learning signal. We find that natural solutions to remedy this exploration problem from classical RL, such as entropy bonuses, more permissive clipping of the importance ratio, or direct optimization of pass@k objectives, do not resolve this issue and often destabilize optimization without improving solvability. A natural alternative is to leverage transfer from easier problems. However, we show that mixing easy and hard problems during RL training is counterproductive due to ray interference, where optimization focuses on already-solvable problems in a way that actively inhibits progress on harder ones. To address this challenge, we introduce Privileged On-Policy Exploration (POPE), an approach that leverages human- or other oracle solutions as privileged information to guide exploration on hard problems, unlike methods that use oracle solutions as training targets (e.g., off-policy RL methods or warmstarting from SFT). POPE augments hard problems with prefixes of oracle solutions, enabling RL to obtain non-zero rewards during guided rollouts. Crucially, the resulting behaviors transfer back to the original, unguided problems through a synergy between instruction-following and reasoning. Empirically, POPE expands the set of solvable problems and substantially improves performance on challenging reasoning benchmarks.

URL: https://openreview.net/forum?id=pbyIoUW7YV

---

Title: CaRT: Teaching LLM Agents to Know When They Know Enough

Abstract: Many tasks require machine learning models to strategically gather relevant information over multiple rounds of interaction before actually acting on a task. Strategic information gathering requires models to know not only how to effectively acquire information, but also when to stop gathering information and make a decision, in order to avoid overthinking or getting derailed when acting. In this paper, we formalize this problem and introduce Counterfactuals and Reasoning for Termination (CaRT), an approach for teaching LLMs when to stop seeking information. To appropriately learn when to terminate, CaRT fine-tunes LLMs using counterfactual pairs of trajectories, one where termination is appropriate and a minimally modified version of the same trajectory where it is not. It trains the LLM to explain the rationale for the termination decision in either case via verbal reasoning, and imbues this capability into the base LLM via fine-tuning. We instantiate CaRT in two domains: interactive medical diagnosis and math problem solving. In both domains, we find that CaRT improves the efficiency of information gathering and task success rate compared to other fine-tuning methods.

URL: https://openreview.net/forum?id=hxGHnbzHqm

---

Title: One Spin at a Time: Sequential Subspace Rotations for Parameter-Efficient Fine-Tuning

Abstract: In this work, we introduce SOARA (Subspace Orthogonal Adaptation via Rotational Alignment), a novel family of parameter-efficient fine-tuning (PEFT) algorithms that navigate the manifold of pretrained weights through geometric alignment. While existing SVD-based methods prioritize magnitude-driven adaptation, for example by modifying singular values (SALT, SVFT) or reparameterizing basis vectors via additive low-rank updates (PiSSA), they often overlook the intrinsic rotational symmetries of the feature space.

In contrast, SOARA treats the pretrained principal singular vectors as a foundational, fixed basis and adapts the model by learning lightweight rotational transformations within these subspaces. This preserves the representational integrity of the pretrained features while allowing for precise alignment with downstream task geometries. We provide two distinct paths for enforcing orthogonal parametrization of rotational matrices: (1) parametrization via regularization, and (2) parameterizations via sequential Givens rotations or butterfly decompositions. By operating purely through rotations, SOARA avoids the representational collapse and shift in feature distribution frequently observed in additive or purely scale-based adaptation schemes.

Our empirical results on large-scale architectures, including ViT-B16 and DeBERTa-v3-base, demonstrate that SOARA matches or exceeds the performance of state-of-the-art PEFT methods across vision and language benchmarks. SOARA achieves these gains with a minimal parameter footprint, suggesting that rotational alignment provides a superior inductive bias for preserving and repurposing the complex feature manifolds of foundation models.

URL: https://openreview.net/forum?id=gvh5FQXr9w

---

Title: Storage-Equivalent Evaluation for Dense Retrieval: Placing Document Selection and Compression on a Common Per-Document Budget

Abstract: Dense-retrieval indexes are increasingly storage-bound. Two families of methods shrink them: compression spends fewer bits per document embedding, and selection (pruning) retains fewer documents. The field evaluates the two under incompatible protocols, compression over a fixed document set and selection at a fixed retention rate, and the gap is consequential: methods at equal retention can differ in per-document storage by up to $16\times$, so the two protocols can return opposite verdicts. We introduce storage-equivalent evaluation, which fixes a per-document bit budget $B$ and lets each method retain $m=\lfloor BN/b\rfloor$ documents, placing pruning, quantization, and dimensionality reduction on one bits-per-document axis. The protocol returns a two-part answer. At a fixed compression code, intelligent selection reliably beats random: significantly, on all four core corpora, at both binary and PQ48$\times$4. But choosing a cheaper code is the larger lever: across nine BEIR corpora and three encoders, a cheap product-quantization code beats every off-the-shelf selector measured at deployment budgets, and a closed-form dominance condition shows the disadvantage is structural, a selector must overcome the cheaper code's documents-per-bit advantage by document choice alone, before seeing the test queries, a bar learned selectors face no less. Among well-powered corpora selection beats compression on only one (BM25 selection on a financial corpus), but no measured corpus property (relevance density, intrinsic dimension, neighborhood similarity, lexical overlap, or the BM25-versus-dense retrieval gap) predicts where it helps.

URL: https://openreview.net/forum?id=AwYSPyXrAR

---

Title: Estimating Rare Events in Language Models with Proper Evaluation

Abstract: Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale deployments, requires estimating probabilities far too small for random sampling. While recent work has formalized Low Probability Estimation, existing pipelines remain fragile in the rarest regimes: estimators can suffer zero-estimate collapse or systematic bias, and standard evaluation losses can become unstable or poorly matched to asymmetric safety costs. In this work, we introduce Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models. Specifically, GA-AMLS uses a gradient-based MCMC kernel to navigate activation space, eliminating the zero-estimate collapse of input-space search and replacing the independence assumptions of prior activation-space estimators with conditional sampling under an explicit, heavier-tailed activation prior. We also propose the Shifted-Power Bregman (SPB) Loss, a proper scoring rule that remains finite for zero-estimates and offers tunable asymmetry between underestimation and overestimation penalties. Experiments on small transformer models reveal a bias-variance tradeoff: GA-AMLS achieves the lowest loss under symmetric evaluation, reducing average log-space squared error relative to the strongest baseline across model sizes, while methods with overestimation bias prevail under asymmetric penalties. Our findings highlight that estimator choice should be matched to deployment context. More broadly, our work establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discrete input-space search.

URL: https://openreview.net/forum?id=4AMA60efWX

---

Title: Harness Engineering for LLM Agents: A Survey of Harness Component Taxonomy, Evaluation, and Model--Harness Coevolution

Abstract: As large language model-driven autonomous agents are increasingly deployed in real-world long-horizon, open-environment tasks, foundation models expose systematic capability gaps in context retention, reliable tool invocation, persistent state management, and multi-step execution robustness. We understand the agent harness as the external execution support structure built around the model and treat it as a distinct performance lever that complements base model capability, positioning harness engineering as a growing area of research and engineering practice. Grounded in the scaffolding perspective from developmental psychology, we structure the field along three nested levels of analysis. At the structural level, we develop a unified taxonomy of harness components, mapping them to the specific capability gaps they compensate for and to their coupling with the core agent loop. At the fit level, we articulate a two-stage evaluation logic that distinguishes native model capability-gap diagnosis from assessments of compensation effectiveness and net benefit, and unpack the inherent multi-objective tradeoffs shaping harness design. At the dynamic level, we delineate the bidirectional coevolution mechanism between models and harnesses, explaining the shifting functional boundary where routine capability-bearing support migrates inward into model weights while constraint-bearing governance functions remain external. By synthesizing studies that are currently scattered across adjacent areas, we provide an organizing framework for understanding LLM agent harnesses in relation to the capability gaps they address. We further discuss open challenges and future directions in harness design, evaluation, and model--harness coevolution.

URL: https://openreview.net/forum?id=1VJLY0hAFT

---

Title: SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do. A standard mitigation hands control to a separate recovery policy whenever the agent leaves a designer-specified safe region (a subset of state space it should stay within), but the resulting mixed-policy rollouts silently bias every on-policy update, and the importance-sampling correction that would remove this bias is ill-defined whenever the recovery policy is deterministic. We address this bias with a drop-in modification of proximal policy optimization (PPO).
Its core is an unbiased policy-gradient estimator that uses the score function only at safe timesteps and never evaluates the recovery policy's density, so it stays valid even when the recovery policy is deterministic, exactly where importance sampling breaks, and it empirically dominates importance sampling even when the recovery policy is stochastic. Because the recovery policy still makes credit assignment slow near the safe-region boundary, two further components accelerate learning: a closed-form value for recovery-triggering states when dynamics and recovery are deterministic, and an imitation loss that copies recovery actions only when recovery succeeds.
On a three-environment, five-seed benchmark, the resulting algorithm reduces training-time falls by factors of $233\times$, $48\times$, and $26\times$ on HalfCheetah, Ant, and Unitree Go1 over standard PPO, while matching or exceeding PPO's final reward, and on Ant, where the recovery policy is unreliable, it is the only method that reaches $80\%$ of the best final reward.

URL: https://openreview.net/forum?id=ppOrbSsMM0

---

Title: LOCO: Local Orbit Consistency for Symmetry-Robust Neural Operators

Abstract: Known Lie symmetries specify how a partial differential equation (PDE) solution operator should transform under group actions, but finite supervised training data constrain this relationship only at observed input--output pairs. To enforce these symmetries without architectural rigidity, we introduce \textbf{LOCO} (\textbf{L}ocal \textbf{O}rbit \textbf{Co}nsistency), a lightweight and flexible training objective for neural operators. LOCO samples a small Lie-group transformation during training and explicitly penalizes the mismatch between the operator's prediction on the transformed input and the transformed prediction of the original input. By leveraging known input and output actions, LOCO leaves the underlying neural-operator architecture entirely unchanged and introduces zero computational overhead at inference time. We evaluate our method on a challenging periodic $64 \times 64$ two-dimensional Navier--Stokes Galilean benchmark. At a highly constrained data regime of 2\% labels, combining data augmentation with local orbit consistency ($\lambda_{\mathrm{orb}} = 0.30$) reduces the orbit out-of-distribution (OOD) relative $L_2$ error by 35.3\% and the equivariance defect by 46.4\% compared to a compute- and time-matched supervised augmentation baseline. Furthermore, a fixed lower-weight setting, $\lambda_{\mathrm{orb}} = 0.10$, yields consistent performance gains across varying label fractions and OOD boost radii. Finally, solver-level closure residuals remain negligible, and action-mismatch ablations confirm that these performance gains depend strictly on the alignment with the intended physical symmetry constraint.

URL: https://openreview.net/forum?id=AQIyD6jSSl

---

Title: On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning

Abstract: Diffusion autoencoders (DAs) are variants of diffusion models that use an input-dependent latent variable to capture representations alongside the diffusion process. The latent variable is typically used for tasks such as downstream classification, controllable generation, and interpolation. However, the latent variable has an effect on unconditional generation, with performance strongly dependent on how this variable is modelled and learnt and how it is used in the denoising process. Here, we draw a connection between DAs and a particular class of improved diffusion models---those that learn their forward (noising) process---to argue that DAs have unrealised potential as generative models. We develop a variant of DAs, termed diffusion models with $z$ (DMZ), to show that the adoption of certain design decisions, such as the choice of latent variables and conditioning method, can help realise this generative potential. Through extensive experiments and ablations, we show that DMZ enables more efficient modelling and generation with fewer denoising steps, as well as providing an effective latent representation for downstream tasks, such as classification and domain transfer.

URL: https://openreview.net/forum?id=KX0RAVVSWz

---

Title: Cascading Bandits: Minimizing Regret with Respect to the Optimal Ordering

Abstract: We study a variant of cascading bandits, motivated by applications such as cognitive radios and carousel-based recommendation systems, where early successes reduce sensing costs or enhance user experience. In this setting, a learner sequentially interacts with a set of $K$ arms over $T$ rounds, where each arm yields a binary reward governed by an unknown success probability. In each round, the learner selects an ordering of the arms and pulls them sequentially until the first success is observed or all the arms are exhausted. The objective is to minimize the cumulative number of arm pulls required to find a success. We define regret as the expected number of additional pulls compared to an optimal ordering of the arms. To minimize regret, we propose two algorithms, O-UCB and O-Greedy, which extend the classical UCB and Greedy algorithms to our setting. For O-UCB, we derive two regret bounds with respect to the time horizon $T$: a logarithmic bound of $O(\log T)$ that depends on suboptimality gaps, and a constant bound of $O(1)$ that additionally depends on a problem-dependent parameter $p_{\text{all}}$, which quantifies the extent of inherent exploration. Furthermore, we show that O-Greedy, despite lacking explicit exploration, also achieves $O(1)$ regret under the same conditions. Extensive experiments on both synthetic and real-world datasets corroborate our theoretical results and delineate the regimes where each algorithm is most effective.

URL: https://openreview.net/forum?id=aHvfCNSJtd

---

Title: Diagnosis as Percolation on an Evidence Graph: Root-Cause Recoverability in Network Fault Diagnosis

Abstract: Automated diagnosis under partial observation recurs across computing and engineering: an agent must infer an unobserved root cause from an incomplete set of symptoms, and sometimes gauge how far to trust that inference. We argue that random-graph theory supplies both a model and a set of operational primitives to solve this problem. We cast diagnosis as percolation on a partially observed evidence
graph, whose vertices are observations, entities, hypotheses and candidate causes and whose edges are derived deterministically from data; a diagnosis is then a connected explanatory subgraph, and recoverability corresponds to whether the evidence has percolated into one giant cluster rather than scattering into fragments. Our primary contribution is to network fault diagnosis, in the lineage of the PLUME (Pradhan et al., 2026) and PROBE (Henry et al., 2026) systems: we work the abstraction out in full on an ensemble of large language models (LLMs) that diagnoses faults from textualized 802.11 packet captures, where an ensemble-plus-reconciliation pipeline lifts weighted evidence F1 from ≈ 0.85 to ≈ 0.96 yet provides no calibrated notion of trust. Probing whether connectivity supplies that missing signal yields a cautionary result: post-hoc connectivity correlates with correctness (Spearman ρ up to 0.67), but under capture-clustered controls, degree-preserving negative-control nulls and multiple-comparison correction, no structural signal retains value beyond capture
size. Connectivity is thus a portable, theoretically grounded substrate for fault recoverability and root-cause localization, but not a reliable proxy for answer-level trust. A controlled benchmark recovers the predicted critical density ϕc ≈ 1/c, where a connectivity-aware personalized-PageRank localizer tracks a Bayesian oracle while degree heuristics collapse under decoy hubs; and studies in five further domains, with a substrate-independent threshold analysis, indicate that the percolation principle transfers beyond networking.

URL: https://openreview.net/forum?id=lzmSUYzSdN

---

Title: Uncertainty Quantification for Regression Using Proper Scoring Rules

Abstract: Quantifying predictive uncertainty is essential to enable reliable decision-making with machine learning models.
Recent work proposed a unified framework for uncertainty quantification (UQ) based on proper scoring rules, but has focused exclusively on classification. Extending these results to regression remains an open problem.
In this paper, we introduce a UQ framework for regression that applies to any proper scoring rule, including the CRPS, logarithmic, squared-error, and quadratic scores.
We derive closed-form expressions under practical parametric assumptions and show how to estimate them using an ensemble of predictors.
The framework allows separating and estimating aleatoric from epistemic types of predictive uncertainty, and it recovers popular, well-established variance- and entropy-based UQ measures as special instances.
Finally, a broad evaluation on synthetic and real-world datasets provides guidance for selecting reliable UQ measures in practice.

URL: https://openreview.net/forum?id=5i6XQ2wmxE

---

Title: On the modality gap and the contrastive loss in multi-modal representation learning

Abstract: We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a mode-failure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer. Code availability: {...}

URL: https://openreview.net/forum?id=P0fmivB4EH

---

Title: SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision–language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on large-scale datasets, explicit 3D inputs, architecture-specific modifications, or sparse Reinforcement Learning (RL) methods that provide insufficient guidance for spatially-grounded reasoning. We introduce \textsc{SpatialThinker}. To our knowledge, it is the first MLLM unifying Scene Graph Generation (SGG) and visual reasoning in a single pass via online RL. The model simulates human-like spatial perception by constructing a mental scene graph of task-relevant objects and relations, and reasoning toward an answer via dense spatial rewards. Our contributions are threefold: (1) SGG-grounded reasoning: integrating SGG directly within the reasoning chain rather than as a disjoint preprocessing step; (2) STVQA-7K: a high-quality spatial VQA training dataset via a scalable synthesis pipeline; and (3) a dense spatial reward design that enforces structured grounding during RL and generalizes to improve broad visual perception. \textsc{SpatialThinker-7B} achieves 3.6$\times$ larger gains over SFT and $1.7\times$ better in- and out-of-distribution generalization than sparse RL. Trained on only 7K samples, \textsc{SpatialThinker-7B} matches GPT-5 and outperforms GPT-4o, while \textsc{SpatialThinker-30B} surpasses both GPT-5 and Claude 4 Sonnet on average across 14 spatial and real-world benchmarks, demonstrating that structured spatial grounding with reward-aligned reasoning enables robust spatial understanding with limited data.

URL: https://openreview.net/forum?id=GrQud1eT6u

---

Title: Self-Distilling Mathematical Reasoning in Small Language Models

Abstract: Small base language models (0.5B–3B parameters) often fail to produce structured, scorable outputs in zero-shot mathematical reasoning, creating a cold-start regime in which reward- or preference-based optimization becomes ill-conditioned. Yet these same models achieve 80–90% scoreability under few-shot prompting, indicating that the underlying reasoning capability exists but is not expressed zero-shot.
We propose Advantage-Weighted Direct Preference Optimization (AWDPO), a self-distillation method that closes this gap by training models to reproduce their own few-shot reasoning behavior without prompts at inference time. AWDPO weights preference updates by the performance gap between outputs and stabilizes training with a dynamically scaled likelihood anchor. On in-domain GSM8K with Qwen-2.5 base models, AWDPO recovers over 90% of the performance of fully supervised chain-of-thought fine-tuning using only four chain-of-thought exemplars, approximately 1/1870th as many reasoning traces (4 vs. 7,473 for SFT). Notably, AWDPO zero-shot accuracy (77.6%) exceeds its few-shot prompting baseline (53.8%), suggesting that parameter-based access to reasoning can surpass prompt-based elicitation. The resulting models further generalize zero-shot to SVAMP, ASDiv, and MATH-500 without additional training.

URL: https://openreview.net/forum?id=dYdt8IFY7U

---

Title: Characterizing the Multiclass Learnability of Forgiving 0-1 Loss Functions

Abstract: In this paper we will give a characterization of the learnability of forgiving 0-1 loss functions in the multiclass setting with effectively finite cardinality of the output and label space. To do this, we create a new combinatorial dimension that is based off of the Natarajan Dimension (Natarajan, 1989) and we show that a hypothesis class is learnable in our setting if and only if this Generalized Natarajan Dimension is finite. We also show how this dimension characterizes other known learning settings such as a vast amount of instantiations of learning with set-valued feedback and a modified version of list learning.

URL: https://openreview.net/forum?id=ThRIaJ00Bu

---

Title: When Structure Doesn't Matter: Loss-Invariant Routing Specialisation in Hybrid Language Models

Abstract: We construct a controlled setting in which an architectural channel in a hybrid state-space/attention language model is free to organise its internal representations without measurably affecting the optimisation objective. The channel is "hormone routing": a token-conditional softmax-gated linear combination of a small bank of frozen direction vectors, added to the residual stream after each attention block. We pretrain a 125M-parameter Mamba/MQA hybrid on FineWeb-Edu under a fixed 2B-token budget across five variants of this channel --- extracted directions, random unit vectors, vectors extracted from a randomly initialised model, and a frozen gate at 1.0. The router develops semantically coherent per-genre routing specialisation in every variant, with maximum pairwise L1 distance between genre routing distributions in 0.6--1.6 (uniform = 0, disjoint = 2). The variant of highest specialisation strength is the one whose gate the model cannot modulate. Specialisation strength does not predict downstream loss: the four properly initialised variants converge to identical validation perplexity (26.13). To probe whether this loss-invariance is a ceiling or a regime the model actively absorbs, a sixth variant adds a forced non-zero gate initialisation and a non-trainable 3x amplification. The model drifts the gates downward and shrinks the per-hormone magnitudes, but the residual perturbation it cannot suppress 2x the unforced regime, does not damage loss either. We refer to the regime in which the gated injection is consistently absorbed without affecting predictions as a "homeostatic envelope". The envelope is wider than the unforced variants naturally select; the only axis on which the mechanism can break the model is the training schedule (introducing it during base pretraining damages the loss by +2.07 PPL even though the same router still develops specialisation). The loss-relevant variable is therefore "when" the perturbation is introduced, not "what" it contains nor how strongly it is gated. We additionally describe "intracellular attention", a parameter-efficient slot-attention mechanism inside the SSM scan, as an architectural extension; full empirical evaluation requires a fused CUDA kernel and is reported as preliminary.

URL: https://openreview.net/forum?id=uZuaQlxfwG

---

Title: Steered Generation via Gradient-Based Optimization on Sparse Query Features

Abstract: Latent steering exploits internal representations of Large Language Models (LLMs) to guide generation, yet interventions on dense states can entangle distinct semantic features. In this paper, we investigate attention query activations as a high-fidelity site for precise control, hypothesizing that manipulating the attention mechanism itself offers sharper steerability than general state interventions. We introduce Prototype-Based Sparse Steering, a framework that applies Sparse Autoencoders (SAEs) specifically to query activations, to decompose them into interpretable features, then apply gradient-based optimization during inference to align the sparse representation with class prototypes of target behaviors. To validate this architectural insight, we first analyze the mechanism in Textualized Gridworld, a controlled environment for verifiable planning constraints. We demonstrate that optimizing sparse query features enables effective navigation of rigid planning requirements (i.e., safe vs. short paths), confirming the method's ability to satisfy objective rules. We then demonstrate the framework's versatility by training SAEs on a high-dimensional educational domain, where the framework steers the cognitive complexity of feedback (i.e., Bloom's Taxonomy). Our experiments establish that sparse query representations provide the necessary disentanglement for unified, interpretable control over both logical planning and stylistic nuance.

URL: https://openreview.net/forum?id=Dqs2guWeQx

---

Title: Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

Abstract: Modern Lean theorem provers achieve strong performance only with substantial training and inference compute. This cost is driven in part by the scarcity of verified proof data and by the long reasoning traces required for formal proof search, which make both supervised fine-tuning and sampling expensive. We introduce Pythagoras-Prover, a compute-efficient open-source family of Lean theorem provers designed to deliver strong performance under practical compute budgets. The family spans two generation paradigms: two autoregressive models with 4B and 32B parameters, and a first proof-of-concept diffusion-based theorem-proving model (4B), which iteratively refines Lean proofs at inference time. To make training more efficient, we construct a Lean-verified corpus stratified into easy, medium, and hard problems and use it for curriculum supervised fine-tuning, allowing the models to acquire proof skills progressively from shorter and simpler proofs to longer and more difficult ones. During supervised fine-tuning, we further apply a dynamic proof-reasoning filtering scheme that preserves informative proof traces while ensuring each training instance fits within an 8k-token context budget. We further introduce Augmented Lean Formalisation (ALF), which expands scarce verified corpora into variants of formal statements; these variants are then populated through self-distillation, providing additional training signal without requiring every mutated instance to be formally verified. By perturbing known problems while preserving their formal character, ALF exposes the model to structured variants of verified problems, reducing reliance on any single statement's surface form. Empirically, Pythagoras-Prover demonstrates strong performance across model scales. Most notably, Pythagoras-Prover-4B surpasses DeepSeek-Prover-V2-671B at pass@32 on MiniF2F-Test ($82.4% \to 86.1%$), despite using roughly $167\times$ fewer parameters. Scaling to 32B further yields state-of-the-art performance among open-source neural theorem provers, with Pythagoras-Prover-32B attaining $93.0%$ on MiniF2F-Test and solving $93$ of $672$ problems on PutnamBench. We additionally release MiniF2F-ALF, an ALF-mutated contamination-sensitive perturbation benchmark on which every evaluated model loses accuracy; on this split, Pythagoras-Prover-32B remains the strongest evaluated prover, while Pythagoras-Prover-4B reaches parity with Goedel-Prover-V2-32B, the prior state of the art. Together, these results show that strong Lean theorem proving need not rely exclusively on frontier-scale models.

URL: https://openreview.net/forum?id=Fn8AtR6keq

---

Title: Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models

Abstract: We revisit a standard accepted practice in the continuous diffusion language model
literature of fixing conditioning prompt tokens clean during training.
Rather, we make a very simple modification: also noise the conditioning prompt tokens during training.
We demonstrate that under this modified training objective, we achieve better generalization
in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants
($3.73\% \to 24.25\%$ solve rate on Sudoku Hard), and increased diversity of generated solutions ($49.89\% \to 73.93\%$ coverage on
10×10 N-Queens). The method can be seen as conditioning augmentation and a continuous relaxation of classifier-free guidance training.

URL: https://openreview.net/forum?id=carxDjFbtO

---

Title: MTL-MHOL: Predicting Online Conversions under Delayed Feedback and Data Sparsity

Abstract: This paper proposes the model-agnostic Multi-Task Learning Multi-Head Online Learning
(MTL-MHOL) framework for conversion rate (CVR) prediction. Existing approaches
typically address key challenges in CVR prediction, such as delayed feedback and data
sparsity, in isolation or lack flexibility and practical applicability. MTL-MHOL adopts a
time bucketing approach to account for delayed feedback and combines it with multi-task
learning of an auxiliary task to mitigate data sparsity. The model is evaluated on datasets
provided by a private company as well as on a dataset from Criteo, where it outperforms all
benchmark models in terms of Negative Log Loss (NLL) and Relative Cross Entropy (RCE),
correctly captures temporal trends in the data, and demonstrates strong performance when
combined with either MLP or DeepFM. In particular, MTL-MHOL achieves up to $87\%$
lift in RCE compared to the best performing benchmark. The code is freely available at
https://anonymous.4open.science/r/MTL-MHOL.

URL: https://openreview.net/forum?id=QQj5zaYQED

---

Title: Spatio-temporal Multivariate Time Series Forecast with Chosen Variables

Abstract: Spatio-temporal multivariate time series forecasting (STMF) leverages historical observations of spatially distributed variables to predict their future values. In real-world applications, however, budget constraints often limit the number of deployed sensors, resulting in scenarios where only a small subset of variables can be observed at inference time. Existing studies on STMF with missing variables typically assume that the observed variables are pre-determined, leaving the critical problem of how to optimally select input variables unexplored. To fill this gap, we study a new problem, termed STMF with chosen variables (STCV), which aims to jointly select $m$ out of n variables as model input and forecast future values for all variables. We propose a unified framework that tightly couples variable selection and model optimization for both forecasting accuracy and efficiency. Its design novelty centers at its learnable ability of jointly pruning model input (which reduces data collection cost with fewer variables) and pruning model parameters (which improves model efficiency), while optimizing model parameters for forecast accuracy, in contrast to the prior STMF models whose technical designs centered around optimizing model parameters alone. Extensive experiments demonstrate that the proposed method consistently outperforms state-of-the-art baselines in both forecasting accuracy and computational efficiency, highlighting the importance of jointly optimizing variable selection and spatio-temporal forecasting under sparse sensing budgets. Our source code is available at: https://anonymous.4open.science/r/NLT-5E56/

URL: https://openreview.net/forum?id=MZKuJKEgGR

---

Title: DivAS: Interactive 3D Segmentation by Depth-Weighted Voxel Aggregation

Abstract: Interactive 3D segmentation of a reconstructed scene should not require a representation-specific optimization loop. We observe that the recipe for lifting 2D foundation-model masks into 3D, namely prompting a few views, refining the resulting masks with rendered depth, and fusing the multi-view evidence into a voxel grid, is shared across scene representations. What remains representation-specific is only the depth signal returned by the renderer and the occupancy prior that gates fusion. We present **DivAS** (Depth-interactive Voxel Aggregation Segmentation), an optimization-free, training-free framework that realizes this recipe as a single interaction-and-fusion skeleton with lightweight, representation-specific adapters, instantiated on both Gaussian Splatting (GS) and NeRF backbones.
On standard forward-facing and unbounded benchmarks, the GS instantiation attains segmentation quality competitive with state-of-the-art optimization-based methods, and the best on LLFF, while being the only one to reach this quality within the consumer-hardware memory envelope at standard resolution. Both instantiations run end-to-end around $2$x faster than feature-field baselines, with a per-update fusion-kernel cost below $70$ ms. Because segmentation evidence is gathered from a small, bounded set of anchor views, user effort and computation remain independent of the training-set size. The same skeleton applied to a NeRF backbone matches or exceeds the performance of optimization-based NeRF baselines, confirming that the recipe transfers across fundamentally different 3D representations.

URL: https://openreview.net/forum?id=GF9wFsdCqn

---

Title: Is It the Audience or the Question? A Neutral-Paraphrase Control for Adjudicating Claimed Moderators of LLM Persona Effects

Abstract: We set out to confirm an appealing hypothesis: that a large language model’s confidence
shifts more for some audiences than others because the underlying evidence is contested. A
control we built to verify this instead dismantled it. Using a comparable-magnitude, identityfree
neutral-paraphrase baseline, we show that the apparent audience-specific moderation
is largely an artifact: content-neutral rephrasings of the same questions move confidence
as much as personas do, and a biased spread statistic (max−min range) had manufactured
an apparent equivalence that an unbiased one (ω2) overturns. What genuinely survives is
narrower, and we report it as such: audience personas induce more systematic confidence
variation than neutral paraphrases across all seven biomedical models tested, and on a current
frontier model (GPT-5.1), with 90% CIs excluding zero, but this effect is biomedicinespecific
and is not organized by contestedness. Where a contestedness relationship appears
(four of seven models), it is carried by the audience-free condition as often as the persona; in
climate the audience-free condition carries it outright (ρ_neutral = +0.44, p = 0.005; persona
n.s.), and in biomedicine the persona-vs-neutral gap is inconclusive. The contribution
is therefore a control protocol for adjudicating any claimed moderator of LLM behavior,
and a cautionary worked example in which the protocol overturned our own headline result,
the paper’s warning enacted on its own data. We do not claim the trap is pervasive; we
show it is real, easy to fall into, and cheap to test for.

URL: https://openreview.net/forum?id=9EnGvzEchQ

---

Title: A Measurement Framework for Decomposing Factual Recall in Language Models: Storage, Access, and Competition

Abstract: Mechanistic analyses of factual recall often conflate whether a fact is encoded in model
parameters (storage), whether a prompt activates an effective route to that fact (access),
and whether alternative answers suppress the target during decoding (competition). We
present a measurement framework that separates these components with prompt-family
controls, matched competition controls, layerwise invariance measurements, and control-
aware hidden-state interventions. We evaluate 6 locally hosted models from 3 families on
a 5-relation benchmark of 90 facts rendered under 4 prompt families. The best prompt
family for each model also maximizes matched-control selectivity, but the preferred fam-
ily is model dependent rather than universal. In the descriptive full-benchmark analysis,
same-fact prompts exhibit near-immediate subject-position invariance in 524 of 540 fact-
model cases (97.0%, 95% CI 95.2–98.2%) with mean depth 0.19 layers, whereas pre-answer
invariance is weaker and later, appearing in 454 of 540 cases (84.1%, 95% CI 80.7–86.9%)
only after mean depth 18.72 layers. Control analyses show that this late effect survives the
removal of chat-template prompts and weakens but remains deep under aggressive length
matching. A null-controlled subspace test strengthens the causal picture: fact-component
projection produces 0.0382 more alignment loss than relation-matched wrong-fact projection
and 0.0491 more than random projection, while fact-component patching outperforms null
patches but rarely restores strong retrieval on its own. These results support a two-stage
view of factual recall and argue that explainability studies should measure storage-related
support, access, and competition separately.

URL: https://openreview.net/forum?id=YLK7CqSYCH

---

Title: Consecutive Preferential Bayesian Optimization

Abstract: Preferential Bayesian optimization allows optimization of objectives that are either expensive or difficult to measure directly, by relying on a minimal number of comparative evaluations done by a human expert. Generating candidate solutions for evaluation is also often expensive, but this cost is ignored by existing methods. We generalize preference-based optimization to explicitly account for production and evaluation costs with *Consecutive Preferential Bayesian Optimization*, reducing production cost by constraining comparisons to involve previously generated candidates. We also account for the perceptual ambiguity of the oracle providing the feedback by incorporating a *Just-Noticeable Difference* threshold into a probabilistic preference model to capture indifference to small utility differences. We adapt an information-theoretic acquisition strategy to this setting, selecting new configurations that are most informative about the unknown optimum under a preference model accounting for the perceptual ambiguity. We empirically demonstrate a notable increase in accuracy in setups with high production costs or indifference feedback, including a real-world task in food science.

URL: https://openreview.net/forum?id=R1ac8MpBt4

---

Title: Preference-Based Reward Learning under Partial Observability with Inexact Dynamics

Abstract: In this paper, we study how partial observability and inexact latent-state inference affect reward learning from preferences. To that end, we study preference-based reward learning under partial observability, where the learner forms latent-state estimates using an inexact learned POMDP model, so model error can accumulate over time. For finite log-linear POMDPs, we characterize this error term by establishing the stability of the belief filter to parametric model error under certain mixing conditions, yielding bounds on the belief mismatch in expectation and in high probability. We further extend this stability mechanism beyond the log-linear setting to neural-softmax POMDP models with overparameterized neural networks. We then propagate these errors into trajectory-level feature perturbations and derive finite-sample guarantees for constrained Bradley--Terry reward estimation from preferences. Our results decouple statistical error from an irreducible model-mismatch bias, and clarify when preference-based reward learning remains feasible under partial observability with imperfect dynamics.

URL: https://openreview.net/forum?id=TUBr2eh2tf

---

Title: Robust Graph Attention for Graph Adversarial Attacks: An Information Bottleneck Inspired Approach

Abstract: Graph Neural Networks (GNNs) have shown exceptional performance in learning node representations for node-level tasks such as node classification. However, traditional message-passing mechanisms solely based on graph structure in GNNs make them vulnerable to adversarial attacks. Attention-based GNNs have been employed to improve the robustness of GNNs due to their ability to selectively emphasize informative signals over noisy or less relevant ones. In this paper, we propose a novel graph attention method termed Robust Graph Attention inspired by Information Bottleneck, or RGA-IB, which explicitly minimizes the IB loss of a multi-layer GNN through a carefully designed graph attention mechanism. In contrast to existing IB-based graph learning methods that rely on unrealistic Gaussian distributional or local-dependence assumptions due to optimizing variational upper bounds of the IB loss, RGA-IB introduces a novel robust graph attention mechanism that directly reduces the original IB loss, thereby avoiding such restrictive assumptions. Extensive experimental results on semi-supervised node classification under various graph adversarial attacks show that GNNs equipped with RGA-IB exhibit lower IB loss, which indicates better adherence to the IB principle, and show significantly improved node classification accuracy under graph adversarial attacks compared to existing robust GNNs. The code of RGA-IB is available at \url{https://anonymous.4open.science/status/RGA-IB}.

URL: https://openreview.net/forum?id=dBkudf5zK9

---

Title: Constrained Markov Chains, Memory Augmentation, and Mediating Variables

Abstract: Imposing constraints on a Markov chain can create global dependencies that violate the Markov property and render standard algorithms inapplicable. A classical solution is memory augmentation, which augments the chain with an auxiliary state that records the constraint-relevant history. However, existing memory-based methods are scattered across several disciplines and are typically specialized to particular constraints or tasks such as inference or sampling. We present a probabilistic formulation of memory augmentation over discrete and continuous spaces: memory augmentation generates a common hidden Markov structure that supports several constrained tasks including inference, learning, and sampling. This perspective ties together various memory constructions in the literature and reduces several constrained tasks to a suite of standard algorithms that can be applied to any memory-augmented model. It also yields a compositional framework for constructing and verifying which constraints admit tractable representations, covering common examples like hitting, precedence, and counting constraints.

URL: https://openreview.net/forum?id=Q445heZCgh

---

Title: LLM Iterative Fine-Tuning as Filtered Expectation Maximization

Abstract: Large language models (LLMs) can solve complex problems better by thinking about them first and rationalizing answers. The rationalization can be viewed as introducing latent variables, and we propose an EM-like algorithm for learning to rationalize correct answers. The main challenge lies in designing a sampling distribution of rationales that justify correct answers. We instantiate and compare three sampling schemes: rejection sampling with a budget, self-taught reasoner (STaR), and prompt posterior sampling (PPS), which only keeps the rationalization stage of STaR that conditions on the correct answer in the prompt. We experiment with LLM-as-a-judge calibration and summarization from feedback tasks, where conditioning on the correct answer provides a strong guidance for generating rationales. Our experiments show the efficacy of PPS over other sampling schemes, and that the sampling scheme can have a significant impact on performance.

URL: https://openreview.net/forum?id=ud3Bauyrhk

---

Title: What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility

Abstract: Despite the success of deep learning, training deep networks in biologically plausible and hardware-efficient ways remains an open challenge. Feedback alignment (FA) methods address this by replacing backpropagation's symmetric backward weights with fixed random matrices, but their effectiveness depends critically on whether they can be accurately evaluated. The standard evaluation relies on two quantities: task accuracy and cosine similarity between the method's credit signal and the backpropagation gradient. We show that this reporting pair is insufficient by identifying two independent failure modes, both silent under current reporting: (1) measurement degeneracy, where the BP reference gradient collapses to the numerical floor in terminal-LayerNorm residual architectures, rendering cosine uninterpretable; and (2) aggregation collapse, where the aggregate cosine masks layerwise heterogeneity that concentrates credit at one end of the network. To address these limitations, we propose a diagnostic evaluation protocol based on three checks---scale stability, reference validity, and depth utility---together with per-layer rather than aggregate cosine reporting. Across multiple architectures and methods, the standard reporting pair gives no signal of failure in any audited case, while our protocol identifies all failures with wide calibration margins. The two failure modes are causally independent: a per-block scale penalty alleviates Mode 1 (residual scale explosion driving reference collapse) without affecting Mode 2 (cosine ranking that contradicts every functional metric we measured). Identifying these silent failures prevents researchers from building on non-functional credit assignment and provides actionable guidance for developing FA methods that genuinely train deep layers.

URL: https://openreview.net/forum?id=2rnAv6ynUr

---

Title: On the Role of the Projector in Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

Abstract: The dimensional collapse of representations in self-supervised learning is an ever-present issue. One notable technique to prevent such a collapse of representations is using a multi-layered perceptron network called Projector. In several works, the projector has been found to heavily influence the quality of representations learned in a self-supervised pre-training task. However, the question still lingers. What role does the projector play? Assuming the projector mitigates dimensional collapse, what prevents the terminal layer of the base encoder from functioning as the projector in the absence of an explicit multi-layer perceptron (MLP) head? In this work, we intend to study what happens inside the projector by examining the rank dynamics of the same and the encoder through empirical study and analysis. Through mathematical analysis, we observe that the effect of rank reduction predominantly occurs in the last layer. Motivated by this insight, we propose a weight regularization strategy applied specifically to the last layer. We demonstrate that this targeted approach yields better performance than applying orthogonal weight regularization across the entire network (WeRank), both with and without a projector. Our method improves Top-1 accuracy by more than 1% on SimCLR on the ImageNet100 dataset and consistently outperforms baseline SimCLR variants on CIFAR datasets, supporting our interpretation of the projector’s role.

URL: https://openreview.net/forum?id=ACuYOxYy8K

---

Reply all
Reply to author
Forward
0 new messages