Weekly TMLR digest for Jul 26, 2026

6 views
Skip to first unread message

TMLR

unread,
Jul 26, 2026, 12:00:11 AM (10 days ago) Jul 26
to tmlr-annou...@googlegroups.com


New certifications
==================

Survey Certification: Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

Pavel Averin, Theodoros Moysiadis, Ioannis Katakis

https://openreview.net/forum?id=3jzafJK8Tz

---


J2C Certification: FedIndex: Federated Domain Adaptation with Continuous Domain Indices

Qingyang Yu, Hao Wang, Qizhen Zhang, Hao Wang

https://openreview.net/forum?id=fnbGFH0330

---


J2C Certification: DecompDreamer: A Composition-Aware Curriculum for Structured 3D Asset Generation

Utkarsh Nath, Rajeev Goel, Rahul Khurana, Kyle Min, Mark Ollila, Pavan K. Turaga, Varun Jampani, Tejaswi Gowda

https://openreview.net/forum?id=3qy4J6QFbn

---


J2C Certification: REPA-FPO: A Fisher Policy Optimization for Efficient Flow Matching Training

Tianyi Zheng, Fengxiang Yang, Yijie Zhong, Jiayang Gao, Lv Tang, Jinwei Chen, Peng-Tao Jiang, Jia Wang, Bo Li

https://openreview.net/forum?id=mRHipMopOC

---


J2C Certification: Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization

Minttu Alakuijala, Ya Gao, Georgy Ananov, Samuel Kaski, Pekka Marttinen, Alexander Ilin, Harri Valpola

https://openreview.net/forum?id=AtCURrC3XA

---


J2C Certification: ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection

Daisuke Yamada, Harit Vishwakarma, Ramya Korlakai Vinayak

https://openreview.net/forum?id=4Kd0VMsL76

---


Accepted papers
===============


Title: FreeEyeglass: Training-free and Target-mask-free Eyeglass Transfer for Facial Videos

Authors: Weng Ian Chan, Yuantian Huang, Xingchao Yang, Fumio Okura, Takafumi Taketomi

Abstract: The rise of e-commerce and short-video platforms has fueled demand for realistic video-based virtual try-on. Unlike virtual try-on of clothing, which has been actively studied to date, virtual try-on of eyeglasses is uniquely challenging: they align closely with facial structure and strongly affect facial identity, making the faithful preservation of unedited regions especially important. Existing generative editing approaches, such as GAN- and diffusion-based methods, lack reconstruction objectives and often rely on inpainting, which fails to ensure identity consistency. We argue that semantic editing requires not only plausible generation but also faithful reconstruction, making autoencoder-based latent spaces a natural fit. We introduce a training-free, reference-guided framework for video eyeglass transfer built on Diffusion Autoencoders (DiffAE). By blending semantic features in the encoder and incorporating spatial-temporal self-attention, our method achieves realistic, identity-preserving, and temporally consistent results, and points to the potential of autoencoder-based latent spaces for local video editing. The project page is available at https://moegi161.github.io/freeeyeglass-project/.

URL: https://openreview.net/forum?id=6aFRoQcm3H

---

Title: Autoregressive Image Generation with Frequency Progression

Authors: Hu Yu, Hao Luo, Hangjie Yuan, Yu Rong, Jie Huang, Feng Zhao

Abstract: Autoregressive (AR) models for image generation typically adopt a two-stage paradigm of vector quantization and raster-scan ``next-token prediction", inspired by its great success in language modeling. However, due to the huge modality gap, image autoregressive models may require a systematic reevaluation from two perspectives: tokenizer format and regression direction.
In this paper, we introduce the frequency progressive autoregressive (\textbf{FAR}) paradigm and instantiate FAR with the continuous tokenizer.
Specifically, we identify spectral dependency as the desirable regression direction for FAR, wherein higher-frequency components build upon the lower one to progressively construct a complete image. This design seamlessly fits the causality requirement for autoregressive models and preserves the unique spatial locality of image data.
Besides, we delve into the integration of FAR and the continuous tokenizer, introducing a series of techniques to address optimization challenges and improve the efficiency of training and inference processes.
We demonstrate the efficacy of FAR through comprehensive experiments on the ImageNet dataset and verify its potential on text-to-image generation.

URL: https://openreview.net/forum?id=cEfd15ouQ1

---

Title: Optimized Graph Structures for Calibrating Graph Neural Networks with Out-of-Distribution Nodes

Authors: Weili Shi, Xueying Yang, Xujiang Zhao, Haifeng Chen, Zhiqiang Tao, Sheng Li

Abstract: edges between ID and OOD nodes can substantially improve calibration. Identifying these edges and assigning appropriate weights is challenging because the identities of OOD nodes are unknown. To address this challenge, we propose Graph Calibration via Structure Optimization (GCSO), a novel framework for calibrating GNNs in the presence of OOD nodes. GCSO introduces an iterative edge-sampling mechanism to capture graph topological information and formulates adaptive structure optimization as a Markov Decision Process (MDP). An actor-critic policy then dynamically adjusts edge weights based on their effects on target-node predictions. We further design a tailored calibration-aware reward to guide the policy toward an adaptive graph structure that suppresses harmful OOD information propagation. Moreover, the optimized graph structure can be seamlessly integrated with existing temperature-scaling methods for further calibration improvements. Experiments on benchmark datasets demonstrate that GCSO significantly reduces expected calibration error while maintaining competitive classification accuracy.

URL: https://openreview.net/forum?id=Y1W3Z3Z6i8

---

Title: Sharpness-Aware Minimization Driven by Local-Integrability Flatness

Authors: Xuanshuo Fu, Lei Kang

Abstract: Sharpness-Aware Minimization (SAM) improves generalization by optimizing for worst-case loss under parameter perturbations, but its max-based objective can be overly conservative, noise-sensitive, and reliant on smoothness assumptions that often fail in modern nonsmooth networks. We propose Lebesgue Sharpness-Aware Minimization (LSAM), a measure-theoretic alternative grounded in the Lebesgue Differentiation Theorem and local Sobolev regularity. Instead of minimizing the worst-case loss, LSAM minimizes the local average loss in a neighborhood of the parameters. This average-case notion of flatness favors Sobolev-regular Lebesgue points with low local loss oscillation and yields a generalization bound depending only on local integrability, a modulus of continuity, and a Sobolev-induced flatness term—without requiring Hessians or global Lipschitz conditions. To make LSAM practical, we introduce a Monte Carlo estimator of the local average that provides an unbiased gradient with modest overhead. Experiments on CIFAR-10/100 with ResNet, ResNeXt, WideResNet, and PyramidNet show that LSAM consistently finds flatter minima and improves test accuracy over both SGD and SAM.

URL: https://openreview.net/forum?id=29Zg9k5NCo

---

Title: VideoScore2: Think Before You Score In Generated Video Evaluation

Authors: Xuan He, Dongfu Jiang, Ping Nie, Minghao Liu, Zhengxuan Jiang, Mingyi Su, Wentao Ma, Junru Lin, Chun Ye, Yi Lu, Keming Wu, Benjamin Schneider, Quy Duc Do, Zhuofeng Li, Yiming Jia, Yuxuan Zhang, Guo Cheng, Haozhe Wang, Wangchunshu Zhou, Qunshu Lin, Yuanxing Zhang, Ge Zhang, Wenhao Huang, Wenhu Chen

Abstract: Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark \VideoFeedback2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling.

URL: https://openreview.net/forum?id=MpkVh4jH44

---

Title: Probing and Controlling Self-Reflection in Language Models

Authors: Xudong Zhu, Jiachen Jiang, Mohammad Mahdi Khalili, Zhihui Zhu

Abstract: Self-reflection, the ability of a large language model (LLM) to revisit, evaluate, and revise its own reasoning, has recently emerged as a powerful behavior enabled by reinforcement learning with verifiable rewards (RLVR). While self-reflection correlates with improved reasoning accuracy, its origin and underlying mechanisms remain poorly understood. In this work, we first show that self-reflection is not exclusive to RLVR fine-tuned models: it already emerges, albeit rarely, in pretrained models. To probe this latent ability, we introduce Reflection-Inducing Probing, a method that injects reflection-triggering reasoning traces from fine-tuned models into pretrained models. This intervention raises self-reflection frequency of Qwen2.5 from 0.6% to 38.9%, revealing a hidden capacity for reflection. Moreover, our analysis of internal representations shows that both pretrained and fine-tuned models maintain hidden states that distinctly separate self-reflective from non-reflective contexts. Leveraging this observation, we then construct a self-reflection vector, a direction in activation space associated with self-reflective reasoning. By manipulating this vector, we enable bidirectional control over the self-reflective behavior for both pretrained and fine-tuned models. Experiments across multiple reasoning benchmarks show that enhancing these vectors improves reasoning performance by up to 13.7, while suppressing them reduces computational cost, providing a flexible mechanism to navigate the trade-off between reasoning quality and efficiency without requiring additional training. Our findings further our understanding of self-reflection and support a growing body of work showing that understanding model internals can enable precise behavioral control. Our code is publicly available at https://github.com/xzAscC/ProbingReflection.

URL: https://openreview.net/forum?id=AwVIfBZwy0

---

Title: Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey

Authors: Pavel Averin, Theodoros Moysiadis, Ioannis Katakis

Abstract: Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and heterogeneous settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and collider orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.

URL: https://openreview.net/forum?id=3jzafJK8Tz

---

Title: FedIndex: Federated Domain Adaptation with Continuous Domain Indices

Authors: Qingyang Yu, Hao Wang, Qizhen Zhang, Hao Wang

Abstract: Federated domain adaptation incorporates source clients’ knowledge to improve the model performance on the target client under the coordination of the server, mitigating the impact of data insufficiency and domain shift. Existing federated domain adaptation (FDA) methods focus on domain adaptation with categorical domain indices (e.g., “source” and “target”), while many real-world tasks involve domains with continuous domain indices. For instance, hospitals need to adapt disease analysis and prediction across patients via age, a continuous domain index in medical applications capturing the underlying relation between patient information and disease analysis. Prior FDA methods struggle with such tasks due to their ignorance of continuous domain indices. This paper proposes FedIndex to enable FDA with continuous domain indices. FedIndex performs adversarial domain adaptation across clients with the help of a global discriminator, aligning all domains’ distributions. Our theoretical analysis demonstrates the capability of FedIndex to generate domain-invariant features across clients using continuous domain indices without accessing data on clients, simultaneously maintaining privacy preservation. Our empirical results show that FedIndex outperforms the state-of-the-art FDA methods on synthetic and real-world datasets.

URL: https://openreview.net/forum?id=fnbGFH0330

---

Title: You Only Prune Once: A Zero-Shot, Data-Free Pruning at Initialization via Low-Rank Residual Saliency

Authors: Sankar Behera, Anuj Kumar, Mahendra Kumar Gurve, Satyadev Ahlawat, Yamuna Prasad

Abstract: Pruning at initialization (PaI) seeks sparse subnetworks that can be trained from scratch without iterative retraining or post-hoc compression. Most existing PaI methods rely on data, gradients, or iterative structural optimization, and their saliency scores are typically coupled to a specific sparsity budget. This work introduces a zero-shot, data and gradient-free pruning criterion based on nonnegative low-rank residual saliency. At random initialization, a once-only ordering of parameters is obtained by measuring their deviation from a low-rank additive template in the absolute weight space. This fixed ordering can be re-thresholded to realize arbitrary sparsity levels without rescoring, decoupling parameter ranking from sparsity budget and dataset.

Structural and dynamical analyses provide insight into the effectiveness of residual-based pruning. Spectral evaluation shows stronger post-pruning low-rank concentration than competing methods, while neural tangent kernel diagnostics indicate alignment between residual magnitude and functional influence. Empirical results across CIFAR-10/100, Tiny-ImageNet, ImageNet, and modern ConvNeXt architectures demonstrate competitive or superior performance relative to gradient-based and topology-driven PaI baselines, particularly at extreme sparsity ($\geq 99\%$), alongside substantial reductions in pruning time. These findings suggest that a once-only, dataset-agnostic saliency ordering can reliably identify trainable sparse subnetworks from intrinsic structural properties of random initialization.

URL: https://openreview.net/forum?id=J78UavuUUA

---

Title: TensorGRaD: Tensor Gradient Robust Decomposition for Memory-Efficient Neural Operator Training

Authors: Sebastian Loeschcke, David Pitt, Robert Joseph George, Jiawei Zhao, Cheng Luo, Yuandong Tian, Jean Kossaifi, Anima Anandkumar

Abstract: Scientific problems require resolving multi-scale phenomena across different resolutions and learning solution operators in infinite-dimensional function spaces. Neural operators provide a powerful framework for this, using tensor-parameterized layers to capture complex, multi-dimensional relationships. However, scaling neural operators to high-resolution problems leads to significant computational demands, making the training of industrial-scale models prohibitive. In this work, we introduce TensorGRaD, a novel method that directly addresses the memory challenges associated with optimizing large tensor-structured weights. Our approach, based on a robust tensor decomposition, factorizes gradients as the sum of a low-rank tensor and a sparse one to efficiently capture information within optimizer states, including outliers. Additionally, we provide a recipe for mixed precision training of TensorGRaD, achieving further memory savings without sacrificing accuracy. We showcase the effectiveness of TensorGRaD on Fourier Neural Operators, a class of models crucial for solving partial differential equations (PDE). We provide theoretical guarantees for TensorGRaD, demonstrating its fundamental advantage over matrix-based gradient compression methods. We empirically demonstrate large improvements across various PDE tasks, including the challenging turbulent Navier-Stokes case at a Reynolds number of $10^5$. TensorGRaD reduces total memory usage by over 50% while maintaining and sometimes even improving accuracy.

URL: https://openreview.net/forum?id=wd1pTrQFv2

---

Title: Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits

Authors: Adam Bayley, Xiaodan Zhu, Raquel Aoki, Yanshuai Cao, Kevin H. Wilson

Abstract: The recent advancement of Large Language Models (LLMs) offers new opportunities to generate user preference data to warm-start bandits. Recent studies on contextual bandits with LLM initialization (CBLI) have shown that these synthetic priors can significantly lower early regret. However, these findings assume that LLM-generated choices are reasonably aligned with actual user preferences. In this paper, we systematically examine how LLM-generated preferences perform when random and label-flipping noise is injected into the synthetic training data. For aligned domains, we find that warm-starting remains effective up to 30\% corruption, loses its advantage around 40\%, and degrades performance beyond 50\%. When there is systematic misalignment, even without added noise, LLM-generated priors can lead to higher regret than a cold-start bandit. To explain these behaviors, we develop a theoretical analysis that decomposes the effect of random label noise and systematic misalignment on the prior error driving the bandit’s regret, and derive a sufficient condition under which LLM-based warm starts are provably better than a cold-start bandit. We validate these results across multiple conjoint datasets and LLMs, showing that estimated alignment reliably tracks when warm-starting improves or degrades recommendation quality.

URL: https://openreview.net/forum?id=tojKjqIOBd

---

Title: Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes

Authors: Louis Grenioux, Maxence Noble

Abstract: Sampling configurations at thermodynamic equilibrium is a central challenge in statistical physics. Boltzmann Generators (BGs) address this problem by pairing a generative model with a Monte Carlo (MC) correction scheme, yielding asymptotically consistent samples from an unnormalized target density. However, most existing BGs rely on classic MC mechanisms such as importance sampling, which (i) impose strong constraints on the backbone model (typically requiring exact and efficient likelihood evaluation) and (ii) suffer from severe scalability issues in high-dimensional, multi-modal settings. This work investigates BGs built around annealed Monte Carlo (aMC) schemes, which mitigate the limitations of classic MC by bridging a simple reference distribution to the target through a sequence of intermediate densities. In this context, diffusion models (DMs) are particularly appealing backbones: they are powerful generative models and naturally induce density paths that have been leveraged in prior aMC-based methods. We provide an empirical meta-analysis of this DM-based aMC-BG design choice on controlled yet challenging synthetic benchmarks based on multi-modal Gaussian mixtures, varying inter-mode separation, number of modes, and dimensionality. To disentangle learning effects from inference effects, we first study an idealized setting in which the DM is perfectly learned, and then turn to realistic settings where the DM is trained from data. Even in the idealized regime, we find that standard aMC integrations of DMs that rely only on first-order stochastic denoising kernels systematically fail in the proposed scenarios. In contrast, incorporating second-order denoising kernels can substantially improve performance when the required covariance information is available. Motivated by this gap, we propose an alternative aMC integration based on deterministic first-order transport maps derived from DMs; empirically, this approach consistently outperforms its stochastic first-order counterpart, albeit at increased computational cost. Overall, while results in the perfect-learning regime suggest that exploiting DM-induced dynamics within aMC is a promising route to building effective BGs, our experiments with learned DMs show that DM–aMC combinations still struggle to produce accurate BGs in practice. We attribute this limitation primarily to inaccuracies in DM log-density estimation. Code available at https://github.com/h2o64/dabg.

URL: https://openreview.net/forum?id=la4FDaeIbw

---

Title: DecompDreamer: A Composition-Aware Curriculum for Structured 3D Asset Generation

Authors: Utkarsh Nath, Rajeev Goel, Rahul Khurana, Kyle Min, Mark Ollila, Pavan K. Turaga, Varun Jampani, Tejaswi Gowda

Abstract: Current text-to-3D methods excel at generating single objects but falter on compositional prompts. We argue this failure is fundamental to their optimization schedules, as simultaneous or iterative heuristics predictably collapse under a combinatorial explosion of conflicting gradients, leading to entangled geometry or catastrophic divergence. In this paper, we reframe the core challenge of compositional generation as one of optimization scheduling. We introduce DecompDreamer, a framework built on a novel staged optimization strategy that functions as an implicit curriculum. Our method first establishes a coherent structural scaffold by prioritizing inter-object relationships before shifting to the high-fidelity refinement of individual components. This temporal decoupling of competing objectives provides a robust solution to gradient conflict. Qualitative and quantitative evaluations on diverse compositional prompts demonstrate that DecompDreamer outperforms state-of-the-art methods in fidelity, disentanglement, and spatial coherence.

URL: https://openreview.net/forum?id=3qy4J6QFbn

---

Title: REPA-FPO: A Fisher Policy Optimization for Efficient Flow Matching Training

Authors: Tianyi Zheng, Fengxiang Yang, Yijie Zhong, Jiayang Gao, Lv Tang, Jinwei Chen, Peng-Tao Jiang, Jia Wang, Bo Li

Abstract: Flow Matching (FM) models are a leading class of generative models, widely used across diverse domains. However, FM models require large-scale training datasets, which makes training computationally expensive. Existing feature alignment (REPA) improves training efficiency but overlooks the role of the data itself, leaving further room for improvement. In this paper, we observe that different samples carry different amounts of Fisher information and thus contribute unequally to parameter learning in FM. This heterogeneity highlights the importance of accounting for sample-wise contributions during training. However, computing per-sample Fisher information accurately is prohibitively expensive in practice. To overcome this limitation, we provide a mathematical analysis showing that the loss magnitude can serve as an effective proxy for the trace of the Fisher Information Matrix (FIM), enabling efficient estimation. Building on this insight, we propose Fisher Policy Optimization (FPO), a strategy that dynamically reweights samples during training by shifting weight from low-FIM samples to high-FIM samples. Extensive experiments demonstrate that FPO improves both training efficiency and generation quality, while generalizing well across inference samplers, model architectures, and diffusion spaces.

URL: https://openreview.net/forum?id=mRHipMopOC

---

Title: CANDERE-COACH: Reinforcement Learning from Noisy Feedback

Authors: Yuxuan Li, Srijita Das, Matthew E. Taylor

Abstract: Reinforcement learning (RL)
has been widely applied to many challenging tasks.
However, in order to perform well, it requires access to a good reward function, which is often sparse or manually engineered with scope for error.
Introducing human prior knowledge is often seen as a possible solution to the above-mentioned problem, such as imitation learning, learning from preference, and inverse reinforcement learning. Learning from feedback is another framework that enables an RL agent to learn from binary evaluative signals describing the teacher’s (positive or negative) evaluation of the agent's action. However, these methods often make the assumption that evaluative teacher feedback is perfect, which is a restrictive assumption. In practice, such feedback can be noisy due to limited teacher expertise or other exacerbating factors like cognitive load, availability, distraction, etc. In this work, we propose the CANDERE-COACH algorithm, which is capable of learning from noisy feedback by a sub-optimal
teacher. We propose a noise-filtering mechanism to de-noise online feedback data, thereby enabling the RL agent to successfully learn with up to 40\% of the teacher feedback being incorrect.
Experiments on three common domains demonstrate the effectiveness of the proposed approach.

URL: https://openreview.net/forum?id=JgP7Wepetn

---

Title: The Gaussian-Multinoulli Restricted Boltzmann Machine: A Potts Model Extension of the GRBM

Authors: Nikhil Kapasi, Mohamed Elfouly, William Whitehead, Luke Theogarajan

Abstract: Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard continuous latent models can struggle to express. We introduce the Gaussian–Multinoulli Restricted Boltzmann Machine (GM-RBM), a generative energy-based model that extends the Gaussian–Bernoulli RBM (GB-RBM) by replacing binary hidden units with $q$-state categorical (Potts) units, yielding a richer latent state space for multivalued concepts. We provide a self-contained derivation of the energy, conditional distributions, and learning rules, and describe the contrastive-divergence training procedure used in our recall and image-generation experiments. To separate architectural effects from parameter count, we evaluate GM-RBM under fixed visible-to-hidden weight budgets and additional hidden-size sweeps against GB-RBM baselines. On hetero-associative recall benchmarks, GM-RBM achieves competitive recall and, in several regimes, improved recall under fixed visible-to-hidden weight budgets in these experiments. The discrete $q$-ary formulation preserves standard RBM block updates. These results clarify when categorical hidden units provide a simple alternative to binary latents for discrete inference within tractable RBMs.

URL: https://openreview.net/forum?id=3QuXwfMcKo

---

Title: Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization

Authors: Minttu Alakuijala, Ya Gao, Georgy Ananov, Samuel Kaski, Pekka Marttinen, Alexander Ilin, Harri Valpola

Abstract: As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks through experience remains a key challenge. Current LLM agents, particularly those based on proprietary language models, typically rely on prompts to incorporate knowledge about the target tasks. This approach does not allow the agent to *internalize* this information and instead relies on ever-expanding prompts to sustain its functionality in diverse scenarios. This resembles a system of notes used by a person affected by anterograde amnesia, the inability to form new memories. In this paper, we propose a novel method to train AI agents to incorporate knowledge and skills for multiple tasks without the need for either cumbersome note systems or prior high-quality demonstration data. Our approach employs an iterative process where the agent collects new experiences, receives corrective feedback from humans in the form of hints, and integrates this feedback into its weights via a context distillation training procedure. We demonstrate the efficacy of our approach by implementing it in a Llama-3-based agent that, after only a few rounds of feedback, outperforms advanced models GPT-4o and DeepSeek-V3 in tasksets requiring correct sequencing of information retrieval, tool use, and question answering.

URL: https://openreview.net/forum?id=AtCURrC3XA

---

Title: Computationally Sufficient Reductions for Joint Multiple Matrix Estimators with Sparsity and Fusion

Authors: Prateek Sasan, Vincent Q. Vu

Abstract: We study a broad class of methods for the joint estimation of multiple sparse
symmetric matrices that incorporates group and fusion penalties for borrowing
strength across related matrices. This class includes extensions of popular
methods for precision and covariance matrix estimation as well as PCA. We show
that these methods can be unified through the lens of computational
sufficiency, a recently proposed theory that can reveal hidden commonalities
between seemingly disparate methods yielding both theoretical insights into
the underlying optimization problems and practical advantages in terms of
computational efficiency. We derive a universal screening rule that applies
simultaneously to all methods in this class, allowing us to reduce the search
space to block diagonal matrices. This enables streamlined algorithms that
drastically reduce the runtime, making the methods far more scalable and
practical for high-dimensional data analysis.

URL: https://openreview.net/forum?id=KK9RHgSbdp

---

Title: Anderson Accelerated Asynchronous Method for Distributed Optimization

Authors: Cong Li, Xuyang Wu

Abstract: Anderson acceleration (AA) is an effective technique for accelerating fixed-point iterations, but it is rarely applied to distributed optimization. In this paper, we apply AA to accelerate an asynchronous distributed gradient method over the master-worker architecture, resulting in the Asynchronous Distributed Gradient Method with Anderson Acceleration (ADGM-AA). In particular, we first transform the asynchronous gradient method into a fixed-point iteration, and then incorporate it with AA. To ensure the global convergence of ADGM-AA, we equip it with a novel reference-path-based safe-guard scheme. We prove that under mild conditions, ADGM-AA converges with fixed step-sizes that are independent of the delays. Compared with the delay-dependent step-size in most existing works, our delay-free step-size is easier to determine and often leads to faster convergence. Numerical experiments on convex classification tasks show that ADGM-AA improves both iteration-count and wall-clock-time convergence over the baselines in most test examples, while achieving comparable performance in the remaining cases.

URL: https://openreview.net/forum?id=Nm7dTogzfa

---

Title: Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity

Authors: Davide Maran, Marcello Restelli

Abstract: We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression provides a strong theoretical foundation, traditional kernel-based estimators often incur high computational costs and memory requirements that scale with the sample size, limiting their utility in real-time applications such as reinforcement learning. To overcome these challenges, we propose a parametric approach based on a finite-dimensional representation that achieves minimax-optimal uniform convergence rates. Our method enables lightweight inference without storing all samples in memory. We provide sharp finite-sample bounds under sub-Gaussian noise, derive second-order Bernstein-type guarantees, and prove matching lower bounds, thereby confirming the optimality of our approach in both estimation error and memory efficiency.

URL: https://openreview.net/forum?id=7AHO204EaZ

---

Title: MSpecTmol: A Multi-Modal Spectroscopic Learning Framework for Automated Molecular Structure Elucidation

Authors: Wenjie Du, Xiaohan Qin, Ye Wei, Jun Xia, Yang Wang

Abstract: Spectroscopic techniques are indispensable for the elucidation of molecular structures, particularly for novel molecules with unknown configurations. However, a fundamental limitation of any single spectroscopic modality is that it provides an inherently circumscribed and fragmented view, capturing only specific facets of the complete molecular structure, which is often insufficient for unequivocal and robust characterization. Consequently, the integration of data from multiple spectroscopic sources is imperative to overcome these intrinsic limitations and achieve a comprehensive and accurate structural characterization. In this work, we introduce \textbf{MSpecTmol}, a novel \textbf{M}ulti-modal \textbf{Spec}trum information fusion learning framework for automated \textbf{Mol}ecular structure elucidation. By extending information bottleneck theory, our framework provides a principled and adaptive approach to fusing spectra. It designates a primary modality to extract core molecular features while leveraging auxiliary inputs to enrich the representation. To validate the end-to-end effectiveness of our framework, we design a two-fold evaluation: molecular substructure classification to probe its discriminative power in identifying substructures, and extends this knowledge to reconstruct plausible 3D structures. Our results not only demonstrate state-of-the-art performance in molecular substructure classification but also achieve near-experimental accuracy (\textasciitilde 0.68\AA) in molecular conformation reconstruction. These findings underscore the model’s capacity to learn interpretable features aligned with chemical intuition, thereby paving the way for future advances in automated and reliable spectroscopic analysis. Our code can be found at \href{https://anonymous.4open.science/r/MspecTmol-6B4D}{https://anonymous.4open.science.}

URL: https://openreview.net/forum?id=kRhf5Z1Cy1

---

Title: PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Authors: Anmol Kankariya, Sercan O Arik

Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four orthogonal agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives—via final candidate selection, semantic synthesis, or neuro-symbolic verification—to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity’s Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved
reasoning performance using similar or fewer inference tokens compared to heavily scaled
homogeneous baselines.

URL: https://openreview.net/forum?id=wApf83NZmh

---

Title: Synth-FAR: A Synthetic Frequency-Autoregressive Driven Framework for Time Series Forecasting

Authors: Liran Nochumsohn, Michal Moshkovitz, Orly Avner, Dotan Di Castro, Omri Azencot

Abstract: Time series forecasting is essential for predicting future values based on observed patterns. Traditional methods perform well in in-domain scenarios with ample data but struggle with scarce data, leading to the rise of zero-shot and few-shot learning. Recent advancements use large-scale models but require extensive data and resources, often learning ineffectively from the available data. This study explores factors influencing effective learning in time series forecasting using Fourier analysis. Findings show that forecasters struggle with data containing multiple frequencies and generalizing to unseen frequencies. To address this, we introduce Synth-FAR, a synthetic data generation framework that enhances or replaces real data by creating a mixture of autoregressive and frequency information, improving model robustness in limited data scenarios. Our method outperforms other popular synthetic data techniques, such as Kernel-Synth, in both generation time and performance, and demonstrates the potential for integration into foundation model data pipelines, thereby enhancing their effectiveness.

URL: https://openreview.net/forum?id=NuIpv0WySf

---

Title: A Mechanistic Analysis of Low-Precision Instabilities in Microscaling Formats

Authors: Huangyuan Su, Mujin Kwun, Stephanie Gil, Sham M. Kakade, Nikhil Anand

Abstract: Training large language models is an expensive, compute-bound process that must be repeated as models scale, algorithms improve, and new data is collected. To address this, next-generation hardware accelerators increasingly support lower-precision arithmetic formats,
such as the Microscaling (MX) formats introduced in NVIDIA’s Blackwell architecture. These formats use a shared scale within blocks of parameters to extend representable range and perform forward/backward GEMM operations in reduced precision for efficiency gains. In
this work, we investigate the challenges and viability of block-scaled precision formats during model training. Across nearly one thousand language models trained from scratch – spanning compute budgets from 2 × 1017 to 4.8 × 1019 FLOPs and sweeping over a broad range of weight–activation precision combinations – we consistently observe that training in MX formats exhibits sharp, stochastic instabilities in the loss, particularly at larger compute scales. To explain this phenomenon, we conduct controlled experiments and ablations on
a smaller proxy model that exhibits similar behavior as the language model, sweeping across architectural settings, hyperparameters, and precision formats. These experiments motivate a simple model in which multiplicative gradient bias introduced by the quantization
of layer-norm affine parameters and a small fraction of activations can trigger runaway divergence. Through in situ intervention experiments on our proxy model, we demonstrate that instabilities can be averted or delayed by modifying precision schemes mid-training.
Guided by these findings, we evaluate stabilization strategies in the LLM setting and show that certain hybrid configurations recover performance competitive with full-precision training. We

URL: https://openreview.net/forum?id=I5bxWT7Xfw

---

Title: Generalized Dirichlet Energy and Graph Laplacians for Clustering Directed and Undirected Graphs

Authors: Harry Sevi, Gwendal Debaussart-Joniec, Malik Hacini, Matthieu Jonckheere, Argyris Kalogeratos

Abstract: Clustering in directed graphs remains a fundamental challenge due to the asymmetry in edge connectivity, which limits the applicability of classical spectral methods originally designed for undirected graphs. A common workaround is to symmetrize the adjacency matrix, but this often leads to losing critical directional information. In this work, we introduce the generalized Dirichlet energy (GDE), a novel energy functional that extends the classical Dirichlet energy to handle arbitrary positive vertex measures and Markov transition matrices. GDE is a unified framework based on random walk diffusion dynamics, applicable to both directed and undirected graphs, and yields a family of generalized Laplacian matrices usable as drop-in operators in broader graph-learning pipelines. Building on GDE, we propose the generalized spectral clustering (GSC) method for clustering weakly connected digraphs without resorting to a random walk with teleportation. A key component of our approach is the utilization of a parametrized vertex measure encoding graph directionality and density. Experiments on real-world point-cloud and network datasets show that GSC consistently matches or outperforms existing spectral clustering methods in both clustering accuracy and robustness, making it a strong tool for graph-based data analysis.

URL: https://openreview.net/forum?id=AA6D7fJ9PN

---

Title: Savaal: Scalable Concept-Driven Question Generation to Enhance Human Learning

Authors: Kimia Noorbakhsh, Joseph Chandler, Pantea Karimi, Mohammad Alizadeh, Hari Balakrishnan

Abstract: Assessing and enhancing human learning through question-answering is vital, yet automating this process remains challenging. We propose Savaal, a scalable question-generation system using large language models (LLMs) with three objectives: (i) scalability, enabling question-generation from hundreds of pages of text (ii) depth of understanding, producing questions beyond factual recall to test conceptual reasoning, and (iii) domain-independent design, supporting various fields without domain-specific training or prompting. Instead of providing an LLM with large documents as context, Savaal improves results with a three-stage processing pipeline. Our evaluation with 76 human experts on 71 papers and PhD dissertations shows that Savaal generates questions that better test depth of understanding by 6.5$\times$ for dissertations and 1.5$\times$ for papers compared to a direct-prompting LLM baseline. Notably, as document length increases, Savaal's advantages in higher question quality and lower cost become more pronounced.

URL: https://openreview.net/forum?id=2DWDQTsz7K

---

Title: Deep Neural Nets in Low Dimensions with Sign Activations are Convex Lasso Models

Authors: Emi Zeger, Mert Pilanci

Abstract: We consider neural networks with sign activations, depths ranging from 2 to an arbitrary but finite number of layers, and rectangular architectures (parallel structures with constant but arbitrary and finite width). We prove that training such neural networks with weight regularization on 1-D data is equivalent to solving convex Lasso problems with discrete, explicitly defined dictionary matrices. The Lasso dictionaries grow richer for 3-layer networks compared to 2-layers, but saturate thereafter. We show that a tree architecture overcomes this depth limitation, allowing the dictionary to expand with every layer. The Lasso model provides intuition and insight, including closed-form solution paths for 1-D data with binary, periodic labels and extensions to certain 2-D data. Numerical simulations support theory.

URL: https://openreview.net/forum?id=weh3w6KPs6

---

Title: Parameterized Adverse Lens Corruptions to Probe Model Robustness to Optical Tolerances

Authors: Kai Bäuerle, Patrick Müller, Ivo Ihrke, Margret Keuper

Abstract: Deep neural networks excel at image classification on benchmarks like ImageNet, yet they remain vulnerable to adverse conditions, including environmental changes and sensor noise, such as lens blur or camera noise. Consequently, the study of these adverse noise corruptions has been extensive. At the same time, image blur, naturally introduced in optical systems, has been widely ignored as a threat to model robustness. In fact, Gaussian blur has even been considered as viable defense against adversarial attacks. In this work, we challenge the common perception of blur as a rather benign data corruption and study optics-driven, blur- based adversarial attacks. Specifically, we introduce Adverse Lens Corruption (ALC), an optics-driven robustness probe that, through adversarial optimization, identifies worst-case lens blurs by optimizing Zernike polynomial-based aberrations. Unlike traditional noise-based attacks, ALC provides a physically-motivated continuous search space. This enables the analysis of model robustness to optics-driven blur corruptions and complements existing noise and corruption benchmarks.

URL: https://openreview.net/forum?id=a93BmQRNxC

---

Title: Merging Feed-Forward Sublayer for Compressed Transformers

Authors: Neha Verma, Kenton Murray, Kevin Duh

Abstract: Pruning is a prevailing model compression method that identifies and removes unimportant parameters based on various importance metrics. In this work, we instead target redundant parameters via parameter merging, proposing a method that combines Transformer feed-forward sublayers through neuron alignment, merging, and weight tying. We find that this method produces compressed models with performance comparable to their original counterparts while tying more than a third of their feed-forward sublayers, and demonstrates improved performance over a strong, generalized layer pruning baseline. For example, this method enables removing 21% of the total parameters from a vision transformer while maintaining 99% of its original performance on ImageNet. We further show our method composes with QLoRA to further shrink base models before fine-tuning, outperforming an equivalent layer-dropping baseline across downstream tasks. Additionally, we observe high activation similarity between different feed-forward sublayers, offering novel insight into their behavior and contextualizing their surprising mergeability.

URL: https://openreview.net/forum?id=t8iuiH46g0

---

Title: $\texttt{DecompSR}$: A Dataset for Decomposed Analyses of Compositional Multihop Spatial Reasoning

Authors: Lachlan McPheat, Navdeep Kaur, Robert E. Blackwell, Alessandra Russo, Anthony G Cohn, Pranava Madhyastha

Abstract: We introduce $\texttt{DecompSR}$, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning ability. The generation of $\texttt{DecompSR}$ allows users to independently vary several aspects of compositionality, namely: productivity (reasoning depth), substitutivity (entity and linguistic variability), overgeneralisation (input order, distractors) and systematicity (novel linguistic elements). $\texttt{DecompSR}$ has been built procedurally in a manner which makes it is correct by construction, which is independently verified using a symbolic solver to guarantee the correctness of the dataset. $\texttt{DecompSR}$ is comprehensively benchmarked across a host of Large Language Models (LLMs) where we show that LLMs struggle with productive and systematic generalisation in spatial reasoning tasks whereas they are more robust to linguistic variation. $\texttt{DecompSR}$ provides a provably correct and rigorous benchmarking dataset with a novel ability to independently vary the degrees of several key aspects of compositionality, allowing for robust and fine-grained probing of the compositional reasoning abilities of LLMs.

URL: https://openreview.net/forum?id=P81p2nTuvA

---

Title: LLM2Prune: Using LLMs as Domain Experts for Search Space Reduction

Authors: Ankur Nath, Alan Kuhnle

Abstract: Combinatorial optimization problems defined over graphs} involve large discrete search spaces where many candidates contribute little due to redundancy or low value. Pruning the ground set to a smaller pool of promising candidates makes heuristics and exact solvers practical for large real-world instances. Classical submodularity-based pruning algorithms do not scale efficiently, while learning-based approaches depend on handcrafted features that require domain expertise and limit generalization. We propose LLM2Prune, a framework that uses large language models (LLMs) to generate features from a task description, which are then used by a downstream classifier to prune the search space. We guide the feature discovery process with feature-importance scores and performance metrics. Across diverse graph optimization tasks}, LLM2Prune prunes over $90\%$ of the ground set while retaining near-optimal solutions, achieving orders-of-magnitude speedups over existing approaches. Code, data, and pre-trained models are available at: \url{https://github.com/ankurnath/LLM_Pruning.git}.

URL: https://openreview.net/forum?id=i56Pxr3btq

---

Title: MVDGC: Joint 3D and 2D Multi-view Pedestrian Detection via Dual Geometric Constraints

Authors: Thinh Phan, Hao Vo, Khoa Vo, Cuong Pham, Thanh Duc Ngo, Ngan Le

Abstract: The core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlusion reasoning. Recent approaches have addressed this by first projecting image-view features onto a Bird's Eye View (BEV) map, where ground localization is then performed. Despite impressive performance, the perspective transformation induces severe distortion, causing spatial structure break and degrading the quality of object feature extraction. The blurred and ambiguous features hinder accurate BEV point localization, especially in densely populated regions. Moreover, the strong mutual relationship between the BEV ground point and image bounding boxes is not capitalized on. Although multi-view consistency of 2D detections can serve as a powerful constraint in BEV space, these detections are commonly treated as auxiliary signals rather than being jointly optimized with the primary task.

In this work, we propose MVDGC, a unified framework that jointly estimates pedestrian locations on the BEV plane and 2D bounding boxes in image views. MVDGC employs a sparse set of 3D cylindrical queries that embraces geometric context across both BEV and image views, enforcing dual spatial constraints for precise localization. Specifically, the geometric constraints is established by modeling each pedestrian as a vertical cylinder whose center lies on the BEV plane and whose projection casts a rectangular box in the image views. These queries function as shape anchors that directly extract 2D features from the intact image-view features using camera projection, eliminating projection-induced distortions. The 3D cylindrical query enables the unification of BEV and ImV localization into a single task: 3D cylinder position and shape refinement.

Extensive experiments and ablation studies demonstrate that MVDGC achieves state-of-the-art performance across multiple evaluation metrics on MVPD benchmarks, including WildTrack and MultiViewX. On the generalized multi-view detection (GMVD) dataset, MVDGC achieves the highest MODP and precision, while maintaining competitive performance on the remaining metrics, highlighting its robustness and generalization to unseen scene configurations. Code is available at: \url{https://github.com/UARK-AICV/MVDGC}

URL: https://openreview.net/forum?id=40cVQX5Mxc

---

Title: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability

Authors: Jianwen Sun, Yukang Feng, Chuanhao Li, Fanrui Zhang, Zizhen Li, Jiaxin Ai, Sizhuo Zhou, Yu Dai, Shenglin Zhang, Kaipeng Zhang

Abstract: Unified multimodal understanding and generation have attracted much attention in the field of vision and language in recent years. Existing unified models (UniMs) aim to simultaneously learn understanding and generation capabilities, which require a large amount of computational resources and have defects in two aspects: 1) difficulty in generating interleaved text-image content; 2) weaker understanding capabilities than multimodal large language models (MLLMs). To bridge this gap, we propose ARMOR, a resource-efficient framework designed to ``upgrade'' rather than ``retrain from scratch'' expert MLLMs. Our core principle is to endow MLLMs with generation capabilities while preventing catastrophic forgetting of their top-tier understanding capabilities. We achieve this goal through three key innovations: (1) an asymmetric architecture that isolates a lightweight generative decoder from the frozen MLLM core via a forward-switching mechanism to enable seamless interleaved generation; (2) a meticulously curated high-quality interleaved dataset; (3) a progressive ``What or How to Generate'' (WoHG) three-stage training algorithm. Experiments demonstrate that ARMOR successfully upgrades a leading MLLM, retaining over 95\% of its original understanding performance while achieving highly competitive image generation at less than 1/70 the cost of training from scratch. This demonstrates the effectiveness of our core idea: ``the efficient paradigm of upgrading and expanding existing expert MLLMs into UniMs.''

URL: https://openreview.net/forum?id=4TLXaJt8Rq

---

Title: ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection

Authors: Daisuke Yamada, Harit Vishwakarma, Ramya Korlakai Vinayak

Abstract: Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deployment---posing serious risks in safety-critical domains. Recent works have focused on designing scoring functions to quantify OOD uncertainty, with score thresholds typically set based solely on ID data to achieve a target true positive rate (TPR), since OOD data is limited before deployment. However, these TPR-based thresholds leave false positive rates (FPR) uncontrolled, often resulting in high FPRs where OOD points are misclassified as ID. Moreover, fixed scoring functions and thresholds lack the adaptivity needed to handle newly observed, evolving OOD inputs, leading to sub-optimal performance. To address these challenges, we propose *ASAT*, a human-in-the-loop framework that *safely updates both scoring functions and thresholds on the fly* based on real-world OOD inputs. ASAT maximizes TPR while controlling FPR at all times under stationary conditions, even as the system adapts over time. Under nonstationary conditions, the method adapts to distribution shifts with only transient FPR violations during the adaptation period. We provide theoretical guarantees for FPR control under stationary conditions and present extensive empirical evaluations on OpenOOD benchmarks to demonstrate that our approach outperforms existing methods by achieving higher TPRs while maintaining FPR control.

URL: https://openreview.net/forum?id=4Kd0VMsL76

---

Title: Mamba-Enhanced Visual-Linguistic Representation for Multi-Label Image Recognition

Authors: Zichang Tan, Hao Tan, Yang Yang, Prayag Tiwari, Shifeng Chen, Jun Wan, Xu Zhou, Zhen Lei

Abstract: Multi-label image recognition stands as a foundational task in computer vision. Recently, vision-language models have achieved significant progress in this domain. However, previous approaches mostly utilized language models in a simplistic manner, without fully leveraging their potential. To address this, we propose a Mamba-enhanced Visual-Linguistic Representation (MVLR) framework for multi-label image recognition, which aims to better leverage the capabilities of the visual-linguistic representations. In our MVLR, we first propose a Prompt-Driven Label Representation learning (PDLR), which consists of both hard and soft prompts for acquiring comprehensive semantic knowledge for all labels from the large language model. After extracting the label representations, we propose an Interaction and Fusion Model (IFM) to interact with those representations and then fuse them together. To be specific, IFM first employs a label attention to explore the label co-occurrence relations and a context-aware attention to adaptively aggregate context information into label representations. Then, IFM further employs a channel attention to fuse the two features together, forming more reliable and effective label representations. Finally, we propose a Quadruplet Mamba-enhanced Visual-Linguistic block (QMVL) to mutually interact with visual and linguistic features with the strong structure of Mamba. Our QMVL simultaneously emphasizes the features of both visual and linguistic modalities, which is greatly different from previous works of taking linguistic information as a secondary supplementary item. Extensive experiments on several popular datasets, including MS-COCO, Pascal VOC 2007 and NUS-WIDE for general multi-label recognition, demonstrate the superiority of our MVLR.

URL: https://openreview.net/forum?id=KCz9Z9VNwr

---

Title: Causally-Aware Information Bottleneck for Domain Adaptation

Authors: Mohammad Ali Javidian

Abstract: We study a common domain adaptation setting in causal systems with local causal knowledge: the target variable is observed in the source domain but is entirely missing in the target domain, and the conditional mechanism of the target given its Markov blanket is assumed stable across domains. We aim to impute the target variable in the target domain from the remaining observed variables under various shifts. Our central transfer mechanism is structural: restricting the predictor to the Markov blanket of the target screens off shift-prone non-blanket variation and yields zero-shot transfer under blanket invariance. On top of this restriction, we frame estimation as learning a compact, mechanism-stable representation, and we instantiate it with the Information Bottleneck (IB) as a principled compression and regularization mechanism. For linear Gaussian causal models, we derive a closed-form Gaussian Information Bottleneck (GIB) solution that reduces to a canonical correlation analysis (CCA)–style projection and is provably lossless relative to using all non-target variables; in this well-specified regime, ordinary least squares on the blanket is already near-optimal, so the value of IB is regularization rather than accuracy gains. For nonlinear or non-Gaussian data, where no closed-form conditional estimator is available, we introduce a Variational Information Bottleneck (VIB) encoder–predictor that scales to high dimensions and can be trained on source data and deployed zero-shot to the target domain. Across synthetic and real datasets, our approach consistently attains accurate imputations, supporting practical use in high-dimensional causal models and furnishing a unified, lightweight toolkit for causal domain adaptation.

URL: https://openreview.net/forum?id=TbcqPEgJ9z

---


New submissions
===============


Title: SAPIENT: Continual Test-time Adaptation via Lightweight plug-and-play Adapters

Abstract: Continual test-time adaptation (TTA) is the problem of adapting a pre-trained source model at inference-time to handle test samples from a non-stationary distribution, while not forgetting the knowledge acquired from earlier domains. Existing continual TTA methods either make unsupervised test-time updates to the entire model, which can be expensive and prone to forgetting, or do so by keeping the base model frozen and adding a small number of learnable adapter modules for better time/memory efficiency and mitigating forgetting. While such adapter-based methods fall within the broader family of parameter-efficient fine-tuning (PEFT) techniques, which PEFT parameterization is the right one for the unsupervised, source-free, continual TTA setting has remained an open question, which we study systematically in this work. We present SAPIENT (continual teSt-time adaPtation vIa lightwEight plug-aNd-play adapTers), a parameter-efficient adapter based approach which not only offers the usual benefits of the adapter based continual TTA methods, but offers additional key benefits, such as (1) its simple plug-and-play design seamlessly integrates with various continual TTA losses, making our approach complementary to existing continual TTA methods, improving their time/memory efficiency and knowledge retention, (2) it does not require access to the source domain data unlike recent adapter based continual TTA methods, and (3) its parameter-efficiency also makes it computationally feasible to design its Bayesian extensions which can help in estimating the uncertainty in adapter weights, which in turn yields more robust predictions. Through extensive experiments on a segmentation task and four classification tasks for continual TTA, including comparisons with recent continual TTA methods such as SANTA, RMT, ViDA, Continual-MAE, and parameter-selective approaches, we demonstrate that, with substantially (∼90%) fewer trainable parameters, our method achieves competitive or better performance compared to the evaluated continual TTA baselines, resulting in efficient and robust adaptation and inference at test-time.

URL: https://openreview.net/forum?id=OKx9Vw9nZM

---

Title: Density Matrix MDPs: A Corrected Formulation and the Limits of Coherence in Learned State Representations

Abstract: Density matrices have been proposed as reinforcement learning state representations on the
intuition that their off-diagonal entries encode interference between competing hypotheses
that a probability vector discards. We test that intuition and find it mostly does not survive
measurement.
We first correct the framework. The action-only transition T(ρ, a) = P
k Ek(a)ρEk(a)

cannot express a Bayesian belief update, since it takes no observation and so depends only
on the action history. We give the observation-conditioned quantum instrument that can
(Theorem 3), with an explicit witness for the expressivity gap (Theorem 4) replacing a
parameter count.
On the corrected framework, off-diagonal terms contribute essentially nothing on real data:
no significant difference against an architecturally identical diagonal baseline on eight retail
datasets at 21 seeds (three statistical tests), and none on regime-switching environments.
Even on a control built so a coherence-blind policy scores near chance, a properly powered
comparison (20 seeds) shows only a medium, non-significant advantage for the full representation (p = 0.08, d = 0.56). Off-diagonal terms are not inert regardless: they inflate Kraus
gradient variance 26–31×, producing catastrophic tail seeds with no compensating benefit.
The explanation generalises. Coherence of a latent state and off-diagonal entries of a learned
representation are different things: a learned basis is arbitrary, so a diagonal state can
encode latent coherence in its populations—in a capacity-matched test, a diagonal agent
reaches fourteen times a coherence-blind reference score. Expressivity gaps in a map class
do not transfer cleanly to advantages in learned representations, a caution we expect applies
well beyond this framework. A further apparent win for our own method—an instrument
recursion beating a recurrent baseline several-fold—dissolved once that baseline’s learning
rate was tuned.

URL: https://openreview.net/forum?id=ShGyK5P1x3

---

Title: Identifiability Is Not Enough: A Causal Discovery Case Study with $r$-Partite Graphs

Abstract: Observational causal discovery methods have grown rapidly in popularity over recent years, with methodological success evidenced by controlled simulation experiments. However, there have been limited successes in causal discovery under real world data conditions. This suggests a gap between theory and practice. In this work, we investigate this mismatch by studying the behavior of causal discovery methods in cases where causal graphs are provably identifiable and recoverable by all studied algorithms. To contextualize performance, we introduce a simple baseline benchmark that isolates performance gained from respecting the problem's structural topology and basic statistical tests. This baseline establishes an empirical performance standard. Our experiments show that across various $r$-partite settings, existing methods do not consistently surpass this baseline. We conclude that these inconsistent performance gains demonstrate that existing causal theory is insufficient to guarantee strong empirical performance. Consequently, our baseline benchmark provides a necessary reality check for evaluating data-driven causal discovery.

URL: https://openreview.net/forum?id=MrsTmGiBRT

---

Title: Tabular Foundation Models Can Do Survival Analysis

Abstract: While tabular foundation models have achieved remarkable success in classification and regression, adapting them to model time-to-event outcomes for survival analysis is non-trivial due to right-censoring, where data observations may end before the event of interest occurs. We utilize a classification-based framework that reformulates both static and dynamic survival analysis as a series of binary classification problems by discretizing event times. Censored observations are naturally handled as examples with missing labels at certain time points. This classification formulation enables existing tabular foundation models (TFMs) to perform survival analysis through in-context learning without explicit training. In contrast to classical approaches that use binary classifiers to model discrete-time hazards, our approach directly models cumulative failure probabilities, which we find empirically to be more robust to the number of discretization bins by avoiding multiplicative accumulation of per-bin errors. We prove that under standard censoring assumptions, minimizing our binary classification loss recovers the true survival probabilities as the training set size increases. We demonstrate through evaluation across $48$ real-world datasets ($43$ static and $5$ dynamic) that off-the-shelf TFMs with this classification formulation outperform classical and deep learning baselines on average over multiple survival metrics.

URL: https://openreview.net/forum?id=nvolj47f3o

---

Title: To Freeze or Not to Freeze? Memory-Constrained End-to-End Training for Whole-Slide Image Classification in Histopathology

Abstract: Training deep neural networks on gigapixel Whole Slide Images (WSIs) poses significant GPU memory challenges. Multiple Instance Learning (MIL) circumvents this by processing a bag of patches per WSI; however, memory limits typically force the feature encoder—whether a Convolutional Neural Network (CNN) or a foundation model—to remain frozen. In this work, we investigate whether end-to-end (E2E) training improves MIL performance and out-of-domain generalisation. To overcome the memory bottleneck, we introduce a memory-efficient strategy that divides computations into manageable chunks. For CNNs, this involves dynamically offloading intermediate activations to CPU RAM. For ViT-based MIL encoders, we process patches in micro-batches and recompute each micro-batch during backpropagation, recovering the full-bag encoder gradient without retaining all encoder activations simultaneously. Across PANDA/TCGA-PRAD and CAMELYON17/16, we find that E2E training is useful but not uniformly beneficial: its external effect depends on the adaptation recipe and evaluation setting. In particular, combining full E2E ResNet-18 (RN18) training with WSI-level augmentation gives the strongest RN18 transfer to TCGA-PRAD, and E2E + Learning without Forgetting allows H0-mini to match or exceed much larger frozen foundation models there. On CAMELYON16, however, frozen foundation models remain stronger. These results suggest that E2E training is a conditional but important tool for WSI model adaptation rather than a universal replacement for frozen encoders.

URL: https://openreview.net/forum?id=ixDBhyLgWs

---

Title: A Theoretical Analysis of Bayes-Optimal Calibration under the Logit-Noise Model

Abstract: We formulate the calibration of predicted probabilities as a problem in Bayesian decision theory. Under an additive Gaussian noise model in logit space, we show that the Bayes-optimal calibration function for the squared-error loss is uniquely characterized as the conditional posterior mean E[p∗ | p̂], and our main theorem gives it in closed form as a sigmoid integral over a Gaussian posterior under a Gaussian prior. We further show that an approximation to this optimal estimator takes the same functional form as temperature scaling, and we derive the correspondence between the temperature parameter and the noise variance. We also prove that the Bayes-optimal estimator under the Brier score reduces to the same posterior mean, providing a unified theoretical basis for squared-type losses. Finally, numerical experiments under two conditions — synthetic logit data in which the assumption holds by construction, and data mediated by logistic regression in which the assumption holds only empirically — corroborate the predicted theoretical behavior and quantify the range over which temperature-scaling-type correction is effective. This verification presupposes that the true logit is analytically known; extending noise-level estimation to real data, where it is not, remains an open problem.

URL: https://openreview.net/forum?id=r8F3dqoopt

---

Title: Calibration-Aware Relevance Estimation for Structured Filter Pruning

Abstract: Structured pruning is widely used to reduce the computational and memory requirements of deep neural networks (DNNs), enabling efficient deployment in resource-constrained computer vision systems. However, in safety-critical applications, predictive confidence must also be well calibrated, such that confidence accurately reflects the probability of correctness. Existing structured pruning methods primarily optimize predictive performance or feature importance, while largely overlooking uncertainty calibration (UC) as a pruning objective. To address this gap, we propose, to the best of our knowledge, the first calibration-aware sample selection strategy for relevance-based structured pruning. Instead of estimating filter importance over the entire validation set, the proposed framework identifies calibration-critical spillover samples from reliability diagrams and performs Layer-wise Relevance Propagation only on this informative subset, enabling pruning decisions that explicitly account for predictive reliability. We evaluate the proposed approach on convolutional neural networks (VGG-16, ResNet-34/56, and DenseNet-121) and a vision transformer (DeiT-Tiny) for the CIFAR-10, CIFAR-100, and ImageNet datasets. Experimental results demonstrate that the proposed method improves UC while maintaining competitive predictive accuracy under comparable sparsity levels, with reductions in expected calibration error (ECE) of over 50\% relative to conventional structured pruning in the best-performing settings. Furthermore, the resulting pruned models remain amenable to post-hoc recalibration, reducing calibration error below 2\% ECE in several pruning settings while preserving predictive performance. These findings establish UC not only as an evaluation metric but also as an effective optimization objective for structured model compression, improving the reliability of compressed deep neural networks without sacrificing deployment efficiency.

URL: https://openreview.net/forum?id=brVxLE6Kal

---

Title: Neural approximation of Procrustes-Wasserstein transport maps

Abstract: Optimal Transport (OT) provides a rigorous mathematical framework for comparing empirical distributions, such as point clouds. However, standard formulations based on the Wasserstein distances are highly sensitive to global rigid transformations in the feature space.
While invariant OT formulations exist, they often disregard geometric discrepancies that are relevant for shape matching and can be computationally prohibitive (e.g. Gromov-Wasserstein). The Procrustes–Wasserstein (PW) distance addresses this limitation in a more scalable way by directly incorporating orthogonal alignments into the optimization problem. However, current approaches remain confined to discrete settings and lack generalization to unseen samples. In this work, we extend the PW framework by formally deriving its dual and semi-dual formulations. Leveraging results in both convex analysis and, more recently, neural optimal transport, we introduce NeuralPW, a framework that jointly learns rigid alignments and transport maps, being optimal whether the assumptions of the Brenier's theorem hold.
We empirically study NeuralPW in a Gaussian setting, where both the optimal alignment and transport map admit closed-form solutions, and demonstrate that it accurately recovers geometric transformations, explicitly decoupling global pose alignment from local non-rigid deformation. Finally, we validate our approach on a real-world archaeozoological application involving the analysis of morphological deformations between biological specimens. The neural formulation enables high-fidelity matching of dense 3D point clouds at scales where discrete primal solvers become computationally intractable.

URL: https://openreview.net/forum?id=4YFOjibNrt

---

Title: Information-Directed Sampling for Causal Bandits

Abstract: Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even though they influence the reward and provide useful information about the underlying causal system. We study contextual causal bandits with non-manipulable variables, where context variables are observed before action selection and additional variables are observed after each intervention. Assuming a known causal graph without latent confounding, we adopt a Bayesian formulation in which the conditional probability tables of the observational distribution constitute the unknown parameter. This representation allows observations collected under one intervention to update reward estimates for other interventions through their shared causal mechanisms. We develop causal variants of Thompson Sampling and Information-Directed Sampling (IDS) for this setting. For Thompson Sampling, we establish an entropy-dependent sublinear Bayesian regret bound. For IDS, we derive an entropy-dependent regret bound that explicitly quantifies the additional error introduced by Monte Carlo approximation of the expected regret and information gain; when these quantities are available exactly, the bound recovers the standard sublinear IDS rate. We further provide high-probability confidence bounds for the Monte Carlo estimates used by the algorithm. Experiments on several synthetic causal bandit tasks show that the proposed methods outperform causal and non-causal baselines by more effectively exploiting information shared across interventions.

URL: https://openreview.net/forum?id=HCJMKOxGjW

---

Title: ASAMO-Prune: Active Surrogate Learning for Adaptive and Hybrid Multi-Objective Neural Pruning

Abstract: Balancing predictive performance and computational efficiency remains a central challenge in deep neural network pruning, traditionally bottlenecked by the rigid separation of structured and unstructured sparsity and the massive cost of evolutionary search. We present Active Surrogate-Adaptive Multi-Objective Pruning (ASAMO-Prune), a unified framework that jointly optimizes structured and unstructured sparsity for convolutional neural networks (CNNs) and Vision Transformers (ViTs). ASAMO-Prune models each pruning configuration as a hybrid layer-wise chromosome containing structural boundaries, unstructured thresholds, and a dynamically learnable blending factor ($\lambda$) that optimizes the local trade-off between gradient saliency and weight magnitude. To bypass the high costs of evolutionary evaluation, we deploy an active Gradient Boosted Regression Tree (GBRT) surrogate model. By framing surrogate training around rank preservation rather than absolute accuracy regression, our model achieves a near-perfect Spearman rank correlation of $\rho = 0.961$, pre-screening candidate masks in only $\approx 0.32\text{ ms}$ and restricting expensive GPU validation to the true Pareto-optimal frontier. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-1K demonstrate that ASAMO-Prune consistently outperforms state-of-the-art compression methods. On ResNet-18, the framework acts as a highly effective structural regularizer, achieving absolute accuracy gains of $+1.49\%$ and $+1.32\%$ at $50.89\%$ and $72.45\%$ sparsity, respectively, while limiting degradation to just $-0.69\%$ under extreme $80.22\%$ compression. On the structurally sensitive DeiT-Small, ASAMO-Prune achieves a $+0.20\%$ accuracy improvement at $40.00\%$ sparsity, and maintains robust performance at $66.42\%$ sparsity, a threshold where standard heuristics suffer catastrophic accuracy drops.

URL: https://openreview.net/forum?id=5KWWJIIjNT

---

Title: A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

Abstract: Advances in Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs). Among these, Large Audio Language Models (LALMs) are essential for realizing universal auditory intelligence. Despite their remarkable performance, the escalation of LALMs' capabilities has significantly outpaced the development of systemic frameworks to ensure their trustworthiness. This survey provides a comprehensive investigation into the endogenous mechanisms of LALMs, detailing the architectural innovations and alignment algorithms that facilitate emergent reasoning. Specifically, we analyze how the transition to unified end-to-end frameworks and the integration of continuous acoustic signals expand the attack surface. To rigorously evaluate the risks within these paradigms, we establish a comprehensive taxonomy of trustworthiness, categorizing critical vulnerabilities such as cross-modal jailbreaking, latent acoustic backdoors, and biometric privacy leakage. We review the state-of-the-art LALMs through six analytical pillars: hallucination, robustness, safety, privacy, fairness, and authentication. The pronounced imbalance between a mature offensive landscape and underdeveloped defenses highlights persistent trustworthiness gaps and multidimensional risks in audio-centric intelligence. Finally, we propose a roadmap advocating for ``Defense-in-Depth'' architectures, causal auditory world modeling, and intrinsic representation engineering to support the development of more reliable and trustworthy audio intelligence.

URL: https://openreview.net/forum?id=FoLjSB06a3

---

Title: QueryEval: A mechanistic framework to evaluate query- aware generations

Abstract: Query-centric tasks such as query-aware summarization and attribute-specific question answering demand evaluation metrics that assess whether a generated response genuinely
addresses the input query. Existing metrics such as ROUGE and BERTScore measure
lexical and semantic similarity against ground-truth references, but systematically fail to
capture query relevance and require costly human annotation to construct references. We
propose QueryEval, a reference-free evaluation metric that scores query-aware generation
quality by analyzing how attention weights and activations corresponding to query tokens
evolve across transformer layers during generation, without requiring any ground-truth la-
bels. This internal-representation perspective allows QueryEval to directly measure the
degree to which a model attends to and integrates query information throughout the gen-
eration process, making it both interpretable and scalable. We evaluate QueryEval across
three benchmarks (QTSumm, MS MARCO, and SQuAD) and three model architectures-
tures, demonstrating that it correlates strongly with LLM judgments and outperforms both
reference-based and reference-free baselines by a consistent margin. QueryEval requires no
supervised annotation, making it practical for large-scale evaluation pipelines where human
references are unavailable.

URL: https://openreview.net/forum?id=tiYMkT1XHy

---

Title: Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Abstract: Adam is often observed to significantly outperform Stochastic Gradient Descent (SGD) on Transformer-based language models, a phenomenon that has puzzled optimization researchers for years and for which several explanations have been proposed. In this work, we revisit this ``optimizer gap'' by training over 1000 Transformer language models and studying how batch size, training budget, model scale, momentum, and gradient clipping affect rankings. The resulting picture is not a monolithic Adam-SGD separation: at a fixed token budget, SGD is inefficient compared to Adam when common batch size values are used; however, while Adam's performance is only marginally affected when reducing the batch size, well-tuned SGD quickly catches up in this regime, drastically shrinking the gap. Interestingly, SignSGD with momentum behaves similarly to Adam as the batch size varies. We revisit existing explanations for the optimizer gap, including heavy-tailed class imbalance, directional sharpness, and heterogeneity. These lenses describe important aspects of SGD failure, but do not explain why increasing stochasticity can shrink the Adam-SGD gap without performance deterioration. Motivated by this, we study SignSGD not as a complete model of Adam but as the simplest algorithm that exhibits the relevant batch-size dependence. Combining intuition from stochastic differential equations with nonconvex bounds, we isolate a simple mechanism directly affecting convergence rates: signed methods have a batch-dependent mean update direction, whereas SGD's drift is unchanged by batch size.

URL: https://openreview.net/forum?id=1spcnOJogl

---

Title: HOPE: Hessian-Optimized Structured Pruning of LLMs for Efficient Inference

Abstract: Large language models (LLMs) deliver strong performance across diverse tasks, yet their heavy compute and memory demands make deployment on real-time edge devices challenging. Structured pruning has become the standard approach to reduce these costs, yet accurately estimating which blocks can be removed remains challenging at scale. Second-order methods such as Optimal Brain Surgeon (OBS) are computationally intractable at LLM scale. Existing approaches rely on static budgets that ignore cross-layer dependencies, and common proxies like FLOPs misestimate real hardware latency. We introduce HOPE, a scalable, Hessian-aware pruning framework for post-training compression of LLMs. HOPE adaptively reallocates budgets across layers using global screening and selective second-order analysis on a candidate set guided by cross-layer sensitivity estimation. It further performs OBS-equivalent batch pruning that certifies and removes multiple blocks at once while exactly matching the greedy OBS sequence, thereby reducing weight updates and numerical drift. A lightweight latency predictor ensures that the compressed model satisfies inference-time constraints. Experiments on LLaMA and OPT models show that HOPE improves accuracy by up to $3\%$ over state-of-the-art structured pruning methods at comparable pruning ratios.

URL: https://openreview.net/forum?id=YlMvm5k5Vv

---

Title: Pixel Decodability Is Not a Compression Signal: Causally Evaluating Importance Proxies for Visual KV-Cache Eviction

Abstract: Vision-language models (VLMs) retain, in the visual key-value (KV) cache of their language model, a substantial amount of pixel-decodable visual content. Yet we show, in our setting, that this retention is task-inert: across our preregistered tests, how much decodable content a unit retains never positively tracks whether the computation that answers the question causally relies on that unit — where retention relates to task structure at all, the association is weak and points the wrong way. We quantify how much decodable content each visual unit retains with a learned pixel-inversion decoder, and how much the model causally uses it with single-super-patch KV ablation (the teacher-forced drop in gold-answer log-probability), and relate the two within images under a preregistered, sign-calibrated, held-out design. Retention is decoupled from attention ($\rho \approx -0.07$ to $-0.09$) and, in a preregistered, well-powered null (held-out $N=72$, power $\approx 0.9$; upper CIs exclude $\rho \gtrsim 0.06/0.01$), from causal utilization ($\rho \approx -0.01$ to $-0.05$). Yet utilization is not inert to every proxy: attention weakly but significantly tracks it ($\rho \approx +0.11$ to $+0.12$, held-out intervals excluding zero) — the only signal we find that does, serving as the design's positive control. We characterize pixel-decodable retention as an informational axis of the visual KV cache, orthogonal to the functional (attention-utilization) one. How much of this task-inert content a cache holds differs by architecture in our model pair: the encoder-free VLM retains $2.68\times$ more than the encoder-based one. The engineering consequence is a controlled negative result: at super-patch granularity, deconfounded pixel-decodable retention ranks KV eviction no better than random; at token granularity it acquires only a weak inverse-importance signal at the larger budgets — dominated at every budget by attention magnitude, the weak-but-real proxy. In our setting, pixel-decodable reconstructability is not a competitive KV-compression signal at any granularity we test.

URL: https://openreview.net/forum?id=KooT89EuFc

---

Title: Fisher Rank Inflation: A Spectral Signature of Memorization under Label Noise

Abstract: Deep networks trained with label noise often learn clean structure before memorizing corrupted labels. We show that this transition leaves a spectral signature in the centered scatter of per-example last-layer gradients. Its effective rank transiently expands during memorization and contracts after corrupted labels are fit. We call this phenomenon Fisher Rank Inflation.

We show that corrupted labels can increase effective rank by injecting spectral mass into low-energy or previously unused eigendirections, thereby increasing the entropy of the gradient spectrum. We derive a first-order leave-one-out attribution formula and identify conditions under which corrupted examples contribute more strongly to rank inflation than clean examples. We further show that once the normalized Fisher-gradient spectrum stabilizes, individual attribution signals vanish, explaining the post-memorization weakening of leave-one-out rank contributions.

We empirically test these mechanistic predictions on CIFAR-10 and CIFAR-100 using SmallCNN, ResNet18, and Vision Transformers under symmetric label corruption, and additionally evaluate the phenomenon on CIFAR-10N with naturally occurring human annotation errors. Across datasets and architectures, Fisher effective rank exhibits a consistent inflation--collapse trajectory aligned with memorization dynamics. At peak-rank checkpoints, corrupted examples are strongly enriched among the highest rank-contributing samples, with top-100 noisy fractions ranging from $69.2\%$ to $96.2\%$ across five-seed experiments under synthetic corruption and reaching $94.4\%\pm1.9\%$ on CIFAR-10N. The first-order spectral attribution closely matches exact leave-one-out rank contributions in the convolutional models and remains enriched in the Vision Transformer. In addition, a seeded corruption sweep shows that peak Fisher effective rank increases monotonically with corruption severity, rising from $28.88 \pm 1.95$ under clean training to $97.09 \pm 1.78$ at $60\%$ corruption. In several settings, the retrospectively identified onset of rank inflation precedes observable test degradation. The persistence of Fisher Rank Inflation under both synthetic corruption and naturally occurring human annotation errors suggests that the phenomenon captures a broader spectral signature of memorization rather than an artifact of a particular noise-generation process.

These results establish Fisher Rank Inflation as a spectral signature of memorization under label noise and connect the dynamics of the last-layer Fisher-gradient spectrum to corrupted-example enrichment, corruption severity, and the transition from structure learning to memorization.

URL: https://openreview.net/forum?id=jBimgV8OJ3

---

Title: Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

Abstract: We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then performs audio editing to align with the visual changes. To achieve this, we present a new video-to-audio generation model that conditions on the source audio, target video, and a text prompt. We extend the model architecture to incorporate conditional audio input and propose a data augmentation strategy that improves training efficiency. Furthermore, our model dynamically adjusts the influence of the source audio based on the complexity of the edits, preserving the original audio structure where possible. Experimental results demonstrate that our method outperforms existing approaches in maintaining audio-visual alignment and content integrity.

URL: https://openreview.net/forum?id=G8ovcTsiZT

---

Title: Where Language Models Break on Actuarial Calculations: A Step-Level Diagnostic on Life Contingencies

Abstract: Insurers and actuaries have started to use large language models for calculations that carry real financial and regulatory weight. Existing benchmarks report how often a model reaches the right final number, but not where a wrong answer goes wrong, whether the same case reworded gives the same result, or whether a simple prompt change helps. We study these questions on six families of actuarial problems: level annuities paid in arrears and in advance, accumulations, pure endowments, term insurance, and temporary life annuities. Each problem is generated with randomized parameters, and both the final answer and every intermediate quantity are computed in code from a validated mortality and interest engine, so the ground truth is exact and the exact items cannot have appeared in any training set. We evaluate four models reached through a single interface. Final-answer accuracy separates a weak small model (0.21, 95% CI [0.12, 0.34]) from three larger models (0.81 to 0.92), and it varies sharply by family, with term insurance the hardest overall (0.56). When a model is wrong, the error is usually early rather than a slip in the final arithmetic: the first incorrect quantity is the actuarial factor itself in 87% of the weak model's 38 errors, and in 13 of the 18 errors the three stronger models make between them. Answers are not stable under rephrasing: only 57% of cases give the same result across three equivalent wordings (95% CI [0.50, 0.64]), and the weak model never does. Finally, two structured prompt scaffolds that ask the model to lay out the formula and check its work both reduce accuracy rather than raise it (drops of 0.16 and 0.26, both intervals excluding zero). We release the generator, the graded responses, and the figures.

URL: https://openreview.net/forum?id=8i2pl09ZkO

---

Title: Thoughts Without Words: A Controlled Study of Task-Resolving Hidden-State Transfer Between Frozen Language Models

Abstract: Collaboration between language models today mostly requires one model to decode its internal computation into text tokens that another model then re-reads. We ask whether task-resolving information can instead be transferred directly as hidden states between two completely frozen language models. A trained bridge of about 2.4M parameters injects the source's late-layer hidden state into the target's residual stream, and an information-gap design—the target receives only a partial prompt from which the answer is undeterminable—isolates transfer causally. On held-out items, injecting the matched hidden state yields accuracy above 0.90 while no-injection and shuffled-source controls remain near chance (0.26 and 0.35; $p<10^{-4}$), under five control conditions and a pre-registered permutation criterion. Analysis shows that transferable information emerges beyond roughly 65% of source depth and is compressed into a single final-token hidden state; that hidden-state transfer exceeds the argmax ceiling of a single-answer textual handoff; and that in a multi-task setting bridge accuracy tracks the source's per-task competence profile. Replacing the source with a model from a different family changes the target's profile accordingly. Transfer across tokenizer families collapses into memorization under the default recipe but succeeds when enlarged bridge capacity, joint anchor alignment, and soft-label distillation are combined—though this optimization is seed-sensitive on small targets—as demonstrated by transfer from Qwen2.5-7B into the Korean model EXAONE, where all seeds succeed. The headline numbers reproduce on a laptop CPU from released precomputed features and bridge weights, without the source model or a GPU.

URL: https://openreview.net/forum?id=JgxtUfFcOM

---

Title: Do Not Imitate, Reinforce: Iterative Classification via Belief Refinement

Abstract: Standard supervised classification trains models to imitate the labels in their training data. Since this typically happens in a single forward pass, models are locked into a fixed compute budget whether an input is simple or complex. Moreover, the imitation objective forces the model to express absolute certainty on its training data, which often carries over into evaluation, leading to overconfident predictions. Rather than forcing a rigid, single-step guess, early learning systems often framed learning as a trial-and-error process of gradual improvement. Reclaiming this perspective for classification, we propose Reinforced Iterative Classification (RIC), a reinforcement learning framework in which a recurrent actor iteratively refines a predictive distribution and a critic estimates the value of continued refinement. We study the properties of RIC and evaluate RIC-SPO, which trains RIC with Simple Policy Optimization. At inference time, the learned value function provides a principled stopping rule, halting when further refinement is no longer useful. Using image classification as a testbed, RIC-SPO matches or exceeds the accuracy of supervised baselines while reducing calibration error, with the largest gains under label noise and fine-grained classes.

URL: https://openreview.net/forum?id=7xXijNfm8j

---

Title: ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

Abstract: Ethereum decentralized-finance (DeFi) transactions are not ordinary table rows: their economic meaning often depends on nested calls, emitted event topics, token-flow structure, call depth, gas use, and block-local order. Existing blockchain-learning benchmarks can therefore reward memorization of contract addresses, protocol tags, or generic event signatures rather than transferable trace understanding. We introduce \textsc{ETH-TraceBench}, a benchmark for trace-native representation learning on Ethereum DeFi event streams under naturally occurring temporal, protocol, contract, and topic shift. The raw event universe covers Ethereum mainnet from January 2021 through December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. The task-labelled benchmark contains 311.97 million labelled transaction instances and 2.33 billion matched raw logs across a primary DEX swap/trade classification task and a liquidation rare-event stress test. The forward split trains on 2021--2024, validates on 2025H1, and tests on 2025H2. The test set includes meaningful transfer settings: Uniswap v4 and Ekubo v1 are absent from training but appear in validation and test, and 26.2 million DEX test transactions occur on pool or liquidity-infrastructure identifiers absent from training. We report verified shallow trace-count, linear topic/emitter, and lightweight neural topic/emitter baselines regenerated from transaction-level prediction artifacts. On fixed supervised evaluation samples, the strongest trace-count baseline reaches 0.953 macro-F1 for DEX classification and 0.845 macro-F1 for liquidation detection; TopicEmitterTrace-SGD reaches 0.924 macro-F1 for DEX and 0.943 macro-F1 for liquidation; and TopicEmitterHashMLP reaches 0.959 macro-F1 for DEX and 0.938 macro-F1 for liquidation. An OOD slice audit shows that symbolic baselines degrade on transactions containing unseen topics or unseen topic--emitter pairs. These results establish \textsc{ETH-TraceBench} as a shortcut-aware benchmark in which aggregate performance must be interpreted alongside temporal, holdout, masking, and symbolic-novelty evaluations. We specify Masked Event Modeling for Blockchain Logs (MEM-BL) as a reference protocol rather than as a reported model-superiority claim.

URL: https://openreview.net/forum?id=41WQ1aDzQP

---

Title: Computing Monetary Risk Measures in Linear Time

Abstract: Monetary risk measures have gained popularity for expressing decision-makers' risk aversion. Value-at-Risk (VaR) and Conditional-Value-at-Risk (CVaR), in particular, are used commonly for this purpose. This paper proposes new efficient algorithms to compute these risk measures for a discrete random variable in expected linear time with respect to the size of its domain. First, we propose a QuickVaR algorithm that computes the VaR of a discrete random variable. Then, we leverage QuickVaR to propose QuickDivergence, an algorithm for computing a class of $\varphi$-divergence risk measures, including the popular CVaR risk measure. The QuickVaR algorithm adapts the well-known Quickselect algorithm, while QuickDivergence builds on polymatroid optimization algorithms. Numerical results show that our new algorithms offer an order-of-magnitude speedup for large domains, and a library implementation of the algorithms is available at [redacted].

URL: https://openreview.net/forum?id=3DI4fAcZZN

---

Title: Benchmarking Expertise Retrieval for Peer Review: Evaluating Statistical and Neural Representations for the Reviewer Assignment Problem

Abstract: The exponential growth of scientific submissions has strained the peer review system. Despite the rapidly expanding global pool of researchers, this scale has rendered the previous approach of manual expert identification impractical. Therefore, institutions have naturally turned to Large Language Models (LLMs) to automate intricate processes like expertise identification. However, the reliability of these new models in accurately identifying domain experts lacks rigorous evaluation. We conduct an empirical evaluation of statistical and AI-driven expertise representation methods to benchmark their reliability and limitations. Framing expert identification as an information retrieval problem, we utilize the Distributed Peer Review system of the European Southern Observatory (ESO) Period 110 call, where proposal authorship serves as our proxy ground truth for domain expertise. Evaluating six retrieval methodologies utilized across observatories and computer science conferences, we find that traditional statistical representations outperform generative AI. Specifically, Term Frequency-Inverse Document Frequency successfully identified a labeled expert within the top \num{25} recommendations 79.5\% of the time, compared to 51.5\% for \textsc{GPT-4o mini}. Our results highlight that distinguishing subfield expertise benefits from fine-grained vocabulary, which is obscured by the semantic smoothing in generative methods. Using operational data from ESO's Distributed Peer Review system, our findings suggest that transparent and reproducible statistical representations are a strong baseline for expertise identification in peer review, outperforming computationally expensive LLMs in this specialized setting. The code, anonymized aggregate scores, and synthetic benchmark functionality are released to enable full reproduction of our evaluation pipeline and to support future work on expertise retrieval benchmarking at (link provided upon acceptance).

URL: https://openreview.net/forum?id=zOSC5mNEhP

---

Title: Smoothness-Based Derandomization of PAC-Bayes Bounds

Abstract: We study PAC-Bayes derandomization for smooth loss functions. Our goal is to obtain generalization bounds that hold with high probability for deterministic predictors by exploiting smoothness properties of both the loss and the predictor class. We show that passing from the Gibbs predictor to the deterministic predictor at the posterior mean has a precise cost, given by the generalization gap of the Jensen gap class. We control this class through its Rademacher complexity, leading to bounds for deterministic predictors that involve flatness quantities expressed in terms of parameter Jacobians and Hessians of the score map. The framework applies to both bounded and unbounded smooth loss functions, and we specialize the results to linear predictors and smooth neural networks. Finally, the Jacobian and Hessian quantities appearing in the theory motivate a practical regularizer. For BatchNorm networks, we compute this regularizer with respect to effective BatchNorm weights obtained by folding the BatchNorm transformation into the adjacent affine weights. Experiments on CIFAR-10 illustrate the behavior of this regularizer under different batch sizes.

URL: https://openreview.net/forum?id=c0BVeMPz1V

---

Title: Weights describe, representations predict: a cross- architecture topological study of learning and generalisation

Abstract: How a neural network stores what it learns in its weights remains poorly understood. Building on topological studies of weight space, we read each weight tensor at each training check-point as a point cloud and ask two questions: does the shape of the weights reorganise during training in predictable, architecture-specific ways, and does that shape predict how well the network generalises?
We probe both with one instrument, a unified five-category topological screen (Mapper graphs, persistent homology, network-science descriptors, discrete Ricci curvatures, and intrinsic dimension), applied to every layer of five architecture families (MLP, plain CNN, ResNet, LSTM, ViT) over three seeds and tracked across training. Our central finding is a dissociation: weight shape describes how learning organises but does not predict generalisation; the held-out representations do.
The shape does reorganise, with a different organising principle in each family. Convolutional networks develop a per-filter topological sweet spot, a single layer whose dominant H1 cycle sharpens during training and whose position shifts systematically with architecture, from the last convolution in a plain LeNet to the first 3×3 layer after the stem in a ResNet. Other families organise analogously, and beneath all of them the intrinsic dimension of the weight clouds contracts during training.
Weight shape does not, however, predict generalisation. A strong apparent link to the generalisation gap (ρ= 0.92) collapses (ρ= 0.10) once the degree of fitting is equalised: like the loss, the weights record how far the network has fit its training data. The predictive signal appears instead in the representations induced on held-out inputs: the same screen, applied there, does predict the gap, label-free and under extrapolation across datasets and architectures.
All code, experiments and figures are released as an anonymised public repository (released upon acceptance) https://anonymous.4open.science/r/TDA-NN-243F/README.md.

URL: https://openreview.net/forum?id=mtJv2ZKZtg

---

Title: Overlapping Schwarz Attention: Hierarchical Attention via Domain Decomposition

Abstract: We propose a hierarchical attention mechanism based on two-level overlapping Schwarz domain decomposition. The method is motivated by domain decomposition methods in partial differential equations which combine local subdomain corrections with a coarse level that communicates global, long-range information. We test its usefulness in the context of finite-dimensional operator learning using a simple, one-dimensional diffusion problem.
Although elementary, this problem provides a controlled sequence-to-sequence setting in which the exact nonlocal solution operator is known. After discretization, learning the solution operator amounts to approximating the inverse of a symmetric positive definite matrix. As a baseline, we use a global softmax-free low-rank attention operator of the form \(QK^T\). The proposed construction replaces this global factorization by a two-level additive structure: local low-rank attention blocks on overlapping subdomains are combined with a coarse attention block. The resulting operator has the form
$$
M_{\theta}^{-1}
=
\Phi Q_0 K_0^T \Phi^T
+
\sum_{i=1}^{N}
R_i^T D_i^{1/2} Q_i K_i^T D_i^{1/2} R_i .
$$
Here, $R_i$ restricts to an overlapping subdomain, $D_i$ is a partition-of-unity weight, and $\Phi$ is a coarse interpolation %(or prolongation)
matrix.
Numerical experiments for synthetic Fourier right-hand sides indicate that the domain-decomposition attention operator can converge faster and can give more accurate approximations than a global low-rank attention baseline while using significantly fewer parameters.

URL: https://openreview.net/forum?id=EU0qvMb8Q5

---

Title: GreenVoice: Language-Agnostic, Sustainability-Aware, and Scalable Evaluation of Synthetic Speech Generation Models

Abstract: Recent advances in synthetic speech generation models have enabled highly realistic speech that emulates human voices. These models generate voice clones that closely mimic target speakers. Consequently, they pose privacy and security risks through audio deepfakes and audio spoofs. Audio deepfakes are synthetic speech created to deceive humans, whereas audio spoofs are voice clones created to deceive speaker verification systems. Combating these threats requires robust and generalizable audio deepfake detection (ADD) and anti-spoofing models. However, existing models often exhibit domain-specific biases. Developing robust models requires large-scale multilingual synthetic speech datasets. Most existing datasets are in English or Chinese. Creating new datasets poses three challenges. First, it often involves fine-tuning generation models for target languages, which is expensive. Second, speech quality assessments rely on human evaluations, which lack scalability. Third, large-scale speech generation can incur substantial carbon emissions. To address these gaps, we propose GreenVoice, a language-agnostic, sustainability-aware, and scalable framework for evaluating synthetic speech generation models. We empirically validate GreenVoice through a human evaluation with 92 participants. Using GreenVoice, we conduct evaluations across 10 generation models, two English accents, and 12 Indian languages. Leveraging evaluation insights, we introduce Indic-EcoSynth, a 2100-hour multilingual synthetic speech dataset for ADD and anti-spoofing research.

URL: https://openreview.net/forum?id=Icx4HKCS39

---

Title: Kernel Complexity Reduced Graph Contrastive Learning for Noisy Node Classification

Abstract: Graph Neural Networks (GNNs) have achieved remarkable success in learning node representations and have demonstrated strong performance on node classification. However, their effectiveness can be substantially compromised by noise in real-world graph data. To address this challenge, we propose Kernel Complexity Reduced Graph Contrastive Learning (KCR-GCL), a principled framework for noisy node classification with a provable transductive generalization guarantee. KCR-GCL introduces a novel KCR-GCL encoder, which incorporates a new KCR self-attention layer that adaptively balances different frequency components of the graph inspired by generalized graph convolution and reduces the kernel complexity for provably improved generalization for transductive learning. The KCR-GCL encoder is optimized with a low-rank regularization term through the truncated nuclear norm (TNN) on the gram matrix of the learned features. The learned low-rank representations are then used to train a linear classifier for transductive node classification in noisy graph data. The design of KCR-GCL is inspired by the Low Frequency Property (LFP) widely studied in general deep learning and node-level graph learning, and is further supported by a sharp generalization bound for transductive learning. To the best of our knowledge, KCR-GCL is among the first to theoretically reveal the benefits of low-rank regularization in transductive settings for noisy graph data. Experiments on standard benchmarks highlight the effectiveness and robustness of KCR-GCL in learning node representations under noisy conditions. The code of KCR-GCL is available at \url{https://anonymous.4open.science/r/KCR-GCL/}.

URL: https://openreview.net/forum?id=2ALQ54bqNw

---

Title: DP-Splat: Bayesian Nonparametric Complexity Control for Gaussian Splatting

Abstract: 3D Gaussian Splatting represents scenes as finite mixtures of anisotropic Gaussians whose
number of components $K$ is governed by heuristic adaptive density control or user-specified
caps. Variational Bayes Gaussian Splatting (VBGS) recast splat fitting as conjugate
variational inference over a finite mixture, but $K$ remains fixed. We replace the finite
symmetric Dirichlet over mixture weights with a truncated stick-breaking Dirichlet-process
prior---and, as a theory-backed alternative, a sparse overfitted finite Dirichlet---so that
the number of \emph{occupied} components adapts to the data while every update remains a
closed-form coordinate-ascent step (Beta/Categorical/Normal--Inverse--Wishart); a
natural-gradient stochastic variant makes the per-step cost independent of the number of
points. We give an exact monotonicity guarantee, a rigorous truncation-error bound that
corrects an anti-conservative large-$\alpha$ approximation in common use, and an honest
account of what the fitted number of components does and does not estimate. Empirically,
(i) the effective complexity $\Khat$ adapts to scene complexity and recovers the true $K$
(within $\pm1$) on well-separated synthetic data with regime-appropriate concentration;
(ii) a deconfounded three-way image comparison shows the DP prior's contribution is
complexity \emph{selection}, not per-component efficiency: converged DP fits exceed
single-pass fixed-$K$ VBGS at matched budgets by $+2.7$\,dB on average yet tie an equally
converged fixed-$K$ baseline, while on 3D scenes DP-Splat matches or exceeds VBGS's
held-out color prediction with $5.9$--$7.6\times$ fewer components (single-pass fits at compute comparable to VBGS's already match with ${\approx}5\times$ fewer); (iii) the posterior-predictive
color variance is well calibrated on model-matched synthetic data (regression-ECE $\approx 3\times10^{-3}$); and (iv) the ordering suggested by exact-posterior asymptotics reverses under mean-field
coordinate ascent at practical $N$: the Dirichlet-process prior \emph{resists}
over-splitting while the sparse finite mixture saturates its truncation---a gap between
variational practice and posterior asymptotics that we document across three orders of
magnitude in $N$. Code, exact reproduction commands, and all experiment records accompany
the submission.

URL: https://openreview.net/forum?id=75Hx0RDPMr

---

Title: BRAVE: Block-wise Structural Regularization via Controlled Evidence Feedback for Reliable Label Aggregation under Sparse Crowdsourcing

Abstract: Modern label aggregation increasingly feeds soft posteriors, rather than hard labels alone, into downstream learners such as preference and reward models, making posterior calibration as important as top-1 accuracy. In sparse crowdsourcing, however, each item receives only a few annotations from workers who may share systematic biases, so correlated agreement can be mistaken for independent evidence and transformed into overconfident posteriors. We identify this failure mode as illusory evidence accumulation and propose BRAVE, a block-wise structural regularization framework for reliable label aggregation. BRAVE represents each worker as a mixture over shared reliability components and partitions the annotation graph into worker-side blocks. Within each round, block-local posteriors update the shared reliability components, while a synchronized global posterior refines worker profiles. Importantly, BRAVE does not attenuate the predictive posterior under fixed parameters; instead, it controls the feedback route through which sharpened cross-block consensus is written back into shared reliability estimates. Across 14 crowdsourcing benchmarks, BRAVE achieves the lowest NLL among external baselines on 9/14 datasets and the best or tied-best ECE on 9/14, while remaining accuracy-competitive on most benchmarks and exposing explicit accuracy–calibration trade-offs. A controlled counterfactual and reward-model study further support the role of decoupled posterior feedback in improving probabilistic reliability.

URL: https://openreview.net/forum?id=iWFI5hO1dZ

---

Title: Do Domain-Specific Curvature Scales Improve Domain Generalization? A Controlled Re-evaluation of HyperDG

Abstract: Domain-specific curvature is an appealing inductive bias for domain generalization (DG): if source domains induce different representation geometries, learning and coordinating one curvature scale per domain might improve transfer. We conduct a controlled re-evaluation of this hypothesis using HyperDG, a curvature-inspired training objective on an ImageNet pretrained ResNet-50. We first state the implemented objective precisely: it rescales tangent features by learned positive curvature magnitudes and regularizes domain means, curvature dispersion, feature-norm dispersion, curvature sensitivity, and local–global loss disagreement. It does not implement a full Lorentz classifier or a Sharpness-Aware Minimization inner loop. We compare the full objective with a matched zero-penalty control, frozen curvature, and no curvature-variance penalty on PACS, VLCS, OfficeHome, and TerraIncognita, using four seeds and all four leave-one-domain-out splits (16 matched runs per cell). The full objective changes mean accuracy by +0.45, −0.15, −1.02, and −0.73 percentage points, respectively. None of the 12 planned variant-by-dataset comparisons is significant after Holm correction. Moreover, domain curvature scales remain within 0.001–0.011 of one another even without the variance penalty. These results do not establish equivalence and do not refute hyperbolic representation learning; they show that this specific curvature-scaling objective provides no detectable benefit under the tested pretrained-backbone protocol.

URL: https://openreview.net/forum?id=p3Yq1sU7qo

---

Title: The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However, consecutive turns within a conversation are not statistically independent — a fact that virtually all current evaluation pipelines fail to correct for in their statistical inference. We systematically characterize the autocorrelation structure of 66 turn-level metrics across 202 multi-turn conversations (11,639 turn pairs, 5 German-speaking users, 4 LLM platforms) and demonstrate that naive pooled analysis produces severely inflated significance estimates: 42% of associations that appear significant under standard pooled testing fail to survive cluster-robust correction. The inflation varies substantially across categories rather than scaling linearly with autocorrelation: three memoryless families (embedding velocity, directional, differential) aggregate to 14%, while the seven non-memoryless families (thermo-cycle, frame distance, lexical/structural, rolling windows, cumulative, interaction, timestamp) aggregate to 33%, with individual category rates ranging from 0% to 100% depending on per-family effect size. We present a two-stage correction framework combining Chelton (1983) effective degrees of freedom with conversation-level block bootstrap, and validate it on a pre-registered hold-out split where cluster-robust metrics replicate at 57% versus 30% for pooled-only metrics. We provide concrete design principles, a publication checklist, and open-source code for the correction pipeline. A survey of ~30 recent papers at major NLP and AI venues that compute turn-level statistics in LLM evaluations finds that only 4 address temporal dependence at all, and 26 do not correct for it.

URL: https://openreview.net/forum?id=5xm3n1bndg

---

Title: Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization

Abstract: Self-supervised learning matches supervised accuracy from a fraction of the labels, but the
labeled-sample efficiency behind this has lacked a theoretical explanation. We provide one. Data
augmentation induces a similarity graph on the unlabeled data, so downstream learning on that
graph is graph-Laplacian-regularized learning. We prove a fast \emph{transductive} rate,
$O(1/n_L)$ in the number of labels, in place of the supervised $O(1/\sqrt{n_L})$, by carrying the
leave-one-out stability apparatus of Johnson and Zhang (JMLR 2007) over to the augmentation graph,
and without the unrealistic assumptions of limit-based analyses (exact kernel, generalizing
features). The bound makes augmentation quality explicit: the expected error is at most
$C/n_L + R_{\mathrm{DA}}(y)$, where the data-augmentation alignment error $R_{\mathrm{DA}}(y)$ is
proportional to the graph-cut mass of augmentations that cross a label boundary, so good augmentations let few
labels suffice. The analysis uses a streamlined loss that drops the projector, negative-sample,
and orthogonality overhead of standard objectives yet still recovers the top-$K$ ideal features in
the infinite-data limit, the augmentation-kernel eigenspace studied by Zhai et al. Rather than only bounding a generalization
gap, the bound gives a mechanistic account of the accuracy-versus-label-count curve through
augmentation quality, verified in a controlled model where the constants are known.

URL: https://openreview.net/forum?id=Ddnx0lOrWG

---

Title: Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

Abstract: Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.

URL: https://openreview.net/forum?id=jRSI2cHKJe

---

Title: Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning

Abstract: Multi-agent latent reasoning composes the KV-cache contributions of multiple agents into a single context for a final agent. Recent work (Agent Primitives) realises this composition by concatenating per-agent caches along the sequence axis with RoPE re-encoding, a construction we name BagMerge. BagMerge is non-commutative, and which input ordering is best is not predictable a priori: it shifts with the deployment regime, the latent-step budget, and even the model scale. We make this cache exchange a convergent replicated state. First, CanonicalMerge fixes the layout: a content-determined ordering by mean K-norm at a middle transformer layer renders the merged cache byte-identical under any permutation of the inputs, verified algorithmically on synthetic tensors (arbitrary arity $N \le 5$) and bit-for-bit on the real KV state of Qwen3-1.7B (28 layers) and Qwen3-4B (36 layers). Second, we separate the replicated state from decode-time layout: the durable object is a set of content-addressed latent fragments whose merge is set union, a state-based CvRDT (commutative, associative, idempotent, absorbing), and CanonicalMerge is its deterministic render. Because the render is byte-equivalent, every $N = 2$ accuracy number is inherited unchanged and re-delivered duplicate fragments are absorbed rather than reconcatenated. On a partitioned-reasoning benchmark (Qwen3-1.7B), CanonicalMerge lands within 4 percentage points of the best BagMerge ordering in every cell of a $12$-cell regime $\times$ budget $\times$ ordering matrix and matches it without needing to know which ordering is best, trading a small, statistically insignificant accuracy margin for an unconditional structural guarantee; a four-cell Qwen3-4B scale check preserves the same qualitative result. The behaviour transfers to real multi-document QA (HotpotQA bridge-$k = 2$: no detectable degradation against a single-agent full-context baseline), while the closest training-free output-fusion baseline (PackLLM) loses by 45 points at matched generation budget, placing cache-level merging in a regime distinct from output-level fusion. Finally, at $k > 2$ we delimit the approach: cache merge transports and colocates latent traces but does not by itself compose them, motivating subsequent work.

URL: https://openreview.net/forum?id=MOoBC8xoEa

---

Title: Task Conditional Adversarial Domain Adaptation for Regression

Abstract: Accurate training of deep neural networks often requires large labeled datasets, which can be a significant limitation. Unsupervised domain adaptation (UDA) addresses this issue by transferring knowledge from a labeled source domain to a similar but distinct target domain for which only unlabeled samples are available. While UDA has been extensively studied for classification, it remains underexplored for regression tasks.
In this work, we present a conditional adversarial adaptation strategy for regression that leverages a multilinear map between features and position-encoded regression predictions. The proposed position encoding keeps the discriminator input scale independent of the predicted regression value, enabling stable conditional alignment across the continuous output space. We show that conditioning the feature space on regressor predictions allows us to minimize the distance between domains in a way that also reduces the joint error. Furthermore, unlike in classification, we show that it is essential to incorporate the regressor into the minimax optimization of adversarial training to enhance the regressor's expressiveness for the target domain. We further introduce a target-focused batch-normalization layer to prevent a behavior mismatch between training and inference.
We evaluate our approach through extensive experiments and ablation studies on four benchmark datasets: dSprites, MPI3D, Biwi Kinect, and our proposed Syn2Biwi adaptation benchmark. Our method outperforms state-of-the-art techniques overall, achieving an average mean prediction error reduction of up to $25\%$.

URL: https://openreview.net/forum?id=6jZSyO9dXb

---

Title: From Intrinsic Quantization Difficulty to Ranking-Aware Distillation for Low-Bit Multimodal Retrieval

Abstract: Low-bit quantization is standard for large-scale retrieval, where the canonical pointwise distillation baseline, Mean Squared Error (MSE), preserves embedding geometry but not retrieval ranking.

We formalize Intrinsic Quantization Difficulty (IQD): under a fixed bit budget, per-sample ranking degradation varies by up to an order of magnitude, and no static float-space predictor tracks it reliably across regimes; only the online contrastive loss gap, computed for free during training, does.

From this we derive Ranking-Aware Distillation (RAD): replace MSE with a KL divergence on similarity distributions, which provably preserves pairwise margins. RAD is the dominant factor, is statistically significant across benchmarks, wins most settings (+11.7 R@1 on image), and transfers zero-shot across domains. On top of RAD, online difficulty weighting (ODARD) adds a conditional further gain where IQD variance is highest (image retrieval).
At 16-64x compression and up to 83x faster search, the method matches or exceeds the float baseline.

URL: https://openreview.net/forum?id=4h8YJsubM1

---

Title: Beyond Discrimination: Evaluating the Calibration of Computed Tomography Visual Encoders

Abstract: Recent progress in supervised multi-label abnormality classification for three-dimensional (3D) computed tomography (CT) has been driven by increasingly powerful visual encoders.
However, current evaluation remains largely limited to discriminative and ranking-based metrics, overlooking the reliability of predicted probabilities.
Calibration, the alignment between predicted confidence and true outcome likelihood, is critical for clinical deployment, where probabilities inform decision-making and risk assessment.
In this work, we present an evaluation of calibration in multi-label abnormality classification from 3D CT scans. We benchmark representative architectures spanning convolutional, transformer-based, and hybrid CT visual encoders under a unified training protocol across two public datasets, and extend our analysis to CT foundation models spanning diverse supervision paradigms and training data domains.
Our results show that architectures achieving the best performance on standard predictive metrics are not consistently the best calibrated, and that calibration is observed to further degrade under distribution shift. We further identify that miscalibration is structured and data-dependent, increasing with abnormality prevalence and volumetric size, varying across anatomical systems, and exacerbated by abnormality co-occurrence. These findings highlight fundamental limitations of current evaluation practices and emphasize the need for calibration-aware modeling in medical image analysis.

URL: https://openreview.net/forum?id=iIHVJCnvfS

---

Title: A Parametric Linear Conic Projection for Structured Sparse Attention

Abstract: Structured sparsity has become a fundamental ingredient for scaling modern Transformer and Large Language Model (LLM) architectures, where dense attention matrices remain one of the primary bottlenecks in both memory footprint and computational cost. While many existing approaches rely on heuristic pruning or fixed sparse attention patterns, they generally lack a principled geometric formulation and strong optimization guarantees.

In this paper, we introduce a novel convex framework for structured matrix sparsification based on the \emph{Cone Alignment Index} (CAI). The CAI defines a conic constraint whose level sets naturally induce a Lorentz-cone geometry, providing a mathematically grounded characterization of structured sparsity. Building upon this formulation, we derive two complementary projections: the \emph{Parametric Linear Conic Projection} (PLCP), an efficient closed-form approximation, and the \emph{Parametric Euclidean Conic Projection} (PECP), the exact Euclidean projection onto the cone boundary.

Our analysis establishes several fundamental theoretical properties. We prove that PLCP and PECP are collinear, differ only by a positive radial scaling factor, and induce exactly the same sparsity threshold, thereby preserving an identical active support. Furthermore, we derive a support-identification theorem showing that, for any prescribed sparsity level $l<n$, the number of coefficients discarded at each iteration is theoretically controlled. Finally, we show that PECP is the proximal operator associated with the CAI cone, connecting our framework to the well-established theory of convex optimization and proximal algorithms.

These results lead to a simple and scalable two-stage algorithm. First, the active support is identified through a provably correct thresholding rule with exponential convergence. Second, the final projection is obtained by a closed-form support-aware extrapolation coefficient, yielding an efficient bilevel projection algorithm whose computational complexity is essentially linear in the active support size.

Building on this theoretical foundation, we develop a bilevel PLCP framework for structured matrix sparsification. The proposed method naturally generates hardware-friendly structured sparsity patterns, including row-wise, column-wise, and diagonal masks, making it particularly well suited for accelerating Transformer attention mechanisms.

Experiments on large synthetic matrices and Transformer attention layers validate both the theoretical and practical advantages of the proposed approach. On GPUs, PLCP achieves a computational complexity comparable to structured pruning while providing up to a quadratic speed-up over GSP-Hybrid. On GLUE and SuperGLUE benchmarks, PLCP reaches up to $90\%$ attention sparsity with only minor accuracy degradation, consistently outperforming GSP, structured pruning, and fixed sparse-attention architectures such as Big Bird at high sparsity levels.

Beyond sparse attention, the proposed framework provides a general geometric and optimization-based methodology for structured matrix sparsification, opening promising directions toward scalable, hardware-aware, and energy-efficient deep learning.

URL: https://openreview.net/forum?id=GR1gGGa0CC

---

Title: Minimax Theory of Neural Ordinal Regression with Application to Metric Learning Models

Abstract: Ordinal regression is a traditional estimation problem of conditional probabilities in which some ordinal structure is assumed to the pair of explanatory and categorical response variables. The recent advances of deep learning have provided the motivation to study the topic from a learning-theoretic viewpoint under more general statistical models, beyond the classical parametric models. In this work, we introduce an extension of the classical ordinal regression model with unknown vector-valued functions and develop the minimax estimation theory, using deep neural networks. Since we aim to estimate the conditional probabilities, we derive general upper and lower bounds of the minimax risk based on $f$-divergence. We apply the general results to the classical parametric model and its nonparametric variant to verify that the minimax rates are attained via the theoretical framework. We further introduce ordinal regression models based on metric learning with the Euclidean and Riemannian distances, where nonparametric feature maps defined with smooth sets are considered. We prove the minimax rates of the metric learning models, up to logarithmic factors, under the minimax risk of $f$-divergence.

URL: https://openreview.net/forum?id=c89GhdBUnw

---

Title: An Order Sensor: Diagnosing Folds, Kinks, and Multilicity in Differentiable Programs

Abstract: A $\varepsilon$-clamp does not repair a singular gradient; it replaces a power-law divergence by an arbitrary plateau, changing the asymptotic law exactly where it matters. We give a small, dependency-free diagnostic that recovers the order the clamp discards. From a differentiable program's own sampled sensitivity $S(t)$ near a degenerate point, a log-log probe estimates the order of vanishing $\nu$ and classifies the failure as a fold (reparametrize by $s=t^{1/\nu}$), a multiplicity (differentiate the invariant/cluster, not the branch), or a kink (use the one-sided directional derivative) - and refuses when the evidence is insufficient. The estimated pair (sign, order) is the classical signed one-sided valuation of the centered local germ; we introduce no new invariant, only the operational layer: a runnable sensor with confidence and refusal modes, a casebook of differentiable-programming failures with a transfer map, and a proof that the universal $\varepsilon$-clamp mis-scales while the order-aware repair does not. The sensor returns the expected classification on six constructed cases (stable under +/-20% noise); on genuinely computed gradients (pure-Python forward-mode AD and jax.grad, in float64/float32) it recovers orders, declines a fold repair on nondiverging objects (wrong-object prevention), and reads $\nu$ off a real $n$-dimensional implicit-layer backward solve, refusing when the singular direction is unexcited. On a deep-equilibrium digit classifier trained on real data, it reads the trained network's own backward solve as regular (every test equilibrium contractive) and a clean fold ($\nu=2.01$) once the weights are pushed onto an actual branch fold. An optimizer-universality result separates magnitude errors, which adaptive normalization tames, from direction errors, which no preconditioned optimizer fixes, yielding a regime law for when fold-aware backward passes matter. Finally, in an experiment registered in advance on a differentiable contact simulator, the $\nu=2$ reading selected a square-root chart from nonsmooth dynamics, which then outperformed the field's tuned soft-contact patch near grazing (the advantage vanishing for interior optima); a second advance-registered transfer returned an honest negative that brackets when such imports win. An at-scale training-loop benchmark remains future work.

URL: https://openreview.net/forum?id=pL9pE5R1AD

---

Title: Breaking the Ouroboros: A Closed-Loop System for Suppressing Model Collapse in Recursively Trained Language Models

Abstract: Large language models are increasingly trained on web-scale corpora that are saturated with their own synthetic outputs. This recursive self-consumption drives model collapse (also called Model Autophagy Disorder): tail knowledge is discarded, outputs regress to a bland mean ("AI slop"), and the model eventually enters a fluent-but-wrong regime before semantic integrity disintegrates entirely. Two lines of defense have been proposed—(i) detecting and filtering synthetic data, and (ii) accumulating real anchor data across generations—but each, used in isolation, leaves critical failure pathways open. We identify four such pathways and show that closing them requires their joint operation as a feedback-controlled loop rather than a pipeline of independent stages. We present a closed-loop framework with four coupled modules: (1) provenance- and watermark-based information purification; (2) diversity maintenance that, crucially, measures entropy only after purification to avoid mis-protecting the "false diversity" of synthetic data; (3) assimilation-aware anchor construction, which introduces the AI Assimilation Degree—a quantitative gate that rejects nominally-human data covertly contaminated by AI assistance—together with an accumulate-never-replace store, drift-triggered dynamic anchor mixing, and a PCP-style adversarial-debate verifier for scalable oversight; and (4) a self-reflection module that endows the model with logical-type discrimination, an internal "logical immune system" against harmful data that evades filtering. Across six controlled recursive-training studies on a 130M-parameter decoder, each module addresses an otherwise-open collapse pathway, and their integration yields a super-linear synergy: test loss remains within $1.05\times$ of its initial value after ten generations, whereas replacement collapses to $4.5\times$ and accumulation-only reaches $1.5\times$. We release our full protocol to support reproduction.

URL: https://openreview.net/forum?id=HLpPFFvHD4

---

Title: Learning From Past Mistakes: Contrastive Reflection Memory for Training-Free Regeneration

Abstract: Verification-guided self-improvement has recently emerged as a promising approach to improving the accuracy of large language model (LLM) outputs.
However, existing approaches face a trade-off between inference efficiency and accuracy: iterative verification–rectification is computationally expensive and prone to being trapped in faulty reasoning, while best-of-N selection requires extensive sampling without addressing internal model flaws.
We propose a training-free regeneration paradigm that leverages an offline-curated contrastive Reflection Memory (RM) to provide corrective guidance, while regenerating from scratch helps break out of faulty reasoning. At inference time, the method performs RM-guided self-verification followed by a single RM-guided regeneration, avoiding both iterative correction and multi-sample selection.
We evaluated our method on nine benchmarks that span algorithmic, reasoning, symbolic, and domain-specific tasks in both small- and large-scale LLMs. Experiment results show that our method outperforms prior methods while maintaining low computational cost.
Code is available at \url{https://anonymous.4open.science/r/Supplementary-code-5B17}.

URL: https://openreview.net/forum?id=Ilif7c3hvj

---

Title: The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

Abstract: Existing AI governance frameworks address safety, transparency, and accountability but, to our knowledge,
none operationalises quantitative constraints on macro-socioeconomic stability. We introduce the
Human Utility Factor (HUF), a di erentiable welfare metric that models the interaction between Agency,
Wellbeing, and Economic Stability through three policy levers: automation depth ha, redistribution intensity
α+β, and employment coverage ρ. HUF yields a closed-form optimal automation level ha and a
minimum redistribution threshold (α+β)∗ below which automation is not welfare-improving, transforming
high-level governance objectives into computable constraints. We evaluate HUF using a three-agent
multi-agent reinforcement learning framework across U.S., Canadian, and Nordic policy regimes. Analytical
agents converge to the closed-form welfare optimum, while PPO agents discover an alternative local
optimum that satisfies the aggregate welfare objective while violating redistribution constraints. These
results suggest that AI governance is fundamentally a constrained optimization problem rather than a
compliance exercise, and that computable welfare constraints are necessary for evaluating automation
policies at scale.

URL: https://openreview.net/forum?id=yVaFqDLxkZ

---

Title: Latent Point Collapse on a Low Dimensional Embedding in Deep Neural Network Classifiers

Abstract: The topological properties of latent representations play a critical role in determining the performance of deep neural network classifiers. In particular, the emergence of well-separated class embeddings in the latent space has been shown to improve both generalization and robustness. In this paper, we propose a method to force the collapse of latent representations belonging to the same class into a single point, which enhances class separability in the latent space while confining
network outputs to a bounded range.
We demonstrate that this phenomenon, which we call \textit{latent point collapse} (LPC), is induced by adding a strong $L_2$ penalty on the penultimate-layer representations and arises from the interplay between the $L_2$ penalty and the
cross-entropy loss.
We note that the added penalty makes the loss strongly convex with respect to the penultimate representations, and that increasing its strength tightens the confinement and reduces the within-class spread of the representations.
In addition, we show the practical utility of applying this compressing loss term to the latent representations of a low-dimensional linear penultimate layer.
LPC can be viewed as a stronger manifestation of \textit{neural
collapse} (NC): while NC only requires within-class representations to
converge around their class means relative to between-class separation,
LPC enforces absolute convergence to single points near the origin,
yielding more pronounced improvements in robustness and generalization.

URL: https://openreview.net/forum?id=h0IFTvq5Uj

---

Title: Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time weighting of per-depth losses. This entangles exit selection with trajectory formation: the gate not only chooses which recurrent state to use, but also determines how strongly each intermediate state is supervised. Consequently, poor adaptive-compute performance can arise from the readout, the induced trajectory, or their interaction. We study adaptive depth in looped Transformers through this trajectory--readout lens, across controlled synthetic tasks (modular arithmetic and binary parity) and large-scale Ouro-1.4B and 2.6B checkpoints. We find that fixed-prior depth supervision, which shapes the trajectory without an input-dependent halting policy, produces difficulty-aware trajectories whose intermediate states expose useful stopping signals, and that simple post-hoc confidence readouts often match or outperform learned linear and MLP gates. Fitting gates on frozen trajectories localizes the failure: it appears to stem mainly from the trajectory induced by joint gate training rather than from limited gate expressivity. The same pattern is present in Ouro evaluations, where pretrained ponder gates are competitive but not uniformly Pareto-optimal, and measured latency confirms that the resulting reductions in average exit depth translate into practical inference-time savings. Our systematic diagnostic evaluation reframes adaptive depth in looped Transformers as a joint problem of trajectory formation and exit readout, rather than gate learning alone, highlighting a distinction that prior learned-halting work has often left implicit.

URL: https://openreview.net/forum?id=FOz5OWi6bb

---

Title: On $\ell_0$-Regularized Sparse Multiple Kernel Learning for Feature Selection

Abstract: This paper studies the direct use of $\ell_0$ regularization in sparse multiple kernel learning (SMKL) problems.
We introduce an explicit $\ell_0$ norm on the kernel combination coefficients, leading to min-max mixed-integer optimization formulations, in which kernel selection is handled in the outer problem and a standard SVM is solved in the inner problem.
We compare $\ell_0$ and the conventional $\ell_1$-based SMKL formulations.
In the case of rank-one linear base kernels, we show that imposing an $\ell_0$ constraint on the kernel coefficients is equivalent to controlling the rank of the combined kernel.
Compared to $\ell_1$ regularization, this provides a more direct control of the learned kernel with an interpretation as features selection.
Furthermore, we investigate the impact of kernel scaling in $\ell_0$-regularized MKL and show that sparsity control and kernel scaling represent distinct modeling objectives, leading to an $\ell_0$--$\ell_p$ MKL framework that controls sparsity through $\ell_0$ regularization and kernel scaling through $\ell_p$ regularization.
Extensive experiments on synthetic and UCI benchmark datasets demonstrate that incorporating $\ell_0$ regularization into SMKL leads to sparser kernels while achieving improved or comparable test accuracy.
This contrasts with common knowledge in MKL where less sparse solutions are often observed to perform better.

URL: https://openreview.net/forum?id=bDgiWKb1Cn

---

Title: Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information

Abstract: Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to deception. Under asymmetric information, deception is seen as strategic shifts in the type distribution such that the omniscient opponent can collude with Nature and condition its play on the observed type. Existing risk-neutral solution concepts, such as Bayesian Nash equilibrium (BNE), are sensitive to distribution shifts, while distributionally robust approaches provide guarantees only within a prescribed ambiguity set. To address these limitations, we introduce Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE), a novel equilibrium concept that combines the distribution-free robustness of minimax-regret reasoning with probabilistic information from a nominal type distribution. PR-MRE minimizes worst-case regret over a high-confidence subset of the type space, providing protection against strategic redistribution of probability mass while avoiding the conservatism of fully distribution-free approaches. We show that, for normal-form Bayesian games, PR-MRE can be formulated as a robust bilinear program and derive a tractable semidefinite relaxation. We then adapt this relaxation into a novel meta-solver within a robust double-oracle framework, PRMRE-PSRO, enabling population-based learning of approximate PR-MRE strategies via deep reinforcement learning best responses. Experiments on graph-structured adversarial team games demonstrate that PR-MRE discovers strategies with substantially improved worst-case performance across hidden types compared to risk-neutral equilibrium solutions, resulting in more robust behavior under strategic distribution shifts.

URL: https://openreview.net/forum?id=Nx0gcKBEFY

---

Title: Frozen Vision Foundation Models Narrow the Sensor-Shift Gap in Non-Visual Spectrogram Recognition

Abstract: WiFi-based micro-Doppler activity recognition offers a privacy-preserving approach to human
sensing, but models trained under one RF acquisition setting often perform poorly when
tested with different sensors, environments, or datasets. A key challenge is that micro-Doppler
spectrograms contain both activity-related motion patterns and unwanted variation from the
sensing and processing pipeline. Activity information includes temporal rhythm, Doppler
spread, and time-frequency motion structure, while unwanted variation may arise from
room clutter, multipath propagation, sensor response, preprocessing, rendering choices, and
dataset-specific artefacts. Existing approaches often address this problem by reducing the
mismatch between source, synthetic, and target recordings. This paper investigates whether
external WiFi micro-Doppler recognition is better addressed by learning stable features than
by adapting input spectrograms toward the target appearance.
We investigate this idea using self-supervised vision encoders pretrained on natural images
and kept frozen during RF training. All models are trained on WiFi micro-Doppler data
collected in our lab and evaluated on two external settings: OPERAnet, a public multimodal
RF benchmark collected outside our lab, and a through-wall test set collected by adding a
wall condition to the same lab setup. Despite using an encoder that is not fine-tuned on RF
data, DINOv2-S/14 with a lightweight MLP classifier substantially improves recognition on
OPERAnet, increasing class-restricted macro-F1 from 0.20±0.01 to 0.46±0.01 compared
with the strongest convolutional baseline trained on the same lab-collected data. On the
through-wall test set, DINOv2 remains competitive but is not uniformly superior, suggesting
that pretrained representations are most beneficial when the test data differ more strongly
from the training data. These results suggest that external RF activity recognition should
be treated not only as a problem of adapting input appearance, but also as a problem of
learning representations that preserve activity-relevant motion structure while being less
sensitive to sensor, environment, and dataset-specific variation.

URL: https://openreview.net/forum?id=akIq8wHuUD

---

Title: Unified Framework for Causal Inference: Causal Calculus, Transfer Entropy, and Convergent Cross-Mapping as Conditional Mutual Information

Abstract: Three dominant methodologies for causal inference---Causal Calculus (CC), Transfer Entropy (TE), and Convergent Cross-Mapping (CCM)---have developed independently, each with its own mathematical foundations. We prove that all three are special cases of a single causal strength measure based on conditional mutual information, with Granger causality recovered as a corollary of the Transfer Entropy instantiation. Building on this unification, we derive precise sample complexity bounds that expose the computational cost of each method
and construct an automated Bayesian model-averaging pipeline that selects, executes, and aggregates the three methods without expert intervention. The approach is validated on synthetic linear systems, coupled chaotic maps, and a simulated climate--ecology benchmark,
where the ensemble asymptotically matches or surpasses the best individual method. The complete framework is available as the
open-source unified_causality Python toolkit with a scikit-learn-compatible API that supports reproducible causal analysis
and scales to large datasets.

URL: https://openreview.net/forum?id=1UU077WP9X

---

Title: The Coverage Regime for Best-of-K Candidate Selection under Distribution Shift

Abstract: Many evaluation and deployment settings allow retaining multiple complete candidate systems while the final operating distribution is unknown. In such settings, selecting the K candidates with highest average validation accuracy can be suboptimal because it may duplicate the same hidden failure modes. We study fixed-library best-of-K candidate selection under distribution shift: given a library of complete candidate predictors and labeled selection scenarios, choose a K-subset whose per-scenario best member performs well on structurally related but held-out shifts. We formalize the pure coverage objective $F(S) = E_t[\max_{i \in S} a_i(t)]$ and show that it is monotone submodular, so greedy selection has a classical $(1 - 1/e)$ approximation guarantee. We introduce COVER, a greedy or exact selector for this objective, and characterize when it should be used. The key diagnostic is validation headroom: the gap between the best achievable K-set under max-over-K scoring and the best single candidate. Across real-support shifts, real-domain shifts, controlled corruptions, and correlation-collapse controls, validation headroom predicts when COVER improves hidden-shift tail performance and when it collapses to no gain. COVER is intentionally narrow: it is fixed-library, label-observed, and relevant to max-over-K or tail-scored regimes. It is not a universal OOD method and does not generally improve mean accuracy; Caruana-style selection remains the strongest mean-oriented baseline.

URL: https://openreview.net/forum?id=PjeMofFpq2

---

Title: The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models

Abstract: Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted. Although fine-tuning methods have been proposed to unlearn a target concept, they struggle to fully erase it while maintaining generation quality on other concepts, leaving models vulnerable to jailbreak attacks. Existing jailbreak methods demonstrate this vulnerability but offer limited insight into how unlearned models retain harmful concepts, limiting progress on effective defenses. In this work, we show that the erased concept persists as a coherent, interpretable linear subspace of the token embedding space, and that both an attack and a defense follow directly from this structure. We introduce SubAttack, a novel jailbreaking attack that reads out this subspace by learning an orthogonal set of attack token embeddings, each being a linear combination of human-interpretable textual elements, revealing that unlearned models still retain the target concept through related textual components. Furthermore, our attack is also more powerful and transferable across text prompts, initial noises, and unlearned models than prior attacks. Conversely, projecting out the same subspace yields SubDefense, a lightweight plug-and-play defense mechanism that suppresses the residual concept in unlearned models. SubDefense provides stronger robustness than existing defenses while better preserving safe generation quality. Extensive experiments across multiple unlearning methods, concepts, and attack types demonstrate that our approach advances both understanding and mitigation of vulnerabilities in diffusion unlearning.

URL: https://openreview.net/forum?id=CwGv55bbj4

---

Title: Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Abstract: Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existing methods regenerate every pixel and often alter elements that should remain unchanged, such as characters or background scenes. We introduce Vera, a layered diffusion framework for content-preserving video editing. Instead of regenerating the entire video, Vera generates an edit layer along with an alpha matte for compositing with the source video, separating creative editing from content preservation by design. To encourage coherent composition with the source video, we extend the text-to-video DiT into a Mixture-of-Transformers (MoT) architecture, with separate DiTs for each layer that interact through joint self-attention. To support the training of Vera, we further construct a high-quality layered dataset with accurate alpha mattes, diverse scenes and dynamics, and visual effects.
Across our quantitative benchmark and human preference study, Vera outperforms leading open-source video editing models in content preservation while remaining competitive in edit quality, using 486K frames of layered training data.

URL: https://openreview.net/forum?id=aDJRKLtUeQ

---

Title: Diversity Is Not Enough: Why Representational Divergence Fails to Predict Fusion Utility in Pathology Foundation Models

Abstract: A common heuristic for choosing which pretrained models to combine is representational diversity: prefer models whose internal representations differ most, on the assumption that divergent representations carry complementary signal. Using Centered Kernel Alignment (CKA) across four foundation-model encoders spanning three architecture families, evaluated on three histopathology classification benchmarks, we show that this heuristic can fail, and characterise when. We first establish, under controlled conditions, that architectural inductive bias—not training objective—governs representational similarity: two ViTs differing only in objective reach CKA 0.83, whereas a ViT and a Swin differing only in architecture reach 0.25. We then show that this measured divergence does not, on its own, predict fusion utility. In our pool the most representationally divergent teacher is also the weakest individually, and fusing it yields a worse result, on every dataset, than fusing a representationally redundant but individually strong teacher; across the teacher additions we examine, individual competence tracks fusion utility while representational distinctness does not. Because divergence and competence are entangled in our encoders, we present this as a concrete demonstration that diversity is not sufficient for utility rather than a general law, and argue that competence must be weighed alongside distinctness in any similarity-based selection. Building on this, a lightweight (11–31 parameter) class-conditional logit fusion improves accuracy over the strongest single model by +2.99% pooled (p < 10^-4, Cohen's d = 3.70) and reduces expected calibration error by 85–96%, with both effects persisting under cross-institutional distribution shift, while fusion architectures with up to 5×10^5 parameters fail to exceed the dominant teacher. All results are reproducible from released features, splits, and code.

URL: https://openreview.net/forum?id=pstfJYpAtl

---

Title: PhyX: Does Your Model Have the "Wits" for Physical Reasoning?

Abstract: Existing benchmarks fail to capture a crucial aspect of intelligence: \textit{physical reasoning}, the integrated ability to combine domain knowledge, symbolic reasoning, and understanding of real-world constraints. To bridge this gap, we introduce PhyX, a large-scale benchmark designed to evaluate models’ physics-based reasoning in high-fidelity visual contexts. Additionally, while existing benchmarks often rely on clean 2D planar schematic diagrams where key relations are explicitly annotated, PhyX evaluates physics reasoning from non-canonical 3D viewpoints, introducing perspective, depth/occlusion, and scale cues that must be inferred before applying physics principles. PhyX comprises 3K carefully curated multimodal questions spanning 6 reasoning types, 25 sub-domains, and 6 core physics areas. Comprehensive experiments reveal that state-of-the-art models struggle substantially: GPT-o4-mini, Gemini-2.5-Pro, and GPT-5 achieve only 45.8%, 62.4%, and 65.2% accuracy, respectively, falling over 10% behind human experts. To assess potential data contamination, we conduct a prefix-only memorization probe and observe only 0.2% accuracy, suggesting negligible contamination. Our analysis highlights three key weaknesses in current models: over-reliance on memorized knowledge, dependence on mathematical formulation, and superficial visual pattern matching rather than physical understanding. We further present fine-grained analysis and detailed case studies to dissect these limitations. To ensure reproducibility, we implement evaluation protocols based on widely used toolkits such as VLMEvalKit and lmms-eval, supporting one-click benchmarking. All code, scripts and datasets used in this work are provided in the supplementary materials.

URL: https://openreview.net/forum?id=KX7GnW0mDs

---

Title: Is the Verifier Worth Its Own Compute? An Equal-FLOP Analysis of Test-Time Verification for Mathematical Reasoning

Abstract: Test-time scaling improves the reasoning accuracy of language models by drawing many samples and selecting one. Two selection strategies dominate: self-consistency (majority vote, free) and verifier-based selection with a learned reward model (here a Process Reward Model, PRM). The verifier almost always wins at a fixed number of samples $N$ — but the verifier is itself a large neural network, and scoring $N$ candidates is not free. We argue that the operative question is not "does the verifier help at equal $N$?" but "does the verifier help at equal compute?" A PRM-scored sample costs the generation FLOPs plus the scoring FLOPs; for a 3B generator and a 7B verifier this is $\approx 3.33\times$ the cost of one generation-only sample, so an honest comparison gives self-consistency that many more samples. We formalize a coverage/realized/verifier-efficiency decomposition with unbiased estimators and bootstrap confidence intervals, and an equal-FLOP comparison on a common compute axis. On GSM8K the verifier earns its compute from a small budget ($k \approx 4$); MATH-500 ($N=64$) is harder and more revealing. Across five generators spanning four families (Qwen, Phi, Llama, Mistral, gemma; the 7B+ models in 4-bit), the verifier's compute-worth follows a simple, measurable rule — its per-sample selection advantage over majority vote, $\Delta$, weighed against its cost ratio $r$: the verifier is worth its compute only once $\Delta$ is large enough to beat $r$. For competent generators $\Delta$ is modest, so at small budgets self-consistency significantly beats the verifier — when compute is scarce, more samples win — and the crossover, where it occurs at all, ranges from $k=8$ to beyond a practical budget. The same rule accounts for the exception: Mistral-7B's self-consistency is nearly useless, so its verifier wins from $k=2$. $\Delta$ is decisive, not the tax: gemma has the lowest $r$ yet its verifier never wins (tiny $\Delta$), while Mistral wins early (real $\Delta$). The finding is robust across four PRM step-aggregations and a second, different-type (outcome) reward model.

URL: https://openreview.net/forum?id=ZADuwPdnl3

---

Title: RocketPFN: Accurate Time Series Classification via In-Context Learning

Abstract: We introduce RocketPFN, a training-free pipeline for time series classification that combines random convolutional feature extraction (Rocket) with in-context classification via a pretrained tabular foundation model (TabPFN v2.5). On 92 UCR datasets (30-resample protocol), RocketPFN matches HC2, the strongest published method on the archive, in mean accuracy (both 0.900, Wilcoxon $p=0.50$), with no training on the target data and a median inference time of 30 seconds per fold. It also significantly outperforms every individual classifier in the HC2 ensemble. On UEA (20 datasets) the difference is likewise not statistically significant. A separate comparison concerns TSC foundation models: when paired with the same downstream classifier, MOMENT, Mantis, and MantisV2 are all significantly outperformed by RocketPFN using fewer extracted features and no learned parameters ($p<0.001$ in each case). This holds even when the encoders were pretrained on corpora that include the UCR training samples. We propose this two-stage pipeline as a reference point for evaluating zero-shot TSC foundation models.

URL: https://openreview.net/forum?id=eDI4hLNKJN

---

Title: UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity

Abstract: The Segment Anything Model (SAM) family has become a widely adopted vision foundation model, but its ability to control segmentation granularity remains limited. Users often need to refine results manually — by adding more prompts or selecting from pre-generated masks — to achieve the desired level of detail. This process can be ambiguous, as the same prompt may correspond to several plausible masks, and collecting dense annotations across all granularities is prohibitively expensive, making supervised solutions infeasible. To address this limitation, we introduce UnSAMv2 which enables segment anything at any granularity without human annotations. UnSAMv2 extends the divide-and-conquer strategy of UnSAM by discovering abundant mask-granularity pairs and introducing a novel granularity control embedding that enables precise, continuous control over segmentation scale. Remarkably, with only 6$K$ unlabeled images and 0.02\% additional parameters, UnSAMv2 substantially enhances SAM-2, achieving segment anything at any granularity across interactive, whole-image, and video segmentation tasks. Furthermore, we introduce a language-guided segmentation pipeline that treats granularity as a predictable variable by grounding it in text semantics via an external VLM, improving previous SOTA in 1-IoU (40.4 $\rightarrow$ 47.3). Across 11+ benchmarks, UnSAMv2 achieves consistent gains over SAM/SAM-2 baselines, improving $\text{NoC}_{90}$ (5.69 $\rightarrow$ 4.75), 1-IoU (58.0 $\rightarrow$ 73.1), and $\text{AR}\_{1000}$ (49.6 $\rightarrow$ 68.3) by a large margin. Our work demonstrates that a small amount of unlabeled data can deliver large, practical improvements to strong supervised baselines!

URL: https://openreview.net/forum?id=TJyPRRrLIV

---

Title: Graph-Based Semi-Supervised Learning via $p$-Conductances

Abstract: We introduce \emph{$p$-conductance learning}, a graph-based semi-supervised learning method in which each class is represented by a probability measure, and predictions are obtained from potentials minimizing a graph $p$-energy subject to an affine measure-separation constraint. From this viewpoint, our approach connects to several familiar graph optimization problems: for $p=1$ it admits a generalized \texttt{maxflow-mincut} characterization, for $p=2$ it admits an effective resistance characterization and a direct connection to Poisson learning, and for $p=\infty$ it is characterized by the reciprocal of a $1$-Wasserstein distance on the graph. To implement the program at scale for general $p$, we develop a semismooth Newton augmented Lagrangian method and prove global convergence of the outer iteration. For the semismooth Newton step, we prove local superlinear convergence for $p\in\{1,2,\dotsc\}$ and $p=\infty$. Experiments on citation and image benchmarks, as well as specialized settings with corrupted and partial label information, show that the method is consistently competitive overall and, for $p=2$ with class-size-aware decoding or MBO refinement, often attains the best or tied-best average accuracy among the compared graph-based baselines.

URL: https://openreview.net/forum?id=AdxTAbz4li

---

Title: Emergence of Preferential Attachment and Glass-Ceiling Effects in Autonomous Networks of LLMs

Abstract: We investigate the emergence of structural disparities in networks comprising large language model (LLM) agents. Each LLM agent refers to a prompted LLM of a specified type determined by its base model, model size, and system prompt. When LLM agents autonomously choose collaborators, the resulting communication network exhibits preferential-attachment dynamics: agents that are already prominent become increasingly likely to attract additional connections. In some cases, weaker LLM agents (agents with smaller base model or older version) can disproportionately occupy central and influential network positions relative to stronger LLM agents. We interpret this misalignment between task capability and network prominence as a type-dependent glass-ceiling effect (GCE).

We model the network of LLM agents as a time-evolving sequence of directed weighted graphs, where the vector-valued edge weights represent cumulative tokens exchanged, number of interaction rounds, and reasoning effort. Using a contraction mapping argument on the mean-field dynamics, we prove that the importance (centrality) of each agent type converges to a unique stable equilibrium. To anchor the model in LLM decision mechanisms, we introduce a cross-attention-inspired utility for collaborator selection. This utility specifies the local connection dynamics and, together with the mean-field model, yields a predictive characterization of the limiting network structure and its type-dependent centrality gaps.

To validate the theory, we develop an experimental testbed with 100 LLM agents. Our experiments show that autonomous network formation can generate persistent centrality disparities, with their magnitude and direction depending on model family, model size, system-prompt design, and task context. They further show that the effect of preferential attachment depends on its alignment with model capability: reinforcing it improves collective performance when stronger agents become central, whereas weakening it improves performance when network dynamics instead favor weaker agents.

URL: https://openreview.net/forum?id=hzDXIjjGv0

---

Title: The Dynamics of Task Interactions in Deep Linear Multi-Task Networks

Abstract: Despite significant empirical progress in multi-task learning (MTL), a theoretical understanding of task interactions and their learning dynamics remains limited. We address this gap by deriving analytical solutions for the gradient-flow dynamics of linear multi-task networks, characterising how shared and task-specific components evolve during training and at convergence. We show that task alignment and magnitude imbalance in the data determine how task-specific functions, losses, and neural representations evolve relative to one another, and that standard measures of negative task interference reflect this structure. Our analysis provides a theoretical explanation for the dynamic nature of task interactions and motivates dynamic, data-dependent loss weighting in linear MTL. We validate these predictions on the Multi-MNIST dataset and nonlinear networks. These results establish a theoretical foundation for understanding task interactions and pave the way toward principled algorithms for task weighting and grouping.

URL: https://openreview.net/forum?id=zUXzsLQt24

---

Title: Structure Matters More than Flexibility: Modeling Covariances for Efficient Approximate Inference

Abstract: Approximate Inference in Bayesian Neural Networks often assumes that increasing posterior flexibility - captured foremost in terms of the number of independent learnable parameters - is the primary driver of performance. We challenge this assumption, demonstrating that covariance structure is more critical than raw parameter count. We introduce the Circulant-Constrained Gaussian, a novel posterior family that leverages the spectral properties of circulant matrices to model rich weight correlations with minimal overhead. By parameterizing the covariance with a single generating kernel and utilizing the Fast Fourier Transform, our method
matches or surpasses the performance of other structured variational approximations in terms of calibration and out-of-distribution detection while requiring up to 20,000 times fewer parameters. Our results suggest that the inductive bias of a structured posterior, rather than its degrees of freedom, is the key to effective and efficient Bayesian Deep Learning.

URL: https://openreview.net/forum?id=wmbKyEhFSh

---

Title: Composition or Magnitude? A Cost-Structure Dichotomy for Epistemic Uncertainty in Deferral Decisions

Abstract: Predictive systems routinely collapse uncertainty into a single scalar of the marginal predictive distribution—entropy, variance, or confidence—and threshold it to decide whether to predict, abstain, or defer. We sharpen a folklore objection to this practice: any statistic of the marginal is insufficient when the fallback action can resolve epistemic uncertainty. This is not, however, an argument for vectors over scalars: a scalar statistic of the ensemble— mutual information I[Y ;M |x] or the epistemic proportion EP = E/(A+E)—already resolves the canonical counterexample. The right question is which decomposition-aware statistic to threshold, and when. We answer with a cost-structure dichotomy: for a binary deferral problem in which resolution removes the epistemic part of the loss at cost κ(x), the Bayes-optimal rule thresholds the raw epistemic uncertainty E(x) when κ is fixed, and the normalized proportion EP(x) when κ scales with total uncertainty; the mismatched statistic incurs regret whenever E and EP are not comonotone on the instance distribution— a measurable property we quantify on CIFAR-10 (≈10% discordant pairs; mismatch regret up to 2.3%). A deliberately double-sided study confirms the boundary. On non-resolution tasks (selective classification on disjoint independent ensembles; OOD detection across three shift types), EP is worse than magnitude scores, as predicted. On DermaMNIST triage, a decomposition-aware three-action policy routes a statistically different class mixture to the expert than abstention does—after Benjamini–Hochberg correction only vascular lesions (4.8×) and basal-cell carcinoma (1.9×) survive as over-referred, and we disclose a clinical misroute (pre-malignant actinic keratoses land in abstention)—while being no cheaper in aggregate and no more unique than an I+H policy. A CIFAR-10H faithfulness check finds estimated aleatoric and epistemic components entangled, with both correlations collapsing to zero on the top-uncertainty decile—a caveat we carry into every claim. Our contribution is a clean statement of when to normalize epistemic uncertainty, and honest evidence for its boundary.

URL: https://openreview.net/forum?id=jOpNCqXY1X

---

Title: Adoption-Aware Crop Recommendation: Expected Realized Benefit as a Direct-Method Offline Policy under Unidentifiable Compliance

Abstract: A recommendation only helps if it is taken. Most recommender systems, though, are trained
on observational data to predict an outcome and then deployed to pick an action for each
user as if that action will be followed. We study the resulting gap in smallholder agricul-
ture, where the standard approach recommends the (crop, input) bundle with the high-
est predicted yield and says nothing about whether the farmer will adopt it. Our ob-
jective, Expected Realized Benefit (ERB), instead recommends the action with the
largest predicted yield gain weighted by the probability the farmer adopts it. We build
an ERB-maximizing recommender from two calibrated models, one for yield and one for
adoption, over a five-tier action space, and we treat it as a direct-method off-policy esti-
mator: farmer-side propensities are not identifiable from survey data, so inverse-propensity
and doubly-robust evaluation simply do not apply. On a real LSMS-ISA Ethiopia panel
(6,770 households), the adoption-aware policy beats a conventional accuracy-only baseline
by +50 kg/ha (cluster-robust 95% CI [+33, +70], cluster permutation p < 10−3
, n = 800
enumeration-area-disjoint test plots). Every number here is a model-believed offline estimate,
not a field-measured return. The gain holds up under four sensitivity sweeps, a Rosenbaum
analysis (Γ ≥ 5), a marginal-sensitivity-model pessimistic policy (out to Λ = 5), a partial-
identification floor anchored to the Duflo–Kremer–Robinson Kenya fertilizer RCT (+26
kg/ha [+13, +39]), and four CATE-metalearner baselines that, without an adoption model,
never recover it. Reported conservatively—propagating the yield split-conformal interval
and a bootstrap uncertainty set on the adoption head jointly—the policy still guarantees a
+18 kg/ha [+9, +29] lift over status quo, and the headline is robust to ±50% perturbation
of the yield-head literature priors (not only the adoption penalties). An education-shuffle
placebo shows the per-subgroup gradient is descriptive rather than identified, so we treat the
aggregate gain as the central claim. Validating the adoption model against the same RCT
then exposes the paper’s principal limitation: the model gets the ordering of the treatment
arms right but is badly miscalibrated in absolute terms on out-of-distribution actions, off by
roughly 50 percentage points. We frame that gap as a concrete benchmark for future work
that fuses observational and experimental evidence, and we release the code, data pipelines,
and RCT-validation harness.

URL: https://openreview.net/forum?id=Pt1GIJYUIM

---

Title: Orchestrating LLMs as Hierarchical Multi-Agent Reinforcement Learning System for Automotive Software Development

Abstract: Software-defined vehicles depend on firmware that must evolve continuously and safely, yet general-purpose LLM coding agents lack the architectural mechanisms required for safety-critical cyber-physical systems. We introduce **AutoEvolve**, a Hierarchical Multi-Agent Reinforcement Learning (H-MARL) framework with three contributions: (i) jointly learned Orchestrator and sub-agent (Data, Requirements, Code) policies, where an *adversarial* Requirements Agent rejects unsafe candidates rather than merely critiquing them; (ii) an offline-to-online curriculum that initializes via SFT on historical development trajectories and refines via PPO in a *shadow-mode* deployment running parallel to human engineers, treating their commits as delayed supervision; and (iii) a dual-reward decomposition $R_{fast} + \lambda R_{slow}$ that anchors policies to deterministic verification while regularizing toward maintainable, human-aligned style. On an internal benchmark, AutoEvolve attains the highest Success Rate (**61.4%** vs. 53.5 - 55.1% for MetaGPT and SWE-Agent at the same Llama-3-70B backbone) and reduces the Requirement Violation Rate to **1.6%** - ~7x fewer violations than task-agnostic multi-agent frameworks and ~10x fewer than a Monolithic ReAct baseline. The architectural contribution alone (SFT-Only AutoEvolve at 4.7% violations) drives most of the safety gain; online RL provides the remaining performance lift. Active Development Cycle Time drops from ~5 days to < 6 hours, framing safety-critical software evolution as *safe* agentic coding, not raw generation.

URL: https://openreview.net/forum?id=Aw3HE36fUP

---

Title: Memorized but Not Realized: Frontier Language Models Precisely Recall Financial Fundamentals, but It Does Not Become Trading Skill

Abstract: LLM "AI-trading" benchmarks are routinely scored by replaying historical markets, raising a contamination concern: a model that has memorized an asset's realized outcome cannot generate honest out-of-sample skill, so backtest performance may be recall rather than foresight. We build a white-box-to-black-box validated leakage instrument and use it to ask the decision-relevant question prior work leaves open: does memorization capacity actually transmit to trading decisions? We report a dissociation. Capacity is large, precise, and universal: across 16 open models (10 vendors) and three commercial APIs, models prefer the realized revenue over magnitude-matched ($\pm 15\%$) counterfactuals at 0.23–0.81 (chance 0.20) and freely generate the figure within 10% for up to 89% of famous firm-years, while no-leak chronologically-consistent baselines sit at chance. Recall is cutoff-bounded: exploiting the models' different training cutoffs, it collapses to chance for facts dated after a model's cutoff (within-model $t=-6.6$, 18/19 models), the design-based signature of memorization rather than estimation. Applied to the realized return itself, the variable a backtest fears, recall is at chance with no cutoff signature: these models memorize fundamentals, not prices. Realization, by contrast, is null across five independent tests, including a cross-sectional ranking task ($N\approx 5{,}700$ per model) and a realistic StockBench-style backtest over six open and frontier-API agents; an apparent agentic signal (+9.9pp, $t=4.07$) is shown to be a firm-prominence confound that vanishes under firm fixed effects ($\beta=-0.018$). A pre-registered, dependence-corrected power analysis bounds any residual post-cutoff skill: on a $\approx$2,000-name post-cutoff universe the minimum detectable mean $|\mathrm{IC}|$ is $\approx$0.012–0.016, below the 0.02–0.05 band where real strategies operate. The contribution is a validated leakage instrument, an identification result showing why memorization is invisible to any pre-cutoff backtest, and the finding that, for the signals we test, LLM memorizatcome measurable market skill.

URL: https://openreview.net/forum?id=jvxqEAiVcj

---

Title: Revisiting Deep BSDE-Based PDE Solvers: A Reproducibility and Robustness Study

Abstract: We revisit deep BSDE-based methods for solving high-dimensional semilinear parabolic partial differential equations through a reproducibility and robustness study of the original Deep BSDE solver and its backward dynamic programming variants. We present a modular JAX-based implementation framework for stochastic differential equation simulation, BSDE discretization, and neural PDE solvers, designed for reproducible experimentation in the modern scientific machine learning ecosystem. Our study reproduces key benchmark problems including Allen-Cahn, Hamilton–Jacobi–Bellman, and nonlinear reaction–diffusion equations, while systematically analyzing optimization stability, discretization effects, gradient parameterizations, random seed sensitivity, and architectural choices across multiple solver variants. In addition to reproducing published results, we investigate failure modes such as divergence and representational limitations in high-dimensional settings. We further compare global forward-shooting approaches against backward dynamic programming formulations and discuss their tradeoffs in stability, scalability, and computational efficiency. Our work aims to provide both a reproducible experimental baseline and an extensible open-source foundation for future research on BSDE-based PDE solvers in JAX.

URL: https://openreview.net/forum?id=zjgNPTaCre

---

Title: Two Routes to Length Generalization: Readout Geometry and the Causal-Mask Latch in Diffusion vs. Autoregressive Recognizers

Abstract: A masked diffusion language model (MDLM) is bidirectional and reads its answer out without a causal final-position bottleneck, so one expects it to help length generalization where an autoregressive (AR) transformer's final-token readout falls outside the trained position range, most clearly on decide-late (suffix) regular languages, which collapse out-of-distribution (OOD) under AR with learned positions. We test this in a matched pipeline: every regular language is a recognition task (one maskable ACCEPT/REJECT label), and the AR baseline is the same model with a single flag flipped (causal vs. bidirectional attention), on a suite of regular-language recognition tasks including new prefix/fixed-position tasks (Appendix A), across three sizes and three seeds. This expectation does not hold: across 250 matched cells the paired test-long difference ($\mathrm{MDLM} - \mathrm{AR}$) is centered below zero (AR beats MDLM in 83 cells vs. 30 at $|\Delta| > 0.05$; sign-test $p < 10^{-6}$), with the clearest gap on prefix tasks. We pre-specified, and report, this failure. The reason is a positive account: OOD success requires the readout to reach the decision-relevant tokens by one of two routes, either an in-range latch (a finite decision carried position-invariantly to the readout; available to a causal model under any encoding, so the causal mask is itself a latch route) or RoPE adjacency (readout at a fixed relative offset). A pre-specified label-position comparison flips sign between suffix and prefix as predicted; a NoPE column confirms the causal-mask latch; probes show the DFA state stays decodable where the readout collapses. The operative variable is attention direction, not diffusion training (a genuine masked-diffusion LM behaves the same, if anything weaker). Separately, for the counting/group class, fixed-depth computation (a weight-tied looped denoiser, one width plus a wider-block check) monotonically rescues in-distribution acquisition of parity/mod-$k$ but never length-generalizes; untied-depth controls show the driver is effective depth, not recurrence, and only a length-scaling, position-indexed recurrence is expected to close the gap.

URL: https://openreview.net/forum?id=VrqgkXrfpd

---

Title: FlexBrain: A Unified Study of Routing in Conditional Computation

Abstract: Conditional computation aims to improve neural network efficiency by routing each input through only the computation it needs. However, comparing routing strategies remains difficult because existing methods are typically evaluated in isolated frameworks, preventing standardized and fair comparisons. To address this, we introduce FlexBrain, a unified framework for comparing conditional computation methods under shared architectures, datasets and compute budgets. Using this standardized setup, we present an empirical analysis, comparing dense baselines and six conditional methods across language and vision tasks. In language modeling, we find that conditional computation is most beneficial when dense models are undertrained, while dense computation becomes increasingly competitive as training compute grows. In the vision domain, conditional computation reduces inference FLOPs while maintaining accuracy close to dense baselines. Overall, our results show that routing choices interact strongly with model scale and compute budget, and that unified evaluation is essential for making these tradeoffs explicit.

URL: https://openreview.net/forum?id=iYtbuUpFqY

---

Title: AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles; Aha-Euler turns Project-Euler-style mathematical ideas into generated taught/held-out tasks with exact validators; and Aha-Vending, an open-source implementation inspired by Vending-Bench, tests whether a simulated vending agent remains profitable while handling delayed feedback and operational incidents. AhaBench reports a three-part scorecard: Initial Score measures starting competence, Post-Experience Score measures the later empirical outcome, and Learning Lift is their difference. This decomposition is the main empirical message: models that use visible support well, models that reach high post-experience scores, and models that improve most during a run are not always the same. On the common eight-model panel, Claude Opus 4.6 leads aggregate Post-Experience Score at 64.3 and aggregate Learning Lift at +25.8, with Gemini 3.1 Pro close behind at 63.4. The component results explain the split: puzzle traces raise supported scores but often fail to become no-hint exploration behavior; Aha-Euler full teaching reaches 78.6-100.0% while answer-only transfer ranges from 0.0 to 73.9%; and Aha-Vending separates profitable incident handling from bankruptcy and no-order failure. We release benchmark tasks, rubrics, generators, validators, simulator code, and model-evaluation interfaces for evaluating new agents.

URL: https://openreview.net/forum?id=7h5Q0cGAEu

---

Title: Optimised Sequential Testing for Binary Ensemble Classifiers

Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote.
A classic example is random forests, which combine the predictions of decision trees.
Ensembles that use more base models can be more accurate but also more costly to train and run.
In this paper, we consider strategies for reducing the computational cost of binary classification using an approach from the field of sequential testing.
Rather than evaluating all the base models and taking a majority vote, we evaluate the base models sequentially and stop execution when a clear majority emerges.
We consider three different notions of optimality for early-stopping strategies that minimize the number of base models executed while controlling the rate of disagreement with the full ensemble.
For each notion of optimality and allowable disagreement rate, we show that a linear program can be constructed and solved efficiently to find the optimal stopping strategy.
We tested these methods on real-world datasets taken from the UC Irvine Machine Learning repository, and on the benchmark datasets proposed by Grinsztajn et al. We found that on most datasets, these methods provide speed-ups of 4x or more while controlling disagreement at 0.1%.

URL: https://openreview.net/forum?id=aUQFMWmW0v

---

Title: Fourier Analysis on the Boolean Hypercube via Hoeffding Functional Decomposition

Abstract: Fourier analysis on the Boolean hypercube is classically defined as the orthogonal decomposition of pseudo-Boolean functions with respect to the uniform probability measure. This uniformity assumption is, however, rarely met in practice, where binary data are governed by dependencies and deterministic constraints, which makes the standard framework too restrictive for most machine learning settings. In this work, we generalize Fourier analysis on the Boolean hypercube to arbitrary probability measures, building on the generalized Hoeffding (or ANOVA) functional decomposition. We first show that, when the probability distribution has full support, this generalization admits a unique and well-posed decomposition, for which we provide an explicit basis extending the Walsh-Hadamard parity functions; we further recast its computation as a least-squares problem. However, this full-support assumption is idealized: real life datasets are sparse, with an effective number of configurations far smaller than $2^d$, which leaves the decomposition under-determined. We propose to address this practical regime by introducing an Elastic Net penalization that smartly exploits the sparsity of the decomposition, which also mitigates the curse of dimensionality. The resulting framework handles the non-uniform configuration spaces that arise in real-world tasks, e.g. one-hot encoded features. Finally, we illustrate the potential of our framework for explainable AI. We observe that a low-order interaction reconstruction (such as GAM or GA\textsuperscript{2}M) already captures most of the signal of the learned models, supporting the sparsity-of-effects assumption in practice. The corresponding ANOVA components then yield feature attributions closely aligned with widely used SHAP-based indicators, obtained directly from the sparse terms of the decomposition rather than through coalition sampling.

URL: https://openreview.net/forum?id=9OVrfzJE6m

---

Title: Learning under Heterogeneous Shifts: Object-Conditional Meta-Alignment for Domain-Generalized Detection

Abstract: Domain generalization for object detection (DGOD) aims to train object detection models that remain robust under unseen data shifts.
Most existing DGOD methods primarily target the single-source setting, where generalization is driven by heavy data augmentation and coarse class-level alignment. However, their performance deteriorates under large, heterogeneous shifts, where distribution changes are localized to object instances and their surrounding context, so class-level alignment fails to address the region-level discrepancies critical for detection. To address this fundamental challenge, we present a novel method for efficiently training the generalized detector under large heterogeneous domain shifts. First, we design a meta-learning based scheme, which efficiently simulates domain shift via episodic meta-train/meta-test splits across source domains, to learn shift-invariant semantic features. Secondly, we introduce an object-centric, likelihood-based similarity for cross-domain samples that underpins a cross-domain feature alignment framework to close the distributional gap induced by heterogeneous domain shifts. The results of comprehensive experiments on a large number of DGOD benchmarks show that our method outperforms previous DGOD methods. Code and prepared data are publicly available on Github (temporally available in supplementary materials).

URL: https://openreview.net/forum?id=JpTRrPUAYf

---

Title: Operator-informed score matching for Markov diffusion models

Abstract: Diffusion models strive to invert a noising process via score matching, a learning objective agnostic to the underlying spatial and temporal dependencies given by an incremental noising schedule. This paper studies the spectral properties of the infinitesimal generators that govern Markov noising processes for better-informed score matching. Notably, we illustrate that it is possible to (i) reformulate the implicit score matching loss of Hyvärinen from an operator-centric viewpoint and (ii) obtain training-free parametric score estimates for all noise levels, using only sample averages with respect to the data distribution. The resulting operator-informed score matching provides both a standalone approach to sample generation for low-dimensional distributions, as well as a recipe for aiding the training of neural score estimators in practical high-dimensional settings.

URL: https://openreview.net/forum?id=25zN49LA92

---

Title: $\bm{\lambda}$-VAE: Variance Equalization for Posterior Collapse

Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a
failure mode in which the approximate posterior converges to the prior,
rendering the latent code uninformative. Despite extensive research, a unified
account of why collapse occurs has remained an open question. We identify and
formalize two logically independent but coupled causes. \emph{Gradient
imbalance} occurs when the decoder's reconstruction signal vanishes faster than
the $\mathbb{KL}$ regularization pressure as the posterior widens.
\emph{Information gap} occurs when the stochastic sampling step discards a
substantial fraction of the encoder's computed representation, attenuating
decoder sensitivity and making collapse inexpensive. Both causes share the same
collapse trajectory, and we show that the information gap is algebraically
equivalent to mismatch between the aggregate posterior and the prior, unifying
two pathologies. Subsequently, we introduce $\lambda$-VAE, which resolves both
causes through a single modification to the reparameterization step: the
sampling noise is scaled by per-dimension exponent, while the $\mathbb{KL}$
penalty retains the original posterior variance. This asymmetry shifts the
stable training attractor away from the degenerate collapsed state, driving all
latent dimensions toward the same equilibrium -- a mechanism we term
\emph{variance equalization}. A closed-form optimal exponent per dimension
follows from a net information gain objective, with a single hyperparameter
controlling the reconstruction--generation tradeoff. We validate on standard
benchmarks (Binary MNIST, Binary Omniglot, CIFAR-10, CelebA-64), showing
consistent reductions in collapsed dimensions, information capacity gains of up
to $2.8\times$ nats, and reconstruction quality improvements of up to $+0.33$ BPD.

URL: https://openreview.net/forum?id=SEPf31zX9P

---

Title: Cognitive Topology: Hallucinations, Transformativeness, Prompt Irrelevance, and Exploratory Generative Cognition

Abstract: The success of state-of-the-art large language models provides a unique opportunity to develop a cognitively grounded approach to exploratory generative cognition—the generation of novel and, in particular, transformative narratives and other artifacts. This requires well-defined and operational notions of key cognitive concepts, including novelty, context, transformativeness, and hallucination. We show that every space of narratives carries a context-relevant cognitive topology and an emergent extended metric derived from that topology. The topology allows us to introduce cognitive indistinguishability, rather than embedding distance, as the primitive relation, so that semantic separation is induced by context rather than imposed by a model-dependent representation. In real contexts, the cognitive topology is generally not a manifold topology. This topology forms a fundamental component of a new analytical framework for studying generative AI processes and artifacts. The framework provides an approach for quantifying the transformativeness of generative models and narrative validity within a context, while establishing formal boundaries between generative creativity, discovery, and hallucinations, which are identified as regions of minimal validity. The analysis reveals new learning tasks: estimating cognitive distance and encounter measures, computing or approximating narrative validity, stabilizing generative models under cognitive indistinguishability, and supporting purpose-driven context restructuring.

URL: https://openreview.net/forum?id=qilkmTy1XT

---

Title: MRS: Multi-Resolution Subgoals for HRL Agents

Abstract: Hierarchical reinforcement learning (HRL) decomposes the policy into a manager and a worker, which enables long-horizon planning but introduces a performance gap on tasks requiring agility.
We identify one of the root causes in subgoal-based HRL: the manager's goal representation is typically learned without constraints on reachability or temporal distance from the current state, preventing precise local subgoal selection.
We further show that the optimal subgoal distance is both task- and state-dependent: nearby subgoals enable precise control but amplify prediction noise, while distant subgoals produce smoother motion at the cost of geometric precision.
We propose Multi-Resolution Subgoals (MRS), which learns multiple goal-prediction modules, each specialized to a fixed temporal horizon, with a jointly trained meta-controller that selects among them based on the current state.
MRS consistently outperforms fixed-resolution baselines and reduces the performance gap between HRL and non-HRL state-of-the-art on DeepMind Control Suite, Gym-Robotics, and long-horizon AntMaze tasks.
[Project page: \url{https://sites.google.com/view/multi-res-skills/home}]

URL: https://openreview.net/forum?id=m05mZ2VYKv

---

Title: The Dog Ate Our Benchmark: Generalizability and Leakage in Fine-Grained Image Classification Benchmarking

Abstract: Dog breed identification is a classic benchmark in fine-grained image classification (FGIC). Several recent studies have reported top-1 accuracies over 90%. To examine whether these results generalize across benchmarks within the same domain, we replicated a recent workflow that achieved 94.45% accuracy on Stanford Dogs and extended it to the Columbia Dogs and Tsinghua Dogs benchmarks. Performance dropped to 91.20% and 87.17%, respectively, suggesting that high accuracy on Stanford Dogs does not necessarily reflect broad domain generalization.

The replicated workflow used frozen features from four convolutional neural networks (CNNs), all distributed as ImageNet-1K-pretrained models. Because Stanford Dogs is derived from ImageNet, we audited image provenance relative to the ImageNet-1K train split using SHA-256 hashing. We found 100% of the unique images in the Stanford Dogs train and test splits were byte-identical to ImageNet-1K training images. This overlap also affects Tsinghua Dogs, where 23.8% of training images and 59.6% of validation images appear in ImageNet-1K train. Overlap in Columbia Dogs was much lower, at 1.5% across the analyzed splits. These results support leakage as a possible explanation for cross-benchmark performance differences.

We further tested this explanation by replacing the CNN feature extractors with OpenCLIP and DINOv2, models with different pretraining provenance. These models achieved higher accuracy on Columbia Dogs than on either ImageNet-1K-contaminated benchmark. These findings raise concerns about a common FGIC practice: combining ImageNet-1K pretraining with evaluation on ImageNet-derived benchmarks. Our provenance audit shows that estimates of FGIC accuracy on dog breed benchmarks can reflect leakage between evaluation sets and pretraining corpora rather than independent task generalization.

URL: https://openreview.net/forum?id=368eHfeW7P

---

Title: Generative Learning on Quotient Spaces: A Geometric Diffusion Framework for Symmetry Reduction

Abstract: Diffusion models are commonly defined in ambient Euclidean spaces, even when the underlying data distribution is invariant to rotations, translations, poses, or other continuous symmetries. This creates a geometric mismatch: the stochastic process propagates noise along both intrinsic data directions and redundant symmetry orbits, forcing the score model to learn nuisance variation that carries no quotient-level information. We propose QDM, a quotient-geometric diffusion framework that makes invariance a property of the generative dynamics rather than only the score architecture. QDM defines diffusion on symmetry-reduced quotient manifolds and realizes the process in ambient coordinates through horizontal projection. The framework unifies exact, approximate, context-dependent, and mode-changing symmetries via soft horizontal projection and stratified quotient spaces. We establish the ambient--quotient measure correspondence, prove well-posedness of the lifted quotient diffusion, and derive a Fisher-information decomposition showing that quotient projection suppresses removable vertical score variance and reduces effective score-estimation complexity. Experiments on molecular conformer generation and legged robot locomotion show consistent gains in convergence, sample efficiency, geometric accuracy, and generation stability over ambient and equivariant diffusion baselines. These results suggest that symmetry-aware generative modeling is best achieved by changing the probability space and diffusion geometry itself, not only by imposing equivariance on the neural score field.

URL: https://openreview.net/forum?id=NFNkxfP8ic

---

Title: Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs

Abstract: LLMs are now an integral part of information retrieval. As such, their role as question-answering chatbots raises significant concerns due to their demonstrated vulnerability to adversarial man-in-the-middle (MitM) attacks. Here, we propose our fundamental attack evaluation on LLM factual memory under prompt injection via Xmera, a theory-grounded factual MitM framework. By perturbing the input to "victim" LLMs across three closed-book, fact-based QA settings, we undermine the correctness of their responses and assess the uncertainty in their generation process. Surprisingly, trivial instruction-based attacks achieve the highest success rate (~85.3%) while simultaneously exhibiting high uncertainty in incorrectly answered questions. To provide a simple defense mechanism against Xmera, we train Random Forest classifiers on the response uncertainty levels to distinguish between attacked and unattacked queries (average AUC of up to ~96%). We believe that signaling users to be cautious about the answers they receive from black-box and potentially corrupt LLMs is a first checkpoint toward user cyberspace safety.

URL: https://openreview.net/forum?id=cL4Ir7kXyZ

---

Reply all
Reply to author
Forward
0 new messages