Skip to content

arXiv cs.LG - 2026-08-31 ​

210 items collected.


1. Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization ​

Author: Junhao Cao, Hongyi Xia, Jianian Wu, Xiaopeng Yi, Lixia Huang, Ping Guo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27507v1 Announce Type: new Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment. However, its pooled team-entropy score measures only collective ...

📖 Read original article


2. Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap ​

Author: Jacopo Dardini, Claudio Stanzione, Giordano Col`o, Giuseppe Fenza
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR

arXiv:2608.27512v1 Announce Type: new Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluatio...

📖 Read original article


3. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization ​

Author: Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27513v1 Announce Type: new Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the K...

📖 Read original article


4. A Deeper Analysis of Block-Sparse Featurizers ​

Author: Alexandru-Iulius Jerpelea, Amith Ananthram
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.27515v1 Announce Type: new Abstract: The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subspace (a block of directions) rather than a single direction. It is designed for features that live on l...

📖 Read original article


5. When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging ​

Author: Shangge Liu, Yuehan Yin, Yinghuan Shi, Lei Wang, Wenbin Li
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27518v1 Announce Type: new Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error. In the literature, these difficulties are me...

📖 Read original article


6. Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics ​

Author: Till Muser, Giovanni Abati, Ivan Dokmani'c
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27521v1 Announce Type: new Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean. Applying these architectures on a regular lat-lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D ...

📖 Read original article


7. Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution ​

Author: Yingqi Feng, Yufei Tang, Min Shi, Xingquan Zhu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27574v1 Announce Type: new Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects. To date, a handful of multi-label graph learning methods exist, but...

📖 Read original article


8. Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification ​

Author: Alexandre L. M. Levada
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ML

arXiv:2608.27634v1 Announce Type: new Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality throughout the feature space. This assumption can be inadequate for data whose local geometry varies subs...

📖 Read original article


9. More Data Cannot Break a Symmetry: Identifiability by Design ​

Author: Jing Xu, Christopher Kanan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27651v1 Announce Type: new Abstract: Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist. The obvious diagnostic for t...

📖 Read original article


10. Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay ​

Author: Pujan Thapa, Alexander Ororbia, Travis Desell
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27662v1 Announce Type: new Abstract: This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributional statistics as well as encoder-decoder models for class-incremental learning. The proposed approach...

📖 Read original article


11. SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning ​

Author: Musa Shams
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27678v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) often uses trajectory structure for future-goal sampling and multi-step targets, yet logged trajectories may be partitioned for administrative reasons that do not correspond to termination. We intr...

📖 Read original article


12. SafeStep: An Interactive Demonstration of Semantic Communication for Pedestrian Safety Monitoring ​

Author: Christian McDowell, Andrea Panebianco, Jeremiah Yang, Sirin Chakraborty, Samuel Chamoun, Travis Ross, Yin Sun
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27688v1 Announce Type: new Abstract: In this paper, we develop SafeStep, an interactive browser-based semantic communication platform for live pedestrian safety monitoring. SafeStep extracts pedestrian information from four live traffic-camera feeds, transmits it through a semantic commun...

📖 Read original article


13. RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing ​

Author: Madhusudan Srinivasan, Namith Nishal Raphae
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE

arXiv:2608.27704v1 Announce Type: new Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version, creating regression faults that are costly to detect because verifying predictions against ground tr...

📖 Read original article


14. DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge ​

Author: Yiming Xie, Pinrui Yu, Geng Yuan, Xue Lin, Ningfang Mi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27713v1 Announce Type: new Abstract: Edge intelligence systems increasingly require model training and online inference to coexist on resource-constrained devices, while inference demand can vary substantially across tasks over time. This creates two coupled challenges: sufficient computa...

📖 Read original article


15. Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning ​

Author: Yiming Xie, Lili Su, Ningfang Mi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27715v1 Announce Type: new Abstract: Federated learning systems are increasingly deployed to facilitate collaborative model training across a heterogeneous client population. Existing practice mostly implicitly assumes that the aggregated client data distribution is representative of the ...

📖 Read original article


16. Leveraging a Foundation Model for the EEG-Based Diagnosis of Alzheimer's Disease ​

Author: Maggie Lin, Chung-Lin Hou, Tzyy-Ping Jung
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, q-bio.NC

arXiv:2608.27719v1 Announce Type: new Abstract: Biological heterogeneity in Alzheimer's Disease (AD) poses a critical diagnostic challenge, particularly for traditional linear methods that fail to capture non-linear neural dynamics. To address this, we propose a diagnostic framework utilizing the La...

📖 Read original article


17. Diffusion Distillation for Efficient Weather Ensembles ​

Author: Yiming Yang, Valentin Brekke, James Briant, Serge Guillas
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.AP

arXiv:2608.27728v1 Announce Type: new Abstract: Diffusion models generate skillful weather ensembles but require costly iterative sampling. We introduce a supervised energy-distance distillation method that compresses a multi-step diffusion teacher into a single-step student by aligning student fore...

📖 Read original article


18. The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs ​

Author: Eric Yeats, Brendan Kennedy, Loc Truong, John Buckheit, Jung Lee, Jesse Friedbaum, John Emanuello, Henry Kvinge
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.27750v1 Announce Type: new Abstract: The hidden states of large language models (LLMs) are known to capture rich information relating to model knowledge and behavior that can be hard to extract from examination of input and output alone. As LLM-based systems increasingly interface with th...

📖 Read original article


19. Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess ​

Author: Szymon Mi{\l}osz, Piotr Duch, Szymon Grabowski
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27757v1 Announce Type: new Abstract: Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc0) released Chessformer, distills the visit counts of an AlphaZero-style Monte Carlo Tree Search (MCT...

📖 Read original article


20. Fast Weight Attention for Continual Learning ​

Author: Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, stat.ML

arXiv:2608.27763v1 Announce Type: new Abstract: Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. Fo...

📖 Read original article


21. Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization ​

Author: Jianing Chen, Vajiheh Farhadi, Yan Li, Thomas La Porta
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2608.27791v1 Announce Type: new Abstract: Short-term load forecasting (STLF) provides essential information for numerous applications in modern power systems. However, accurate STLF often relies on fine-grained smart-meter data from distributed users, raising increasing concerns about data pri...

📖 Read original article


22. Node-wise Feature Encoding for Neural Performance Prediction ​

Author: Matthew Grenier, William Hammer, Andrew Heuer, Nikhil Krishna, Yi Wang, Ramtin Zand
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27794v1 Announce Type: new Abstract: As neural networks are increasingly deployed on resource constrained edge devices, accurate prediction of latency and energy is critical for efficient neural architecture search. Existing GNN and transformer based predictors achieve strong results but ...

📖 Read original article


23. Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning ​

Author: Sejong Oh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27821v1 Announce Type: new Abstract: Explainable artificial intelligence (XAI) increasingly calls for actionable counterfactual recourse, yet current methodologies face challenges related to causal invalidity, excessive cognitive burden, and predictive failure. Exhaustive causal search al...

📖 Read original article


24. FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling ​

Author: Jun Bai, Ruilin Wang, Yue Li
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2608.27856v1 Announce Type: new Abstract: Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows. However, agents deployed at individual hospitals remain constrained by institution-speci...

📖 Read original article


25. SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning ​

Author: Hao Wang, Siyu Zhang, Wei Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.27882v1 Announce Type: new Abstract: Tabular foundation models based on in-context learning have recently emerged as strong alternatives to task-specific model fitting. However, the current performance frontier remains dominated by attention-heavy architectures, where attention is used th...

📖 Read original article


26. Beyond Pairwise Graphs in Science: Hypergraph Adaptive Wavelet Operators for Parametric PDEs ​

Author: Rajat Sarkar, Venkataramana Runkana, Souvik Chakraborty
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph

arXiv:2608.27883v1 Announce Type: new Abstract: Physical systems are often modeled by solution operators that map input fields, parameters, geometries, or past states to steady or future physical states. Learning these maps is difficult, especially for time-dependent systems that must assimilate his...

📖 Read original article


27. There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation ​

Author: Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27885v1 Announce Type: new Abstract: Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are...

📖 Read original article


28. TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision ​

Author: Ji'an Lei, Jian Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27911v1 Announce Type: new Abstract: Agents with smaller language-model backbones are less expensive but can drift into persistent failure modes, whereas those with larger backbones are generally more reliable but more costly. This reliability-cost trade-off motivates routing methods that...

📖 Read original article


29. TI$^2$PS: A Topology-Informed Inverse Design Framework for Stochastic Multicellular Pattern Formation ​

Author: Kenji Komiya, Andrew Kailiang Jin, Ryo Nishikimi, Kunio Kashino
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27931v1 Announce Type: new Abstract: This study proposes a novel framework to estimate parameters for reproducing target multicellular patterns using an agent-based model (ABM). Two major challenges in multicellular ABMs are estimating cell-level parameters (agent-specific variables) and ...

📖 Read original article


30. Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs ​

Author: Bo Li, Xin Zheng, Ming Jin, Can Wang, Shirui Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27948v1 Announce Type: new Abstract: Test-time adaptation (TTA) on graphs aims to adapt a graph neural network (GNN) that is well-trained on the training graph to the test graph, which involves potential distribution shifts that may harm model generalization and test-time inference. While...

📖 Read original article


31. PhyMamba: Physics-Modulated Mamba for Robust Battery Health Prognostics ​

Author: Sara Sameer, Yunyi Zhao, Wei Zhang, Minggang Zeng, Wenqing Li, Man-Fai Ng, Yonggang Wen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2608.27978v1 Announce Type: new Abstract: Battery health prognostics is a core function in battery management systems (BMSs), yet long-horizon health forecasting from BMS signals remains challenging due to operating-condition dependency and sensor noise. In this paper, we propose PhyMamba, a t...

📖 Read original article


32. Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control? ​

Author: Xianyi Wu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27985v1 Announce Type: new Abstract: Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo (MC) control are usually presented as different methods. MCTS is described in the language of search (selection, expansion, simulation, and backup), whereas MC control is described in the langu...

📖 Read original article


33. A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint ​

Author: Artem Safronov
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28003v1 Announce Type: new Abstract: This paper proposes a layer bit allocation method for Gemma-3-1B, formulating the problem as performance maximization (latency decrease) given a degradation budget constraint (allowable level of generation quality loss). This approach is different from...

📖 Read original article


34. Exact Risk Ratios for Weighted Data Selection in Linear Regression ​

Author: Guangjian Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.TH

arXiv:2608.28007v1 Announce Type: new Abstract: Hanneke, Moran, Shlimovich and Yehudayoff (COLT 2025) posed the following open problem. A selector sees a finite dataset $D \subseteq \mathbb{R}^d \times \mathbb{R}$, picks at most $n$ examples together with nonnegative weights, and hands the weighted ...

📖 Read original article


35. When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood? ​

Author: Yansen Han, Hongxin Sun, Tao Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28010v1 Announce Type: new Abstract: Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and their old/new differences as log-likelihood ratios. We characterize when the...

📖 Read original article


36. Explainable Uncertainty Estimation for Reliable Medical AI ​

Author: Li Rong Wang, Jamie Duell, Xinran Xu, Thomas C. Henderson, Yu Yue Hew, Pik Wan Erica Chiang, Xiao Wei Alstar Ang, Bingwen Eugene Fan, Xiuyi Fan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28052v1 Announce Type: new Abstract: Artificial intelligence has strong potential to support clinical decision-making, yet its adoption in healthcare remains limited due to a lack of trust. Uncertainty estimation can signal unreliable predictions, and explainable AI (XAI) can clarify how ...

📖 Read original article


37. Comparing Classical and Quantum Machine Learning for Regression in High Energy Physics Collision Data ​

Author: Tariq Mahmood, Zain ul Abidin, Itzel Luviano Soto, Alfredo Raya
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, hep-ex, hep-ph, hep-th, quant-ph

arXiv:2608.28084v1 Announce Type: new Abstract: The classification and regression of particle collision events constitute a persistent computational challenge in experimental high energy physics, where large volumes of simulated data must be processed with both speed and precision. This work carries...

📖 Read original article


38. Generalized Gibbs Ensemble Weighting for Forecast Combination ​

Author: Prasen R. Nuthanakaluva, Nava K. Gaddam
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2608.28116v1 Announce Type: new Abstract: Forecast combination is a reliable way to improve predictive performance when several forecasting models are available. Simple aggregation rules such as the mean, median, trimmed mean, inverse-loss weighting, and exponential weighting are often strong ...

📖 Read original article


39. VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning ​

Author: Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28128v1 Announce Type: new Abstract: Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable terminal rewards by broadcasting each sparse outcome to every action in a traj...

📖 Read original article


40. Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting ​

Author: Morad Laglil, Younes Hlal, Marouane El Hadari, Emilie Devijver, Eric Gaussier
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28134v1 Announce Type: new Abstract: Reliable long-horizon time series forecasting is an important yet difficult problem. Trends and seasonality introduce complex temporal structure that challenges learning-based forecasting models. Differencing, which subtracts nearby past values to remo...

📖 Read original article


41. Conditional Diffusion Models for Energy-Efficient Driving ​

Author: Hemanth Neelgund Ramesh, Andr'e Snoeck, Chyi-Fu Hong, Shijing Sun
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28142v1 Announce Type: new Abstract: Electrification of commercial delivery fleets is shifting fleet routing from distance- and time-based optimization toward energy-aware decision-making. Existing sequence models primarily provide deterministic point estimates or limited uncertainty summ...

📖 Read original article


42. The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension ​

Author: Yuhe Sui, Jianing Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28150v1 Announce Type: new Abstract: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row-$\ell_1$ approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support ...

📖 Read original article


43. HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees ​

Author: Boyuan Meng (Ant Group, China), Peihua Bao (Ant Group, China), Hong Liu (Ant Group, China), Xiaowei Zhu (Ant Group, China), Chao Wang (Ant Group, China), Gen Li (Ant Group, China), Zhenxuan Pan (Ant Group, China)
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.DC

arXiv:2608.28158v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes. Existing systems primarily target full-attention models and lack dense,...

📖 Read original article


44. Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons ​

Author: Florin Leon
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28184v1 Announce Type: new Abstract: Grokking is a delayed transition from memorization to generalization that is often accompanied by substantial reorganization of internal representations. This paper studies whether biologically inspired mechanisms, many of which are not commonly incorp...

📖 Read original article


45. Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits ​

Author: Jingyi Zhou, Zhengyuan Shi, Jiaying Zhu, Ziyang Zheng, Qiang Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2608.28188v1 Announce Type: new Abstract: Circuit Representation Learning (CRL) offers a powerful paradigm to guide and optimize core Electronic Design Automation (EDA) tasks, but its practical adoption is hindered by the immense scale of industrial netlists and a failure to explicitly model r...

📖 Read original article


46. Performative Privacy: When Differential Privacy Maximizes Utility ​

Author: Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.28198v1 Announce Type: new Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning...

📖 Read original article


47. Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data ​

Author: Kazi F. Akhter, Ibna Kowsar, Manar D. Samad
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28209v1 Announce Type: new Abstract: Unlike images and text, applying transfer learning to tabular data is challenging due to heterogeneity in feature types, structures, and semantics across disparate domains. Existing methods assume shared features across data tables to enable knowledge ...

📖 Read original article


48. D-TAIA: Domain-Aware LLM Adaptation for Multi-Task Predictive Process Monitoring ​

Author: Sjoerd van Straten, Christine Jacob, Marwan Hassani
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28236v1 Announce Type: new Abstract: Predictive Process Monitoring (PPM) enables organizations to forecast future process behavior, such as the next activity and remaining time of ongoing cases. In practice, three conditions cause existing methods to degrade, namely data scarcity, high pr...

📖 Read original article


49. Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring ​

Author: Sjoerd van Straten, Marwan Hassani
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28237v1 Announce Type: new Abstract: Predictive Process Monitoring (PPM) models are increasingly deployed in dynamic environments where concept drift causes the underlying process distribution to shift over time. While recent work has moved toward online continual learning, existing metho...

📖 Read original article


50. Spectral Features Dominate BCG Respiratory-Event Detection: A Large-Scale Patient-Independent Comparison of Feature Groups in Sleep Apnea Patients ​

Author: Israel Campero Jurado, Zoe Bousraou, Lara Benning, Sara Padilla Neira, Alexander Breuss, Robert Riener, Esther Irene Schwarz, Elisabeth Wilhelm
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28242v1 Announce Type: new Abstract: Unobtrusive ballistocardiographic (BCG) sensing is a promising modality for long-term sleep-apnea monitoring, yet it remains unclear which signal features are most discriminative for respiratory-event detection. We present a literature-guided, patient-...

📖 Read original article


51. SinkSLOT: Sinkhorn via Sparse Lifted Optimal Transport ​

Author: Ian Hsieh, Soumya Snigdha Kundu, Tom Vercauteren, Reuben Dorent
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28262v1 Announce Type: new Abstract: Entropic optimal transport (EOT) has been shown to offer a computationally tractable approximation to exact optimal transport. However, the standard Sinkhorn-Knopp algorithm has two main limitations. First, given discrete measures with $N$ points, each...

📖 Read original article


52. Residual-Guided Randomized Neural Networks ​

Author: Mushir Akhtar, M. Tanveer, Mohd. Arshad
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28267v1 Announce Type: new Abstract: Randomized neural networks enable fast and analytically tractable training by fixing the input to hidden layer parameters at random and learning the output weights in closed form; however, their performance critically depends on a single uninformed dra...

📖 Read original article


53. Learning to Transfer Across Modes: Towards Unified Urban Mobility Forecasting ​

Author: Yixuan Zhao, Man Luo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28273v1 Announce Type: new Abstract: Urban transportation systems consist of multiple mobility modes that coexist within the same city and exhibit complex interdependencies, leading to correlated demand dynamics across modes. However, forecasting demand jointly across different modes rema...

📖 Read original article


54. An algebraic proof of Colombo's difference-power determinant conjecture ​

Author: Kun Li, Li Tie, Peng Wang, Zihan Liu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, math.RA

arXiv:2608.28274v1 Announce Type: new Abstract: Let $n\ge2$ be even, let $\lambda=(\lambda_1,\ldots,\lambda_n)\in\mathbb{R}^n$ have pairwise distinct coordinates, and define the difference-power matrix [ A_d(\lambda) := \bigl[(\lambda_r-\lambda_s)^d\bigr]_{r,s=1}^n, \qquad d\in\mathbb{N}. ] In 192...

📖 Read original article


55. Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation ​

Author: Linze Wu, Xinrui Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28276v1 Announce Type: new Abstract: Structured generation underpins large language model (LLM) agents that produce JSON, SQL, and function calls, where a single wrong field can cause the downstream action to fail. Constrained decoding already tracks parser transitions to enforce formal v...

📖 Read original article


56. VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation ​

Author: Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua, Haoyu Cao, Linli Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.28306v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, standard OPSD treats the teacher di...

📖 Read original article


57. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss ​

Author: Niccol`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28308v1 Announce Type: new Abstract: We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling \textit{jointly optimal} learning rates and batch sizes, we investigate their \textit{marginal} evolut...

📖 Read original article


58. SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data ​

Author: Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28408v1 Announce Type: new Abstract: Tabular data, as a core data format in machine learning, often lacks the discriminative power needed for high-performance modeling due to insufficient feature informativeness. Automated Feature Engineering (AutoFE) overcomes this by automating feature ...

📖 Read original article


59. Euclidean Fourier Neural Operators ​

Author: Nathanael Bosch, Niklas Frederik Schmitz, Michael F. Herbst
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.mtrl-sci, physics.comp-ph

arXiv:2608.28425v1 Announce Type: new Abstract: Fourier neural operators (FNOs) provide an efficient framework for learning mappings between function spaces as they are, by construction, independent of the grid resolution at which they are trained and evaluated. However, FNOs are not independent of ...

📖 Read original article


60. Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining ​

Author: Shuchen Zhu, Yuxin Fang, Mingze Wang, Kun Yuan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28442v1 Announce Type: new Abstract: Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring severe challenges. Although modern adaptive optimizers such as AdamW and Mu...

📖 Read original article


61. How Proper Scoring Rules Shape LLM Forecasting ​

Author: Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop"a"a, Philip E. Tetlock
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28482v1 Announce Type: new Abstract: This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same th...

📖 Read original article


62. REPLICANT: Learning Policies for Evading and Hardening Malware Detectors ​

Author: Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks, Lorenzo Cavallaro, Fabio Pierazzi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CR

arXiv:2608.28499v1 Announce Type: new Abstract: To determine the real-world effectiveness of machine learning based malware detection, it is vital to evaluate its robustness against highly capable adversaries. However, state-of-the-art attacks do not effectively model realistic adversaries, as they ...

📖 Read original article


63. An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models ​

Author: Javier Aguilar Mart'in
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28541v1 Announce Type: new Abstract: A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mo...

📖 Read original article


64. DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging ​

Author: Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28547v1 Announce Type: new Abstract: Model merging combines multiple task-specific fine-tuned LLMs into a single multi-task model without additional training. However, merged models are known to suffer from representation bias: systematic drift between the merged model's hidden states and...

📖 Read original article


65. Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining ​

Author: Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.28552v1 Announce Type: new Abstract: As a precursor to high-dimensional biomedical data modeling, reliable feature selection can reduce computational expense, improve modeling performance, and yield simpler, more interpretable models. However, most filter-based feature selection methods s...

📖 Read original article


66. Blog: Survey of Optimizers ​

Author: Ruoran Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.28557v1 Announce Type: new Abstract: Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design space has expanded from coordinates to matrices and layers, from fixed training horizons to policies over time, and from mathematical ...

📖 Read original article


67. QGPINNs: A Physics-Informed Neural Network Framework for Nonlocal Differential Equations on Quantum Graphs ​

Author: Vaibhav Mehandiratta, Saket Ramchandra
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2608.28589v1 Announce Type: new Abstract: We propose QGPINNs, a physics-informed neural network framework developed in PyTorch for the numerical solution of nonlocal differential equations on quantum graphs. The framework is designed as a general computational implementation in which the solut...

📖 Read original article


68. Accelerating LLM Inference via Vector Index Based Output Embeddings ​

Author: Martin Loretz, Sepp Hochreiter
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.27460v1 Announce Type: cross Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection ...

📖 Read original article


69. SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction ​

Author: Nilay Yilmaz, Naga Sai Abhiram Kusumba, Stella Wenxing Liu, Yezhou Yang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.27461v1 Announce Type: cross Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capt...

📖 Read original article


70. Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields ​

Author: YuJie Huang, WenWu He, ZhuoEr Lin, Congcong Liu, Dong Liang, Zhuo-Xu Cui
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.27475v1 Announce Type: cross Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it. These tasks are coupled: changing field placement changes the differential law, while a sufficiently flex...

📖 Read original article


71. Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations ​

Author: D. K. C. Senevirathna, A. A. E. Nanayakkara, H. M. C. K. Kulathunga, J. K. D. P. Nadula, R. M. Mapatuna, Malithi Nawarathne, Jaliya L. Wijayaraja, P. D. Senanayake, Samitha Vidhanaarachchi, Kalpani Manathunga
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SD

arXiv:2608.27480v1 Announce Type: cross Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected. This study proposes an IoT-enabled acoustic monitoring fr...

📖 Read original article


72. Multiscale Community-Based Fingerprinting of Signed Functional Networks ​

Author: Sema Athamnah, Selin Aviyente
Published: 8/31/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG, eess.SP

arXiv:2608.27483v1 Announce Type: cross Abstract: Objective: Recent studies demonstrate that functional connectomes contain subject-specific signatures, or \textit{fingerprints}, that can identify individuals across repeated sessions and tasks. Existing methods mostly rely on edge-level features tha...

📖 Read original article


73. Optimal Transport for Network Comparison: A Review with Machine Learning Applications ​

Author: James Hyun, Fran\c{c}ois G. Meyer
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.SI

arXiv:2608.27500v1 Announce Type: cross Abstract: Network comparison using optimal transport is a growing area of research in network science. Unlike standard graph metrics, optimal transport computes both network dissimilarity and a transport plan that explains how one graph morphs into another. In...

📖 Read original article


74. How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution ​

Author: Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Sch"olkopf, Zhijing Jin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.27510v1 Announce Type: cross Abstract: Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which instead trains attribution graphs with respect to a linear probe dire...

📖 Read original article


75. Destroy Me: Automatic Artifact Generation for Histopathology Images ​

Author: Zuzanna Krawczyk-Borysiak, Adam Krawczyk, Mateusz Miller, Gabriela Kaczmarek, S{\l}awomir Paku{\l}o, Ma{\l}gorzata Sok'o{\l}, .Zaneta Swiderska-Chadaj
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG

arXiv:2608.27516v1 Announce Type: cross Abstract: Deep learning's diagnostic utility in pathology is constrained by model vulnerability to real-world data imperfections. While current strategies favor "perfect data" by filtering low-quality regions, which can lead to the loss of valuable diagnostic ...

📖 Read original article


76. Ab initio Modeling of MoS2/Oxide Device Interfaces with Machine Learned Electronic Structures ​

Author: Manasa Kaniselvan, Mauro Dossena, Denghui Lu, Alexander Maeder, Nicolas Vetsch, Alexandros Nikolaos Ziogas, Mathieu Luisier
Published: 8/31/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.LG

arXiv:2608.27533v1 Announce Type: cross Abstract: We introduce a new ab initio approach to simulate semiconductor devices that integrates scalable machine-learned (ML) electronic structure models with an advanced quantum transport (QT) solver. The developed framework enables 10,000X speedups over de...

📖 Read original article


77. Towards a mathematical theory of superposition ​

Author: Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.IT, cs.LG, math.CO, math.IT

arXiv:2608.27540v1 Announce Type: cross Abstract: We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector (x) of active features is encoded through an overcomplete dictionary (W), and feature ...

📖 Read original article


78. Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study ​

Author: Nathaniel Chen, Kouroche Bouchiat, Peter Steiner, Azarakhsh Jalalvand, SangKyeun Kim, Egemen Kolemen
Published: 8/31/2026, 4:00:00 AM
Categories: physics.plasm-ph, cs.LG

arXiv:2608.27578v1 Announce Type: cross Abstract: Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such ...

📖 Read original article


79. Physics-informed learning for the inverse problem in resonant ultrasound spectroscopy ​

Author: Alejandro Cubillos Mu~noz, Manuela Rivas, Julian Rincon
Published: 8/31/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.LG, physics.app-ph, physics.comp-ph

arXiv:2608.27590v1 Announce Type: cross Abstract: Inferring elastic constants from resonant ultrasound spectra is a nonlinear and typically overdetermined inverse problem based on finite spectral data. We formulate the Rayleigh-Ritz inverse problem as a constrained inverse-isospectral problem on the...

📖 Read original article


80. Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates ​

Author: Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
Published: 8/31/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG

arXiv:2608.27612v1 Announce Type: cross Abstract: In neuroevolution, indirect encoding generates neural network connectivity from a compact genome rather than specifying each connection. ES-HyperNEAT automatically discovers where to place hidden nodes by examining CPPN output patterns: it recursivel...

📖 Read original article


81. Quantum SEDONet: Spectrally-Embedded Quantum Deep Operator Networks for Partial Differential Equations ​

Author: Muhammad Abid, Arth Sojitra, Bipin Tiwari, Omer San
Published: 8/31/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2608.27626v1 Announce Type: cross Abstract: Quantum DeepONet accelerates neural-operator inference by evaluating an orthogonally parameterized network on a quantum computer, reproducing in ideal simulation the accuracy of its classical counterpart at asymptotically lower inference cost. Its tr...

📖 Read original article


82. Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge ​

Author: Md Monjurul Ahsan Prodhan, Md Nour Hossain
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.27633v1 Announce Type: cross Abstract: Pothole detection and its severity measurement is still an important challenges in urban infrastructure management, where late maintenance directly contributes to vehicle damage, road accidents, and escalating repair costs. Existing automated approac...

📖 Read original article


83. CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT ​

Author: Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour, Michael Galarnyk, Abanoub Abdelmalak, Marly van Assen, Carlo N. De Cecco, Arshed Quyyumi, Ali Adibi
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG, cs.RO

arXiv:2608.27690v1 Announce Type: cross Abstract: Cardiovascular risk prediction remains limited by incomplete clinical data and imaging biomarkers that reduce computed tomography (CT) to a small number of handcrafted features. We developed CARDINAL (Cardiovascular Assessment via Representation lear...

📖 Read original article


84. On the Computational and Statistical Efficiency of the Empirical Maximum Entropy on the Mean Method ​

Author: Matthew King-Roskamp, Gabriel Rioux, Rustum Choksi, Tim Hoheisel
Published: 8/31/2026, 4:00:00 AM
Categories: math.OC, cs.LG, stat.ML

arXiv:2608.27705v1 Announce Type: cross Abstract: The Maximum Entropy on the Mean (MEM) method provides a flexible computational framework for solving inverse problems by combining data fidelity with entropy-based regularization. In practice, however, the prior distribution is typically unknown but ...

📖 Read original article


85. Beyond Procrustes distances: a multilinear Gromov-Wasserstein distance capturing chirality ​

Author: Cl'ement Soubrier, Geoffrey Woollard, Andrew Warren, Khanh Dao Duc
Published: 8/31/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.27774v1 Announce Type: cross Abstract: Efficiently and robustly analyzing shape data is critical across many scientific disciplines. While chirality is a fundamental property in numerous applications - most notably in molecular science - existing shape analysis metrics fail to distinguish...

📖 Read original article


86. Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots ​

Author: Xujun Che, Depeng Xu, Shuhan Yuan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.CL, cs.LG

arXiv:2608.27782v1 Announce Type: cross Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is treated as a proxy against all of them at once. We pin down the exact DP constant for the two that car...

📖 Read original article


87. CURA: Certified Runtime Alarms for Computer-Use Agents ​

Author: Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.27808v1 Announce Type: cross Abstract: Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. On 361 OSWorld tasks our pipeline, a read-only feasibility gate, a planner, and a GUI executor, reache...

📖 Read original article


88. Personalized and Multi-View Representation for Federated Cold-Start Recommendation ​

Author: Jaehyung Lim, Wonbin Kweon, Woojoo Kim, Junyoung Kim, Dongha Kim, Hwanjo Yu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.LG

arXiv:2608.27826v1 Announce Type: cross Abstract: Federated recommendation (FedRec) enables personalized modeling without centralizing users' interaction histories, but most existing methods assume a fixed item pool and thus overlook the practical cold-item setting where new items continuously arriv...

📖 Read original article


89. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests ​

Author: Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.27831v1 Announce Type: cross Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To c...

📖 Read original article


90. Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design ​

Author: Chuhan Yang, Chenxi Wang, Linhan Wu, Yuyang Liu
Published: 8/31/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.27873v1 Announce Type: cross Abstract: Early discovery of at least one valid design satisfying a target requirement is a central objective in failure-prone closed-loop inverse design. A natural batch baseline ranks candidates by a product-form marginal valid-hit score, but selecting the h...

📖 Read original article


91. What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection ​

Author: Parishruthi Ganesh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.27879v1 Announce Type: cross Abstract: Articulated human pose provides detailed body-configuration information beyond coarse spatial relationships, but whether this detail yields greater discriminative information when the downstream pipeline is held fixed remains unclear. We examine this...

📖 Read original article


92. OpenStamp: A Watermark for Open-Source Language Models ​

Author: Miroojin Bakshi, Saksham Rastogi, Danish Pruthi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.27899v1 Announce Type: cross Abstract: With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle ...

📖 Read original article


93. Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification ​

Author: Cameron Wilding, Mina Shaker, Fatemeh Ganji
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.27954v1 Announce Type: cross Abstract: Post-deployment changes to large language models can alter behavior while leaving routine outputs largely unchanged, creating a challenge for AI governance when model weights are proprietary. We present a privacy-preserving zk-SNARK-based audit frame...

📖 Read original article


94. Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning ​

Author: Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.28018v1 Announce Type: cross Abstract: Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention...

📖 Read original article


95. Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms ​

Author: Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton, John Paul Shen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.PF, cs.DC, cs.LG

arXiv:2608.28044v1 Announce Type: cross Abstract: Large language model (LLM) inference serving is priced by tokens, but GPU energy is consumed over inference windows. This accounting mismatch makes token-normalized metrics incomplete, since average output-token energy can decrease even when total re...

📖 Read original article


96. Emergent aggregation from collective foraging ​

Author: Gorka Mu~noz-Gil, Andrea L'opez-Incera, Vide Ramsten, Giovanni Volpe, Thomas M"uller, Hans J. Briegel
Published: 8/31/2026, 4:00:00 AM
Categories: cond-mat.stat-mech, cs.LG, cs.MA, nlin.AO, physics.bio-ph

arXiv:2608.28046v1 Announce Type: cross Abstract: Collective behaviour in living systems is usually modelled as the outcome of a \emph{direct} social drive: agents are rewarded, or hard-wired, to align with or approach their neighbours. Here we show that aggregation can instead emerge from an \emph{...

📖 Read original article


97. Landau theory of quenched criticality in linear in-context learning ​

Author: Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han
Published: 8/31/2026, 4:00:00 AM
Categories: cond-mat.dis-nn, cs.LG

arXiv:2608.28059v1 Announce Type: cross Abstract: In-context learning (ICL) allows a pretrained model to infer a new task from examples supplied in its prompt without updating its parameters. In linear models of ICL, the prediction error develops a double-descent singularity when the number of pretr...

📖 Read original article


98. Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations ​

Author: Naren Akash, Neeraja Ramanan
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG

arXiv:2608.28092v1 Announce Type: cross Abstract: Interpreting a CT scan means comparing structures on either side, judging how far apart organs sit, and knowing where each one belongs. Medical vision encoders are evaluated on diagnostic accuracy, or through assembled multimodal systems where a fail...

📖 Read original article


99. CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs ​

Author: Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG

arXiv:2608.28137v1 Announce Type: cross Abstract: We present CheXtriev, a graph-based, anatomy-aware framework for chest radiograph retrieval. Unlike prior methods focussed on global features, our method leverages graph transformers to extract informative features from specific anatomical regions. F...

📖 Read original article


100. Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards ​

Author: Zhen Liu, Marta Bono, Robbe Decloedt, Ajda Flisar, Maarten Van Den Bossche, Maarten De Vos
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.28152v1 Announce Type: cross Abstract: Agitation fluctuates over short time horizons in people living with dementia, yet continuous physiological information for anticipating next-day risk is limited. We assessed whether contactless under-mattress signals from the preceding night inform n...

📖 Read original article


101. Empowering Local Agriculture: A Deep Learning-Powered Web System for Identifying Bangladeshi Mango Varieties ​

Author: Monowar Islam, Safaruzzaman Shovo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.28161v1 Announce Type: cross Abstract: Mango variety identification in Bangladesh is challenging because closely related cultivars can have similar visual characteristics and images are often captured under varying real-world conditions. This work presents a deep learning-based web system...

📖 Read original article


102. Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control ​

Author: Amirmohammad Farzaneh, Osvaldo Simeone
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.IT, cs.LG, math.IT

arXiv:2608.28179v1 Announce Type: cross Abstract: We study risk-averse decision making, in which an agent selects actions while being uncertain about the true system state. The risk is measured via optimized certainty equivalent (OCE) metrics, which generalize popular criteria such as mean-variance ...

📖 Read original article


103. EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders ​

Author: Anja Witte, Maximilian Lennartz, Jan Baumbach, Guido Sauter, Stefan Bonn, Patrick Fuhlert, Marina Zimmermann
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.28191v1 Announce Type: cross Abstract: Vision Foundation Models (VFMs) are widely used in computational pathology but remain sensitive to domain shifts arising from variations in staining, tissue preparation, and scanner hardware. A key limitation is that VFM embeddings entangle biologica...

📖 Read original article


104. Explainable Diabetic Retinopathy Classification Using Vision Foundation Models ​

Author: Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.28207v1 Announce Type: cross Abstract: Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple tran...

📖 Read original article


105. Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance ​

Author: Vincenzo Collura, Karim Tit, Eleonora Giunchiglia, Mike Papadakis, Maxime Cordy
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.FL, cs.LG

arXiv:2608.28229v1 Announce Type: cross Abstract: Grammar-constrained decoding helps large language models produce syntactically valid structured outputs, such as code, JSON, and SQL. For context-free grammars, many practical decoders enforce local prefix feasibility: each token must keep the curren...

📖 Read original article


106. I-FLOP: Fast Learning of Order and Parents from Interventional Data ​

Author: Liuting Chen, Alex Markham
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.28245v1 Announce Type: cross Abstract: We extend the FLOP (fast learning of order and parents) algorithm recently proposed by Wien"obst et al. (2026) from observational to interventional data. In particular, we use the interventional BIC score of Hauser and B"uhlmann (2012), adapting it...

📖 Read original article


107. BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla ​

Author: Rowzatul Zannat, Abdullah Al Shafi, K. M. Azharul Hasan, Atia Shahnaz Ipa
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.28329v1 Announce Type: cross Abstract: Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To ad...

📖 Read original article


108. GRACE:Gradient-guided Coreset Selection for LLM Unlearning ​

Author: Praveen Bushipaka, Andrea D'Angelo, Lucia Passaro, Tommaso Cucinotta
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.28361v1 Announce Type: cross Abstract: Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests may provide only a few examples of undesired behavior, requiring forget and retain sets to be inferre...

📖 Read original article


109. Real-Time Monitoring of MHD Liquid Metal Flows with Shallow Recurrent Decoders ​

Author: Claudio Scardino, Stefano Riva, Carolina Introini, Matteo Lo Verso, Eric Cervi, Antonio Cammi, Laura Savoldi
Published: 8/31/2026, 4:00:00 AM
Categories: physics.comp-ph, cs.LG, physics.flu-dyn

arXiv:2608.28366v1 Announce Type: cross Abstract: State estimation in magnetohydrodynamic flows is critical for real-time monitoring of liquid metal blankets in tokamak fusion reactors. Due to the multiphysics nature of these phenomena, high-fidelity simulations are computationally prohibitive for r...

📖 Read original article


110. Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection ​

Author: Tommaso dorigo
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, physics.data-an

arXiv:2608.28375v1 Announce Type: cross Abstract: Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each ...

📖 Read original article


111. Quantum Federated Learning Based on Bures--Uhlmann Geometry for Heterogeneous Noisy Clients ​

Author: Haruki Emori, Masaki Uchihara, Yuuki Tokunaga
Published: 8/31/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2608.28379v1 Announce Type: cross Abstract: Quantum federated learning enables collaborative model training across quantum devices without sharing raw data, and it faces the data and hardware heterogeneity inherent to noisy quantum devices. Utilizing the quantum geometric tensor is a natural r...

📖 Read original article


112. Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation ​

Author: Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.28393v1 Announce Type: cross Abstract: Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with s...

📖 Read original article


113. Post-Training VLMs for Video Mistake Detection ​

Author: Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.28406v1 Announce Type: cross Abstract: Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed-...

📖 Read original article


114. Sliding-window beats linear attention ​

Author: Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.28444v1 Announce Type: cross Abstract: Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which i...

📖 Read original article


115. Generalized Splines and Gaussian Processes ​

Author: Michael Unser
Published: 8/31/2026, 4:00:00 AM
Categories: math.ST, cs.LG, math.FA, stat.ML, stat.TH

arXiv:2608.28446v1 Announce Type: cross Abstract: For finite-dimensional linear inverse problems where the variables are Gaussian, it is well-known that the minimum-mean-square error estimator takes the form of a regularized least-squares data fit. In this chapter, we show that this equivalence exte...

📖 Read original article


116. Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents ​

Author: Nan Li
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.28458v1 Announce Type: cross Abstract: Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Cha...

📖 Read original article


117. Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy ​

Author: Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.28564v1 Announce Type: cross Abstract: We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $\alpha\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum a...

📖 Read original article


118. On two proofs of $d^2$ mixing of weighted Dikin walks ​

Author: Yuansi Chen, Yunbum Kook
Published: 8/31/2026, 4:00:00 AM
Categories: cs.DS, cs.LG, math.OC, math.PR, stat.CO

arXiv:2608.28566v1 Announce Type: cross Abstract: We study the mixing time of weighted Dikin walks for sampling from exponential distributions on polytopes and truncated positive-semidefinite (PSD) cones. Our first result gives a general total-variation mixing bound under strong self-concordance, $...

📖 Read original article


119. Learning a Size-Weight Frontier for Synthetic-Augmented Inference ​

Author: Chengpiao Huang, Kaizheng Wang
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG, stat.ML

arXiv:2608.28576v1 Announce Type: cross Abstract: Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference acro...

📖 Read original article


120. Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning ​

Author: Nan Wang, Mohit Yadav, Jonathan Wulff, Aidan Rosenbaum, Kezhou Chen, Yuvan Sharma, Xu Dong, Yiwei Tao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.28578v1 Announce Type: cross Abstract: Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit insi...

📖 Read original article


121. Trajectory balance: Improved credit assignment in GFlowNets ​

Author: Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio, Chen Sun, Yoshua Bengio
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2201.13259v4 Announce Type: replace Abstract: Generative flow networks (GFlowNets) are a method for learning a stochastic policy for generating compositional objects, such as graphs or strings, from a given unnormalized density by sequences of actions, where many possible action sequences may ...

📖 Read original article


122. Diffusion models as plug-and-play priors ​

Author: Alexandros Graikos, Esmeralda S. Whitammer, Nebojsa Jojic, Dimitris Samaras
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2206.09012v4 Announce Type: replace Abstract: We consider the problem of inferring high-dimensional data $\mathbf{x}$ in a model that consists of a prior $p(\mathbf{x})$ and an auxiliary differentiable constraint $c(\mathbf{x},\mathbf{y})$ on $x$ given some additional information $\mathbf{y}$....

📖 Read original article


123. Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey ​

Author: Juan Zhong, Yuhang Shi, Zukang Xu, Xi Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.RO, cs.SY, eess.SY

arXiv:2304.10891v4 Announce Type: replace Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi-agent interactions, and multimodal context across perception, prediction, and planning. At the same time,...

📖 Read original article


124. Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNets ​

Author: Dinghuai Zhang, Hanjun Dai, Esmeralda S. Whitammer, Aaron Courville, Yoshua Bengio, Ling Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DM, stat.ML

arXiv:2305.17010v4 Announce Type: replace Abstract: Combinatorial optimization (CO) problems are often NP-hard and thus out of reach for exact algorithms, making them a tempting domain to apply machine learning methods. The highly structured constraints in these problems can hinder either optimizati...

📖 Read original article


125. Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network ​

Author: Tristan Deleu, Mizu Nishikawa-Toomey, Jithendaraa Subramanian, Esmeralda S. Whitammer, Laurent Charlin, Yoshua Bengio
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2305.19366v3 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets), a class of generative models over discrete and structured sample spaces, have been previously applied to the problem of inferring the marginal posterior distribution over the directed acyclic graph (DAG) of a B...

📖 Read original article


126. Biases in Expected Goals Models Confound Finishing Ability ​

Author: Jesse Davis, Pieter Robberechts
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.AP

arXiv:2401.09940v2 Announce Type: replace Abstract: Expected Goals (xG) has emerged as a popular tool for evaluating finishing skill in soccer analytics. It involves comparing a player's cumulative xG with their actual goal output, where consistent overperformance indicates strong finishing ability....

📖 Read original article


127. Improved off-policy training of diffusion samplers ​

Author: Marcin Sendera, Minsu Kim, Sarthak Mittal, Pablo Lemos, Luca Scimeca, Jarrid Rector-Brooks, Alexandre Adam, Yoshua Bengio, Esmeralda S. Whitammer
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2402.05098v5 Announce Type: replace Abstract: We study the problem of training diffusion models to sample from a distribution with a given unnormalized density or energy function. We benchmark several diffusion-structured inference methods, including simulation-based variational approaches and...

📖 Read original article


128. Amortizing intractable inference in diffusion models for vision, language, and control ​

Author: Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Esmeralda S. Whitammer
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2405.20971v3 Announce Type: replace Abstract: Diffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable posterior inference problem. This paper studies amortized sampling ...

📖 Read original article


129. Meta-Prompt Optimization for LLM-Based Sequential Decision Making ​

Author: Mingze Kong, Zhiyong Wang, Yao Shu, Zhongxiang Dai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2502.00728v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been employed as agents to solve sequential decision-making tasks such as Bayesian optimization and multi-armed bandits (MAB). These works usually adopt an LLM for sequential action selection by providing ...

📖 Read original article


130. RegCL: Compact Continual SAM Adaptation for Visual Grounding in Multi-Sensorial Media ​

Author: Yuan-Chen Shu, Zhiwei Lin, Xiaoyu Zhou, Yongtao Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2507.12297v2 Announce Type: replace Abstract: Multi-sensorial media systems, including AR/VR, remote operation, and embodied AI, require visual grounding modules that remain reliable as sensing environments and application domains evolve. The Segment Anything Model (SAM) provides a strong foun...

📖 Read original article


131. Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers ​

Author: Mu Qiao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.NC

arXiv:2508.08289v3 Announce Type: replace Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major lin...

📖 Read original article


132. Class Incremental Continual Learning with Self-Organizing Maps and Synthetic Replay ​

Author: Pujan Thapa, Alexander Ororbia, Travis Desell
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2508.21240v2 Announce Type: replace Abstract: This work introduces a novel generative continual learning framework based on self-organizing maps (SOMs), a brain-inspired natural computing model, extended with learned distributional statistics and encoder--decoder models for class incremental c...

📖 Read original article


133. Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning ​

Author: Bastien Dubail, Stefan Stojanovic, Alexandre Prouti`ere
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2509.05193v3 Announce Type: replace Abstract: Low-rank structure is a common implicit assumption in many modern reinforcement learning (RL) algorithms. For instance, reward-free and goal-conditioned RL methods often presume that the successor measure admits a low-rank representation. In this w...

📖 Read original article


134. Large Reasoning Models Learn Better Alignment from Flawed Thinking ​

Author: ShengYun Peng, Pin-Yu Chen, Eric Smith, Song Jiang, Hongyuan Zhan, Haozhu Wang, Mahesh Pasupuleti, Duen Horng Chau, Jianfeng Chi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2510.00938v3 Announce Type: replace Abstract: Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flawed premise is inject...

📖 Read original article


135. One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms ​

Author: Xiang Li, You Li, Yazhou Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2511.08444v2 Announce Type: replace Abstract: EEG-based emotion recognition is hampered by profound dataset heterogeneity (channel/subject variability), hindering generalizable models. Existing approaches struggle to transfer knowledge effectively. We propose 'One Model for All', a universal p...

📖 Read original article


136. Aspiration-based Perturbed Learning Automata in Games with Noisy Utility Measurements. Part A: Stochastic Stability in Non-zero-Sum Games ​

Author: Georgios C. Chasparis
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.GT, cs.MA, math.OC

arXiv:2511.11602v3 Announce Type: replace Abstract: Reinforcement-based learning has attracted considerable attention both in modeling human behavior as well as in engineering, for designing measurement- or payoff-based optimization schemes. Such learning schemes exhibit several advantages, especial...

📖 Read original article


137. The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior ​

Author: Erik Larsen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2512.12066v3 Announce Type: replace Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption by investigating the stabil...

📖 Read original article


138. Bayesian Experimental Design for Model Discrepancy Calibration: A Rivalry between Kullback--Leibler Divergence and Wasserstein Distance ​

Author: Huchen Yang, Xinghao Dong, Jin-Long Wu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2601.16425v2 Announce Type: replace Abstract: Designing experiments that systematically gather data from complex physical systems is central to accelerating scientific discovery. While Bayesian experimental design (BED) provides a principled, information-based framework that integrates experim...

📖 Read original article


139. Simplex-to-Euclidean Bijection for Conjugate and Calibrated Multiclass Gaussian Process Classification ​

Author: Bernardo Williams, Harsha Vardhan Tetali, Arto Klami, Marcelo Hartmann
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2603.16621v2 Announce Type: replace Abstract: We propose a conjugate and calibrated Gaussian process (GP) model for multi-class classification by exploiting the geometry of the probability simplex. Our approach uses Aitchison geometry to map simplex-valued class probabilities to an unconstrain...

📖 Read original article


140. InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model ​

Author: Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.18031v2 Announce Type: replace Abstract: Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, wherea...

📖 Read original article


141. Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning ​

Author: Moritz G"ogl, Christopher Yau
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.20111v2 Announce Type: replace Abstract: The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emphasizing prediction in representation space rather than reconstruction in observation space. We argue ...

📖 Read original article


142. PolicyLong: Towards On-Policy Context Extension ​

Author: Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.07809v2 Announce Type: replace Abstract: Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic verification, selecting contexts that reduce a base model's predictive...

📖 Read original article


143. SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning ​

Author: Halil Ibrahim Gulluk, Olivier Gevaert
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2604.09887v2 Announce Type: replace Abstract: Medical vision-language datasets are often limited in size and biased toward negative findings, as clinicians report abnormalities mostly but might omit some positive/neutral findings because they might be considered as irrelevant to the patient's ...

📖 Read original article


144. Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making ​

Author: Abhirami Pillai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, econ.EM, stat.ML

arXiv:2604.26169v2 Announce Type: replace Abstract: Treatment allocation under budget constraints is a central challenge in digital advertising. The standard approach trains an offline uplift model on historical data, then solves a constrained optimization to allocate budget. This fails in cold-star...

📖 Read original article


145. ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space ​

Author: Gabe Guo, Thanawat Sornwanee, Lutong Hao, Elon Litman, Stefano Ermon, Jose Blanchet
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.27443v3 Announce Type: replace Abstract: Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), suffer ...

📖 Read original article


146. ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation ​

Author: Haoan Feng, Xin Xu, Leila De Floriani
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.22556v3 Announce Type: replace Abstract: Digital elevation models (DEMs) underpin terrain analysis in Geographic Information Systems (GIS), but commonly as raster representation, they rely on interpolation for off-grid sampling and finite-difference operators for derivative-based analysis...

📖 Read original article


147. Learned Relay Representations for Forward-Thinking Discrete Diffusion Models ​

Author: Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.22967v3 Announce Type: replace Abstract: When Masked Diffusion Models (MDMs) generate sequences through iterative refinement, the rich internal computation over masked positions is discarded, forcing every subsequent refinement step to recompute the valuable internal information stored as...

📖 Read original article


148. More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations ​

Author: Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2605.26647v2 Announce Type: replace Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLMs). Despite the evolution from ReLU and GELU to gated variants such as SwiGLU, most FFN designs st...

📖 Read original article


149. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models ​

Author: Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2605.26895v2 Announce Type: replace Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood ...

📖 Read original article


150. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis ​

Author: Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu, Lei Liang, Ningyu Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.MA

arXiv:2605.30434v2 Announce Type: replace Abstract: Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a benchmar...

📖 Read original article


151. Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks? ​

Author: Ojas Nimase, Jiate Li, Yue Zhao, Yushun Dong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.30470v2 Announce Type: replace Abstract: Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. However, this transparency creates exploitable vulnerabilities for model extraction attacks. We pr...

📖 Read original article


152. ERP-XTTN: Interpretable Prototype-Guided Cross-Attention for Cross-Subject ERP Classification ​

Author: Charlotte Genevier Wyman, Leanne Hirshfield
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, eess.SP

arXiv:2606.02939v2 Announce Type: replace Abstract: Interpretable brain-computer interface classifiers that generalize across subjects without calibration remain an open challenge. We evaluated whether prototype-based cross-attention can provide competitive, inherently interpretable ERP classificati...

📖 Read original article


153. The Discrete-Log Clock: How a Transformer Learns Modular Multiplication ​

Author: Huu Danh Nguyen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.17399v2 Announce Type: replace Abstract: When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies suffices. ...

📖 Read original article


154. SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity ​

Author: Liyang Yuan, Yibo Yang, Dandan Guo, Peter Richtarik, Zhouchen Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.04189v2 Announce Type: replace Abstract: Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this dri...

📖 Read original article


155. On the Depth Scalability of Logic Gate Networks ​

Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO

arXiv:2607.21633v3 Announce Type: replace Abstract: Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased depth. We identify two causes: optimization collapse and topology-induced degradation of output-specific credit ...

📖 Read original article


156. Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction ​

Author: Mehrdad Shoeibi, Niloofar Yousefi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.00152v2 Announce Type: replace Abstract: Predicting how held-out target genes respond to CRISPRi perturbation, and whether such predictions transfer across biological screens, is hard to evaluate: a representation can be informative within one screen yet fail across screens, while endpoin...

📖 Read original article


157. ED-CSP: Crystal Structure Prediction from Electron Diffraction ​

Author: Germain Poloudenny, Arnaud Demorti`ere, Ya"el Fr'egier
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06448v3 Announce Type: replace Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures ...

📖 Read original article


158. Recirculation ​

Author: Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal, Rosanne Liu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.17981v2 Announce Type: replace Abstract: We describe an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency during gen...

📖 Read original article


159. How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness ​

Author: Yi Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD

arXiv:2608.18025v2 Announce Type: replace Abstract: GPT-style models have achieved remarkable success with finite vocabularies of reusable tokens, making the token interface a central component of modern sequence modeling. Symbolic music appears naturally compatible with this paradigm: it consists o...

📖 Read original article


160. How Architecture and Training Affect TPC Representations Across Experiments ​

Author: Tyler Wheeler, Michelle P. Kuchera, Raghuram Ramanujan, William Sieland, Ryan Krupp, Yassid Ayyad, Daniel Bazin, Connor L. Cross, Hoi Yan Ian Heung, Andrew J. Jones, Ruchi Mahajan, Saiprasad Ravishankar, Pranjal Singh, Benjamin Votaw, Chris Wrede
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, nucl-ex, physics.ins-det

arXiv:2608.21756v2 Announce Type: replace Abstract: Deep-learning efforts have increasingly shifted toward foundation model approaches. In experimental physics, this allows models and learned representations to be reused beyond the experiments in which they were developed. This work evaluates the re...

📖 Read original article


161. RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling ​

Author: Ziyuan Wang, Bohao Tang, Fei Zhang, Shuo Han, Pengfei Liu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, q-bio.GN

arXiv:2608.22849v2 Announce Type: replace Abstract: Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bi...

📖 Read original article


162. JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics ​

Author: Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.24044v3 Announce Type: replace Abstract: Latent world models plan by predicting how candidate actions advance learned latent dynamics. In self-predictive models, however, the encoder and predictor are optimized jointly and can co-adapt to latent transitions that are easy to predict but we...

📖 Read original article


163. On-policy Distillation with Verifiable Reward ​

Author: Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24696v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense...

📖 Read original article


164. Trust the Mass: Forced Weights in KV-Cache Eviction ​

Author: Jack Shi, Jerry Gu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.25230v2 Announce Type: replace Abstract: Every deployed sparse-attention or KV-cache-eviction rule keeps a subset of the keys, discards the rest, and renormalizes the attention weights over the kept set. Enumerating the exact best subset under that constraint on $168{,}192$ attention rows...

📖 Read original article


165. TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development ​

Author: Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.26086v2 Announce Type: replace Abstract: Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions stil...

📖 Read original article


166. Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs ​

Author: Rupesh Sapkota, Louis Mozart Kamdem Teyou, Moshood Yekini, Caglar Demir, Axel-Cyrille Ngonga Ngomo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.26729v2 Announce Type: replace Abstract: In recent years, transductive knowledge graph embedding models have been applied to tasks such as link prediction and query answering. Although knowledge graphs often contain rich numerical attributes, most embedding models neglect them, limiting t...

📖 Read original article


167. Accurate prediction is not profitable advice: profit-based evaluation of machine learning nitrogen recommendations in winter wheat ​

Author: Xulong Wang, Po Yang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27205v2 Announce Type: replace Abstract: Nitrogen rates for winter wheat are set before the season, under unknown prices and weather. The standard UK advice does not respond to prices, yet recent price swings moved the most profitable rate by tens of kilograms per hectare. Machine learnin...

📖 Read original article


168. Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO ​

Author: Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.27351v2 Announce Type: replace Abstract: Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-...

📖 Read original article


169. Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity ​

Author: Ismail Rasim Ulgen, John H. L. Hansen, Carlos Busso, Berrak Sisman
Published: 8/31/2026, 4:00:00 AM
Categories: eess.AS, cs.LG

arXiv:2407.04291v4 Announce Type: replace-cross Abstract: Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech systems, but they are typically trained for speaker recognition, where intra-speaker variability is supp...

📖 Read original article


170. Mixture of Multicenter Experts in Multimodal AI for Debiased Radiotherapy Target Delineation ​

Author: Yujin Oh, Sangjoon Park, Xiang Li, Pengfei Jin, Yi Wang, Jonathan Paly, Jason Efstathiou, Annie Chan, Jun Won Kim, Hwa Kyung Byun, Ik Jae Lee, Jaeho Cho, Chan Woo Wee, Peng Shu, Peilong Wang, Caiwen Jiang, Nathan Yu, Jason Holmes, Jong Chul Ye, Quanzheng Li, Wei Liu, Woong Sub Koom, Jin Sung Kim, Kyungsang Kim
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.CV, cs.LG

arXiv:2410.00046v4 Announce Type: replace-cross Abstract: Clinical decision-making reflects diverse strategies shaped by regional patient populations and institutional protocols. However, most existing medical artificial intelligence (AI) models are trained on highly prevalent data patterns, which r...

📖 Read original article


171. Off the Normal Path: Learning Spatial Density Models of Node Mobility ​

Author: Wanxin Gao, Ioanis Nikolaidis, Janelle Harms
Published: 8/31/2026, 4:00:00 AM
Categories: cs.NI, cs.LG, stat.ML

arXiv:2411.10997v2 Announce Type: replace-cross Abstract: We consider the problem of learning models of spatial density functions, representing the steady-state density of mobile nodes moving on a two-dimensional terrain. Deriving such models can assist in network design and optimization problems, e...

📖 Read original article


172. Ampere: Communication-Efficient and High-Accuracy Split Federated Learning ​

Author: Zihan Zhang, Leon Wong, Blesson Varghese
Published: 8/31/2026, 4:00:00 AM
Categories: cs.DC, cs.LG

arXiv:2507.07130v2 Announce Type: replace-cross Abstract: A Federated Learning (FL) system collaboratively trains neural networks across devices and a server but is limited by significant on-device computation costs. Split Federated Learning (SFL) systems mitigate this by offloading a block of layer...

📖 Read original article


173. Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning ​

Author: Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas
Published: 8/31/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD

arXiv:2509.00094v2 Announce Type: replace-cross Abstract: Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is enabled by the rigorous recitation rules (Tajweed) established through the e...

📖 Read original article


174. Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests ​

Author: Somjit Roy, Pritam Dey, Bani K. Mallick, Debdeep Pati
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, cs.SC, math.ST, stat.ML, stat.TH

arXiv:2509.19710v3 Announce Type: replace-cross Abstract: Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expressions that reveal governing relationships directly from data. Existing methods, however, oft...

📖 Read original article


175. Examining the robustness of Physics-Informed Neural Networks to noise for Inverse Problems ​

Author: Aleksandra Jekic, Afroditi Natsaridou, Signe Riemer-S{\o}rensen, Helge Langseth, Odd Erik Gundersen
Published: 8/31/2026, 4:00:00 AM
Categories: physics.comp-ph, cs.LG, cs.NA, math.NA

arXiv:2509.20191v2 Announce Type: replace-cross Abstract: Approximating solutions to partial differential equations (PDEs) is fundamental for the modeling of dynamical systems in science and engineering. Physics-informed neural networks (PINNs) are a recent machine learning-based approach, for which...

📖 Read original article


176. OceanGym: A Benchmark Environment for Underwater Embodied Agents ​

Author: Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.RO

arXiv:2509.26536v3 Announce Type: replace-cross Abstract: We introduce OceanGym, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments. Unlike terrestrial or aerial domains, underwater settings present extr...

📖 Read original article


177. GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis ​

Author: Subrat Kishore Dutta, Yuelin Xu, Piyush Pant, Xiao Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2510.09260v3 Announce Type: replace-cross Abstract: Recent work has shown that RLHF is highly susceptible to backdoor attacks. However, existing methods often rely on rare tokens or fixed triggers, limiting their impact in realistic scenarios. In this work, we develop GREAT, a novel framework ...

📖 Read original article


178. Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity ​

Author: Anastasiia Tokareva, Judith Dineley, Zoe Firth, Pauline Conde, Faith Matcham, Sara Siddi, Femke Lamers, Ewan Carr, Carolin Oetzmann, Daniel Leightley, Yuezhou Zhang, Amos A. Folarin, Josep Maria Haro, Brenda W. J. H. Penninx, Raquel Bailon, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins, The RADAR-CNS Consortium
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2511.07011v2 Announce Type: replace-cross Abstract: Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (...

📖 Read original article


179. Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning ​

Author: Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.PF

arXiv:2511.08577v4 Announce Type: replace-cross Abstract: Improving the reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications. Looped transformers address this by performing multiple latent iterations to refine each token ...

📖 Read original article


180. Prequential posteriors ​

Author: Shreya Sinha-Roy, Richard G. Everitt, Christian P. Robert, Ritabrata Dutta
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2511.17721v2 Announce Type: replace-cross Abstract: Data assimilation is a fundamental task in updating forecasting models upon observing new data, with applications ranging from weather prediction to online reinforcement learning. Deep generative forecasting models (DGFMs) have shown excellen...

📖 Read original article


181. Aligning Agentic World Models via Knowledgeable Experience Learning ​

Author: Baochang Ren, Yunzhi Yao, Rui Sun, Shuofei Qiao, Ningyu Zhang, Huajun Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.MM

arXiv:2601.13247v2 Announce Type: replace-cross Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural grounding to respect the immutable laws of the physical world. Consequently, while these agents implicitly ...

📖 Read original article


182. Learning Fast Monomial Orders for Gr\"obner Basis Computations ​

Author: R. Caleb Bunch, Alperen A. Erg"ur, Melika Golestani, Jessie Tong, Malia Walewski, Yunus E. Zeytuncu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SC, cs.LG, math.AC, math.AG

arXiv:2602.02972v2 Announce Type: replace-cross Abstract: The efficiency of Gr"obner basis computation, the standard engine for solving systems of polynomial equations, depends on the choice of monomial ordering. Despite a near-continuum of possible monomial orders, most implementations rely on sta...

📖 Read original article


183. SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models ​

Author: Hyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang, Seongwon Cho, Jonghyun Choi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2602.04208v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require...

📖 Read original article


184. Robust Assortment Optimization from Observational Data ​

Author: Miao Lu, Yuxuan Han, Han Zhong, Zhengyuan Zhou, Jose Blanchet
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.OC, math.ST, stat.TH

arXiv:2602.10696v3 Announce Type: replace-cross Abstract: Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue under complex customer choice behaviors. While recent advances in...

📖 Read original article


185. Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval ​

Author: Jiaqi Xi, Raghav Saboo, Luming Chen, Johny Rufus, Aditya Dodda, Ved Sampath, Kenny Chi, Elyse Winer, Akshad Viswanathan, Martin Wang, Sudeep Das
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.LG

arXiv:2602.17654v2 Announce Type: replace-cross Abstract: Embedding-based retrieval (EBR) for large-scale e-commerce search faces three intertwined challenges: graded (non-binary) relevance where engagement signals are noisy and intent-varying while business relevance guidelines admit acceptable-but...

📖 Read original article


186. FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation ​

Author: Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2602.22056v3 Announce Type: replace-cross Abstract: Generative manipulation policies can fail catastrophically under deployment-time distribution shift, yet many failures are near-misses: the robot reaches almost-correct poses and would succeed with a small corrective motion. We propose FlowCo...

📖 Read original article


187. Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling ​

Author: Hamed Hamzeh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.DC, cs.LG, cs.MA

arXiv:2603.12031v4 Announce Type: replace-cross Abstract: Cloud-native container orchestration requires resource schedulers capable of balancing infrastructure expenditure, fault resilience, and node utilisation. Conventional reinforcement learning approaches typically rely on monolithic single-agen...

📖 Read original article


188. The Autonomy Tax: Defense Training Breaks LLM Agents ​

Author: Shawn Li, Yue Zhao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2603.19423v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks. Practitioners deploy defense-trained models to protect against prompt...

📖 Read original article


189. Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence ​

Author: Peter Fasogbon, Ugurcan Budak, Patrice Rondao Alface, Hamed Rezazadegan Tavakoli
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2603.21933v3 Announce Type: replace-cross Abstract: The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream processing. However, most of the existing pruning strategies depend on camera parameters, rendered images, ...

📖 Read original article


190. Deflation-PINNs: Learning Multiple Solutions for PDEs and Landau-de Gennes ​

Author: Sean Disar`o, Ruma Rani Maity, Aras Bacho
Published: 8/31/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2603.27936v3 Announce Type: replace-cross Abstract: Nonlinear Partial Differential Equations (PDEs) are ubiquitous in mathematical physics and engineering. Although Physics-Informed Neural Networks (PINNs) have emerged as a powerful tool for solving PDE problems, they typically struggle to ide...

📖 Read original article


191. Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions ​

Author: Doan Nam Long Vu, Simone Balloccu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2603.28387v3 Announce Type: replace-cross Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and co...

📖 Read original article


192. Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks ​

Author: Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2604.12379v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and current benchmarks fo...

📖 Read original article


193. G-Loss: Graph-Guided Fine-Tuning of Language Models ​

Author: Aditya Sharma, Vinti Agarwal, Rajesh Kumar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.25853v4 Announce Type: replace-cross Abstract: Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language models such as BERT, operate only within local neighborhoods and fail to account for the ...

📖 Read original article


194. DiffAnon: Diffusion-based Prosody Control for Voice Anonymization ​

Author: Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn, Berrak Sisman
Published: 8/31/2026, 4:00:00 AM
Categories: eess.AS, cs.LG, cs.SD

arXiv:2604.26281v2 Announce Type: replace-cross Abstract: To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled me...

📖 Read original article


195. D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery ​

Author: Hanane Nour Moussa, Yifei Li, Zhuoyang Li, Yankai Yang, Cheng Tang, Tianshu Zhang, Nesreen K. Ahmed, Ali Payani, Ziru Chen, Huan Sun
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.27977v3 Announce Type: replace-cross Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we int...

📖 Read original article


196. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces ​

Author: Chang Jin, An Wang, Zeming Wei, Kai Wang, Biaojie Zeng, Qiaosheng Zhang, Chao Yang, Jingjing Qu, Xia Hu, Xingcheng Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2605.12015v3 Announce Type: replace-cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that...

📖 Read original article


197. Online Learning-to-Defer with Varying Experts ​

Author: Dang Hoang Duy, Yannis Montreuil, Maxime Meyer, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2605.12340v5 Announce Type: replace-cross Abstract: Learning-to-Defer (L2D) methods route each query either to a predictive model or to external experts. Real-world deployments require handling streaming data, changing expert availability, shifting expert reliability, and feedback observed onl...

📖 Read original article


198. WINO: A Weak-Form Physics Informed Neural Operator for Hyperelasticity on Variable Domains ​

Author: Bokai Zhu, Yizheng Wang, Qinghui Zhang, Timon Rabczuk
Published: 8/31/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2605.24651v3 Announce Type: replace-cross Abstract: We propose a Weak-form Physics-Informed Neural Operator (WINO), a data-free framework that combines the efficiency of neural operators with the geometric flexibility of the $\varphi$-finite element method ($\varphi$-FEM). $\varphi$-FEM is an ...

📖 Read original article


199. ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs ​

Author: Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber, Sahil Bansal
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.LG

arXiv:2606.12451v2 Announce Type: replace-cross Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, parametric tool r...

📖 Read original article


200. RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation ​

Author: Hashir Ahmed, Zahra Khotanlou, Chenghao Tan, Ahmed Abdelaal, Amir-Hossein Karimi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.16113v2 Announce Type: replace-cross Abstract: Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing...

📖 Read original article


201. TokenPilot: Cache-Efficient Context Management for LLM Agents ​

Author: Buqiang Xu, Zirui Xue, Dianmou Chen, Chenyang Fu, Chiyu Wu, Caiying Huang, Chen Jiang, Jizhan Fang, Xinle Deng, Yijun Chen, Yunzhi Yao, Xuehai Wang, Jin Shang, Gong Yu, Ningyu Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA

arXiv:2606.17016v2 Announce Type: replace-cross Abstract: As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; however, their unconstrained sequence mutat...

📖 Read original article


202. Closing the Operational Gap in Semantic Caching ​

Author: Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.CL, cs.LG

arXiv:2606.19719v3 Announce Type: replace-cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usab...

📖 Read original article


203. An End-to-End Hybrid Quantum--Classical Sampling Workflow for Discrete Markov Random Fields: A Reproducible Case Study ​

Author: Arul Rhik Mazumder
Published: 8/31/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2607.09893v2 Announce Type: replace-cross Abstract: Sampling from discrete Markov random fields (MRFs) is a hard problem. We study amplitude-encoded i.i.d. sampling for small MRFs where $2^n$ target probabilities are precomputed classically. This removes quantum exponential speedup but allows ...

📖 Read original article


204. Robust Chance-Constrained Optimization using a Continuous Parameter Space Wasserstein-2 Ambiguity Set of Gaussian Mixtures ​

Author: Shibshankar Dey, Sanjay Mehrotra
Published: 8/31/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2607.17018v2 Announce Type: replace-cross Abstract: We study distributionally robust linear chance-constrained problems in which uncertainty is modeled by a Gaussian mixture model (GMM). Finite-support distributionally robust (FDR) formulations, widely used in data-driven robust optimization, ...

📖 Read original article


205. Where Steering Signals Come From: Activation Source Selection in Activation Steering ​

Author: Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.25270v2 Announce Type: replace-cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation...

📖 Read original article


206. Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization ​

Author: Kuangyu Ding, Kim-Chuan Toh
Published: 8/31/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.07248v3 Announce Type: replace-cross Abstract: Sequence convergence to a boundary Karush--Kuhn--Tucker (KKT) point has long remained unclear for nonconvex mirror descent with Legendre kernels. The difficulty arises from the blow-up of the gradient of the Legendre kernel at the boundary. R...

📖 Read original article


207. JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification ​

Author: Tianxin Zhou, Ruixi Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.20607v2 Announce Type: replace-cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind...

📖 Read original article


208. What Neural Network Field Theory Can and Cannot Realise on a Computer ​

Author: Thomas R. Harvey
Published: 8/31/2026, 4:00:00 AM
Categories: hep-th, cs.LG, hep-lat, hep-ph, math-ph, math.MP

arXiv:2608.21523v2 Announce Type: replace-cross Abstract: One aim of neural network field theory is to put a quantum or effective field theory on a computer, with the network ensemble itself as the theory. We ask how far that aim can be pushed for a function class regular enough to be computed with....

📖 Read original article


209. GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets ​

Author: Saif Ahmed, Asadullah Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Mohammad Abdul Qayum, Mohsin Sajjad, Mohammad Ashrafuzzaman Khan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.22272v2 Announce Type: replace-cross Abstract: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a...

📖 Read original article


210. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors ​

Author: Joshua Penman
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.LG

arXiv:2608.23873v2 Announce Type: replace-cross Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read lik...

📖 Read original article