Skip to content

arXiv cs.LG - 2026-08-14 ​

268 items collected.


1. LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining ​

Author: Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu, Zeng You, Ryan Dong, Simeng Ma, Yaofo Chen, Mingkui Tan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12419v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for localit...

📖 Read original article


2. Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods ​

Author: Matthew Kahn, Milan Arjel, Nirmala Adhikari, Mingmar Sherpa, James Pope
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.AP

arXiv:2608.12422v1 Announce Type: new Abstract: Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slowly sagging, and satellite weather marks the weeks when a primed lake is under stress. A companion feas...

📖 Read original article


3. MARCH: Scaling Recurrent Memory with Content-Routed State Anchors ​

Author: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12435v1 Announce Type: new Abstract: Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key--value cache that grows linea...

📖 Read original article


4. Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach ​

Author: Jiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Qian Zhu, Ying Liu, Zhenyu Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.MA, cs.NI

arXiv:2608.12436v1 Announce Type: new Abstract: Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although mu...

📖 Read original article


5. Unifying Generative Models with Path Integrals ​

Author: Ramon Winterhalder
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, hep-ph, stat.ML

arXiv:2608.12438v1 Announce Type: new Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form ...

📖 Read original article


6. Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection ​

Author: Iyad Assaad Nekka, Hamida Seba, Khaled Walid Hidouci, Karima Amrouche
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12441v1 Announce Type: new Abstract: Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information syste...

📖 Read original article


7. Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts ​

Author: Seyyed Ali Hoseini, Javad Baseri, Hamid Saadatfar, Edris Hoseini Gol, AmirHossein Eshghi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12446v1 Announce Type: new Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. This study investigates whether...

📖 Read original article


8. Geometric and Behavioral Stratification in Transformer Residual Streams ​

Author: Nelson Guda
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction o...

📖 Read original article


9. Exemplar-based objective classification of gust-induced loads across multiple flight conditions ​

Author: Paolo Olivucci, Kowshik Srivatsan, David E. Rival
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, physics.data-an

arXiv:2608.12448v1 Announce Type: new Abstract: Is it possible to find an objective classification criterion that organizes the complexity of gust-induced loads across many flight conditions? And one that remains as interpretable as a labelling based on coarse parameters, such as the flight attitude...

📖 Read original article


10. Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication ​

Author: Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12477v1 Announce Type: new Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. As a case study of this ...

📖 Read original article


11. When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide ​

Author: Binshuang Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.ME, stat.ML

arXiv:2608.12489v1 Announce Type: new Abstract: Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes al...

📖 Read original article


12. Exploring Oversmoothing with Householder Matrices ​

Author: Bhaskar Karol
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12514v1 Announce Type: new Abstract: Deep graph neural networks(GNNs) suffer from oversmoothing- a progressive collapse of node representation towards a low information subspace as network depth increases because the normalized graph propagation operator is repeatedly applied directly to ...

📖 Read original article


13. GENADA: efficient generative time series adversarial attack framework ​

Author: Michael Baronov, Denis Vorobev, Margarita Rusanova, Petr Sokerin, Alexey Zaytsev
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12535v1 Announce Type: new Abstract: Deep learning models are widely used for time series analysis in domains such as healthcare, finance, energy systems, and environmental monitoring. However, these models remain vulnerable to adversarial attacks, where small input perturbations cause se...

📖 Read original article


14. Scaling Automatic Research Agents via World Models ​

Author: Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Chenlei Guo, Jingrui He, Zhenyu Liao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12564v1 Announce Type: new Abstract: Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcome...

📖 Read original article


15. Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection ​

Author: Tadeusz Dziarmaga, Witold Sikora, {\L}ukasz Struski, Jacek Tabor, Marcin Mazur
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12573v1 Announce Type: new Abstract: Top-k selection is a fundamental computational primitive with applications spanning databases, information retrieval, signal processing, and modern machine learning workloads, including sparse activations and attention pruning. As data sizes grow, exis...

📖 Read original article


16. Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness ​

Author: Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Knoz, Tianlong Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12592v1 Announce Type: new Abstract: Continuous physiological time series underpin modern clinical monitoring, yet many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. Conditional generation offers a remedy: an absent signal can be synth...

📖 Read original article


17. Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks ​

Author: Andrew Cheng, Ali Eslamian, Jie Cheng, Mehdi Zargham, Qiang Cheng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12597v1 Announce Type: new Abstract: Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent...

📖 Read original article


18. The Boolean Power of ReLU ​

Author: Pablo Barcel'o, Floris Geerts, Matthias Lanzinger, Klara Pakhomenko, Jan Van den Bussche
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.LO

arXiv:2608.12617v1 Announce Type: new Abstract: We prove that, on finite simple undirected graphs equipped with a single Boolean node feature, the Boolean queries expressible in $\Sigma$-MPLang, for any collection $\Sigma$ of eventually constant activation functions and with arbitrary real coefficie...

📖 Read original article


19. Structure-preserving uncertainty quantification for GENERIC dynamics ​

Author: Zequn He, Celia Reina
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph

arXiv:2608.12624v1 Announce Type: new Abstract: Structure-preserving machine learning embeds physical structure directly into model architectures, yet uncertainty quantification (UQ) for such hard-constrained models remains limited because standard UQ methods may violate the encoded admissibility co...

📖 Read original article


20. CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution ​

Author: Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12629v1 Announce Type: new Abstract: GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce. Agents usually treat the compiler as a fixed black box and receive only errors, correctness outcomes, and timing, while existing DS...

📖 Read original article


21. Interpretable Causal Discovery via Causal-Effect Constraints ​

Author: Cixuan Zhang, Guy Van den Broeck, Benjie Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.12640v1 Announce Type: new Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized p...

📖 Read original article


22. Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks ​

Author: Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2608.12655v1 Announce Type: new Abstract: A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer. We introduce Training Under Challenge, an executable-certificate framework in w...

📖 Read original article


23. Demand Transfer Estimation at Scale via Restricted Logit Modeling ​

Author: Lakshya Garg, Deep Narayan Mishra, Swapnil Yadav, Haoan Wang, Sujal Alugubelli, Karthik Kumaran, Anupriya Sharma
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12680v1 Announce Type: new Abstract: Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e. expected demand) with...

📖 Read original article


24. Finding the Needle in a Haystack: Test-Time Analog Circuit Representation Adaptation for Bayesian Optimization ​

Author: Fin Amin, Sounak Dutta, Paul D. Franzon
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12687v1 Announce Type: new Abstract: Bayesian optimization (BO) is a sample-efficient framework for analog circuit topology search, where evaluating each candidate topology can require costly simulation. However, representation-based BO methods typically treat circuit embeddings as fixed ...

📖 Read original article


25. The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning ​

Author: Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, eess.SP

arXiv:2608.12695v1 Announce Type: new Abstract: Self-supervised electrocardiogram (ECG) models are often trained on a few seconds of ECG signal and, increasingly, on discretized token sequences. It remains unclear whether these choices sacrifice information needed for rhythm inference and longitudin...

📖 Read original article


26. A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family ​

Author: Rishi Shah, Rishav Shrestha
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AR, cs.DC

arXiv:2608.12700v1 Announce Type: new Abstract: Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can ...

📖 Read original article


27. Federated Compositional Muon Optimizer for Matrix-Wise Models ​

Author: Wang Yan, Feihu Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2608.12710v1 Announce Type: new Abstract: Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are still not particularly well-suited for hierarchical structured problems. To fill this gap...

📖 Read original article


28. Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia ​

Author: Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.12717v1 Announce Type: new Abstract: Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human ...

📖 Read original article


29. MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning ​

Author: Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi, Shijie Li, Kangkang Lu, Bharadwaj Veeravalli, Nancy F. Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12724v1 Announce Type: new Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled mul...

📖 Read original article


30. A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings ​

Author: Hei Ting (Una), Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng, Luis Filipe Nakayama, Michael G. Morley, Liyue Shen, Jiasi Chen, Z. Morley Mao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.DC

arXiv:2608.12745v1 Announce Type: new Abstract: Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating...

📖 Read original article


31. Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry ​

Author: Larissa Xu, King Bi, William Chang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (...

📖 Read original article


32. Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents ​

Author: Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12764v1 Announce Type: new Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPSD)...

📖 Read original article


33. CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility ​

Author: Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identifica...

📖 Read original article


34. HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models ​

Author: Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static de...

📖 Read original article


35. Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing ​

Author: Yuxiao Wen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12831v1 Announce Type: new Abstract: Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design gives eac...

📖 Read original article


36. A Compositional Theory of Curvature in Probabilistic Circuits ​

Author: Hrithik Suresh, Sahil Sidheekh, Shelar Parth Vijay, Yasir Z, Sriraam Natarajan, Narayanan Chatapuram Krishnan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12869v1 Announce Type: new Abstract: Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes th...

📖 Read original article


37. Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning ​

Author: Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12874v1 Announce Type: new Abstract: Plasticity loss has emerged as a critical challenge in continual learning that significantly hinders the acquisition of sequential tasks. While optimizing activation designs offers a potential solution, current fixed-form functions suffer from an inher...

📖 Read original article


38. Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection ​

Author: Pongpisit Thanasutives, Yoshinobu Kawahara
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, math.DS

arXiv:2608.12879v1 Announce Type: new Abstract: Fractional partial differential equations describe nonlocal dynamics, but discovering them from noisy data is difficult because fractional differentiation amplifies high-frequency measurement noise and the derivative orders are unknown. We propose Weak...

📖 Read original article


39. Adaptive $k$ Nearest Neighbors Classifier via Granular Ball Computing ​

Author: Xiaoyu Lian, Shuyin Xia, Hongxuan He, Lifeng Shen, Guoyin Wang, Xinbo Gao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12903v1 Announce Type: new Abstract: The $k$-Nearest Neighbor~(KNN) algorithm is widely used across various tasks. The selection of the $k$ value is a key issue because it significantly impacts performance. In this paper, an adaptive and efficient KNN approach via granular-ball computing ...

📖 Read original article


40. EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction ​

Author: Danyu Li, Ling Zhou, Rubing Huang, Xian Zhong, Bin Zou, Kui Jiang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12906v1 Announce Type: new Abstract: RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI Pre...

📖 Read original article


41. Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection ​

Author: Pu Li, Tao Tan, Hong Xie, Xiaoyu Shi, Mingsheng Shang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12912v1 Announce Type: new Abstract: This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find that the large action space increases the randomness in Q-value estimati...

📖 Read original article


42. Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling ​

Author: Takieddine Soualhi (CHROMA), Jacques Saraydaryan (CPE, CHROMA), Laetitia Matignon (UCBL)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.RO

arXiv:2608.12917v1 Announce Type: new Abstract: Developing effective robot navigation methods in crowded environments is essential for real-world applications. Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments, they often focus pri...

📖 Read original article


43. Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization ​

Author: Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, Shengbo Eben Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12925v1 Announce Type: new Abstract: Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain only partially understood. We develop an $\textbf{A}$DMM-$\textbf{I}$nspired $\textbf{M}$omentum (AI...

📖 Read original article


44. H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities ​

Author: Julius Broermann, Oliver M"uller, Michael D"oring, Jochen Baumeister
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12926v1 Announce Type: new Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions ...

📖 Read original article


45. Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance ​

Author: Chukwunonso Henry Nwokoye, Blessing Oluchi Iloka, Chikwue V. Umeugoji, Christopher Anene Egemba, Nnenna D. Duroha
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.NI

arXiv:2608.12929v1 Announce Type: new Abstract: The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. We compared the predictive power of network, environmental, device, and vision feature groups for 6GBO. ...

📖 Read original article


46. Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency ​

Author: Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can ...

📖 Read original article


47. CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation ​

Author: Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, eess.IV, stat.ML

arXiv:2608.12944v1 Announce Type: new Abstract: Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology ac...

📖 Read original article


48. I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization ​

Author: Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.12957v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with dense token supervisio...

📖 Read original article


49. The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use ​

Author: Joyjeet Singh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.12959v1 Announce Type: new Abstract: Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective ...

📖 Read original article


50. Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice ​

Author: Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, Ka-Ho Chow
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CR

arXiv:2608.12962v1 Announce Type: new Abstract: Vertical Federated Learning (VFL) enables organizations holding complementary features of shared entities to collaborate and train models. In this setting, the initiator can withhold information about the learning task, while other contributors partici...

📖 Read original article


51. Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks" ​

Author: Orr Well, Idan Tarshish, Nur Lan, Roni Katzir
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.12974v1 Announce Type: new Abstract: McCoy & Griffiths (2025, henceforth M&G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al., 2017). They support this empirically by showing that meta-trained ne...

📖 Read original article


52. Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices ​

Author: Rekha, Santosh Singh, S. K. Neogy
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12982v1 Announce Type: new Abstract: In this research work, we are constructing the sensing matrix, which is essential for the success of the compressive sensing technique. We have chosen a learning-based technique for the construction of the sensing matrix. The novelty and uniqueness of ...

📖 Read original article


53. Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data ​

Author: Mahboobe Jadid, Melika Rezaye Garkani, Ali Mousavi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12989v1 Announce Type: new Abstract: Pretrained tabular foundation models have demonstrated strong predictive capability; however, their application to large-scale datasets remains constrained by the limited inference context. This paper introduces Balanced Adaptive Prototype Selection (B...

📖 Read original article


54. Incremental Evaluation and Training in Relational Deep Learning ​

Author: Jakub Pele\v{s}ka, Gustav \v{S}'ir
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.DB

arXiv:2608.13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, prevailing RDL evaluation practices rely on static, single-episode dataset snapshots, overlooking the ...

📖 Read original article


55. On the global feature importance for interpretable and trustworthy heat demand forecasting ​

Author: Milan Zdravkovi'c
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2608.13039v1 Announce Type: new Abstract: The paper introduces the ante-hoc Explainable AI methodology to assess the global feature importance of the Machine Learning models used for heat demand forecasting in intelligent control of District Heating Systems, with motivation to facilitate their...

📖 Read original article


56. Latent On-Policy Self-Distillation ​

Author: Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren, Shuicheng Yan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.13040v1 Announce Type: new Abstract: Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervisi...

📖 Read original article


57. A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits ​

Author: Gurbhit Chaurakoti, Harshit Kumar, Hani Kumar, Anurag Singh, Ram Asrey
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13073v1 Announce Type: new Abstract: Significant health risks are associated with the illegal, yet commonly practiced use of industrial-grade Calcium Carbide (CaC2) for ripening climacteric fruits like mango and banana, which leaves behind trace residues of arsenic and phosphorus. To addr...

📖 Read original article


58. Learning Discrete Decisions for MIPs with Constraint-Aware Diffusion ​

Author: Vincenzo Di Vito, Mehdi Taghizadeh, Deepjyoti Deka, Kaarthik Sundar, Ferdinando Fioretto
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13079v1 Announce Type: new Abstract: This paper proposes a novel learning-based approach to approximately solve instances of mixed-integer optimization problems. These problems are computationally challenging, as they require jointly determining discrete and continuous decisions while sat...

📖 Read original article


59. Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization ​

Author: Jinhyung Bae
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC

arXiv:2608.13087v1 Announce Type: new Abstract: Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything ha...

📖 Read original article


60. FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching ​

Author: Zhuohan Wang, Andreea Bacalum, Ollie Olby, Carmine Ventre, Namid Stillman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CE, q-fin.CP, q-fin.TR

arXiv:2608.13096v1 Announce Type: new Abstract: Limit order book (LOB) simulators are most useful to practitioners when they combine realistic market dynamics, computationally efficient sampling, controllable scenario generation, and the ability to generalize beyond the instruments seen during train...

📖 Read original article


61. Branch and Bound for Relational Verification of Neural Networks ​

Author: Kota Fukuda, Zhenya Zhang, Guanqin Zhang, Jianjun Zhao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13118v1 Announce Type: new Abstract: Verification of neural networks against relational specifications, such as global robustness, is crucial for safety-critical applications of cyber-physical systems (CPS), given their increasing adoption of AI components. Compared to simple trace proper...

📖 Read original article


62. ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning ​

Author: Qianqian Wang, Yunshan Li, Dawei Huang, Wenwu Gong, Lili Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.13190v1 Announce Type: new Abstract: Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable. However, existing methods often infer environments from a separate reference model and select representations before fitting th...

📖 Read original article


63. Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity ​

Author: Timilehin B. Aderinola, Ilaria D'Ascanio, Luca Palmerini, Lorenzo Chiari, Jochen Klenk, Clemens Becker, Brian Caulfield, Georgiana Ifrim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13197v1 Announce Type: new Abstract: Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. However, real-world falls are extremely rare: collecting 100 of them requires an estimated 100,000 d...

📖 Read original article


64. TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures ​

Author: Orkun Irsoy, Leman Akoglu, Osman Yagan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.SI

arXiv:2608.13212v1 Announce Type: new Abstract: Networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity. A node whose load exceeds its capacity fails and sheds its load onto its neighbors, which can trigger a system-wide cascade. We ...

📖 Read original article


65. History-informed Lagrangian Neural Networks ​

Author: Tianshuo Zhang, Xianglei Xing, Wenzhe Zhai, Jia Gao, He Cao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13215v1 Announce Type: new Abstract: Forecasting the long-horizon evolution of mechanical systems from position-only observations is a pivotal yet difficult task, as hidden velocities and trajectory-specific physical properties must be inferred simultaneously. Although physics-guided neur...

📖 Read original article


66. Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations ​

Author: Gaute Johannessen, Geert Roelof van der Ploeg, Evrim Acar
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13234v1 Announce Type: new Abstract: In order to understand complex systems such as the human metabolome or human brain, different sensing technologies are used, generating complex data. These datasets are often multiway, i.e., with more than two axes of variation such as a subjects by me...

📖 Read original article


67. Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data ​

Author: Francesca Pia Panaccione, Sofia Mongardi, Marco Masseroli, Pietro Pinoli
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.13256v1 Announce Type: new Abstract: As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data genera...

📖 Read original article


68. Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations ​

Author: Berk Hadzhamolla, Alexander Johannes Stasik, Signe Riemer-S{\o}rensen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13260v1 Announce Type: new Abstract: Accurate modeling and forecasting of power transformer thermal behavior are critical for reliability, asset lifetime, and optimized power system operation. Numerical approaches such as finite element methods (FEM) and computational fluid dynamics (CFD)...

📖 Read original article


69. Into the ORBIT for Time Series: Training Regimes for Foundation Models ​

Author: Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.13262v1 Announce Type: new Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often poorly controlled with...

📖 Read original article


70. EEG Decoding Using CNN and LSTM Network ​

Author: Athanasios Karagounis
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.HC

arXiv:2608.13285v1 Announce Type: new Abstract: Motor imagery (MI) brain--computer interfaces (BCIs) have emerged as a promising approach for establishing flexible communication pathways between the human brain and external devices , particularly for individuals affected by stroke or neurodegenerati...

📖 Read original article


71. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks ​

Author: Wojciech Zarzecki, Jaros{\l}aw Arabas
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.13296v1 Announce Type: new Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. This causes a risk of biasing the development of global optimization methods. We argue that t...

📖 Read original article


72. The Time Value of Evolution ​

Author: Matthew Siper, Ahmed Khalifa, Julian Togelius
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13297v1 Announce Type: new Abstract: In evolutionary search, a weak child can be a valuable ancestor that makes high-fitness regions reachable. Immediate-return control is blind to this delayed utility, penalizing mutations through their immediate offspring even when they open productive ...

📖 Read original article


73. A Probe Direction Is a Property of Its Prompt ​

Author: Valentin No"el
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts th...

📖 Read original article


74. Training AI Scientists to Replicate Research ​

Author: Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.13331v1 Announce Type: new Abstract: The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, a...

📖 Read original article


75. Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws ​

Author: Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cond-mat.stat-mech

arXiv:2608.13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training losses instead follow smooth power laws. Variants of both behaviors occu...

📖 Read original article


76. Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation ​

Author: Valentin No"el
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13337v1 Announce Type: new Abstract: Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be measured at one of them. ...

📖 Read original article


77. Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples ​

Author: Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.13341v1 Announce Type: new Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is diffi...

📖 Read original article


78. When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation ​

Author: Shuhan Wang, Yilin Luo, Nan Xu, Chi Wang Cheung
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13365v1 Announce Type: new Abstract: Rotation-based post-training quantisation commonly applies an orthogonal transform across an entire attention head to reduce outlier-induced error. RoPE instead partitions each head into two-dimensional frequency pairs, raising the question of whether ...

📖 Read original article


79. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference ​

Author: Zixuan Lan, Yanhong Li, Jiawei Zhou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.13426v1 Announce Type: new Abstract: Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method ...

📖 Read original article


80. Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion ​

Author: Van Khoa Nguyen, Alexandros Kalousis
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13457v1 Announce Type: new Abstract: Generating crystals has recently attracted significant interest due to their broad applications in materials science. However, existing generative models struggle to produce complete crystallographic specifications, limiting their ability to capture gl...

📖 Read original article


81. Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization ​

Author: Jiayi Dan, Bo Li, Lu Deng, Yong Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13461v1 Announce Type: new Abstract: Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of...

📖 Read original article


82. Concept Drift Detection and Adaptive Retraining of Malware Classification Models ​

Author: Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2608.13465v1 Announce Type: new Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performanc...

📖 Read original article


83. Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling ​

Author: Yuchen Xin, Zhihua Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13467v1 Announce Type: new Abstract: We study the Moreau--Yosida unadjusted Langevin algorithm (MYULA) for the nonsmooth composite target [ \pi(dx)\propto \exp{-f(x)-g(x)},dx, \qquad x\in\mathbb R^d, ] where (f) is (m)-strongly convex with (L_f)-Lipschitz gradient and (g) is ...

📖 Read original article


84. Synthetic Persona Pretraining: Alignment from Token Zero ​

Author: Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.13482v1 Announce Type: new Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once ...

📖 Read original article


85. Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration ​

Author: Sabin Roman, Ljupco Todorovski, Saso Dzeroski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13504v1 Announce Type: new Abstract: We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. SORT estimates expansion coefficients directly from observations using L1-regul...

📖 Read original article


86. Intern-S2-Preview: Scientific Agentic Foundation Model ​

Author: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, cs.CV

arXiv:2608.13505v1 Announce Type: new Abstract: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a...

📖 Read original article


87. Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology ​

Author: Yunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng, Mary M. Maleckar, Nassir Marrouche, Jihun Hamm
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.13518v1 Announce Type: new Abstract: Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, medication chan...

📖 Read original article


88. The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity ​

Author: Martin J. Wainwright
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT, math.ST, stat.ML, stat.TH

arXiv:2608.13520v1 Announce Type: new Abstract: We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}). Its local increments directly control Kullback--Leibler (KL) discretization e...

📖 Read original article


89. Vero: Can AI Agents Build Formally Verified Software Repositories? ​

Author: Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO, cs.PL, cs.SE

arXiv:2608.13522v1 Announce Type: new Abstract: AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers...

📖 Read original article


90. DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees ​

Author: Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.13524v1 Announce Type: new Abstract: Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise...

📖 Read original article


91. Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure ​

Author: Mingyuan Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2608.13549v1 Announce Type: new Abstract: The per-instance Jaccard score, or intersection over union (IoU), is standard in multi-label classification and binary segmentation. With $s$ labels, its loss matrix has $2^s$ outcomes and reports. Under the convention $\mathrm{Jac}(\varnothing,\varnot...

📖 Read original article


92. Defensive Boosting for Online Probabilistic Forecasting ​

Author: Georgy Noarov, Aaron Roth
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CC, cs.DS, stat.ML

arXiv:2608.13554v1 Announce Type: new Abstract: We study online probabilistic forecasting of binary outcomes chosen by an adaptive adversary. Given an online learning algorithm for a weak hypothesis class $H$, we would like to efficiently obtain two incomparable guarantees that existing online boost...

📖 Read original article


93. Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability ​

Author: Frederick Hayes III
Published: 8/14/2026, 4:00:00 AM
Categories: cs.NE, cs.LG

arXiv:2608.11506v1 Announce Type: cross Abstract: Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation. Drawing on Barrett and Miller's account of categorization as predictive, compressive, functionally organized, and ...

📖 Read original article


94. RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning ​

Author: Yibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DC, cs.LG

arXiv:2608.12146v1 Announce Type: cross Abstract: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on ...

📖 Read original article


95. Position: Reasoning is a Learnable Rule-Based Process ​

Author: Rachel Lawrence, Jacqueline Maasch
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.12325v1 Announce Type: cross Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and...

📖 Read original article


96. Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition ​

Author: Suman Paudel, Sarbin Sayami
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and...

📖 Read original article


97. Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation? ​

Author: Hyowon Wi, Noseong Park
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.12332v1 Announce Type: cross Abstract: In recent years, low-rank adaptation (LoRA) has emerged as a significant paradigm that freezes pre-trained weights and introduces small, learnable adapters instead of fine-tuning the full set of parameters. In this work, we uncover several key insigh...

📖 Read original article


98. Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents ​

Author: Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.12342v1 Announce Type: cross Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock pri...

📖 Read original article


99. Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices ​

Author: Ahmed M Salih, Oliver D'iaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.LG

arXiv:2608.12360v1 Announce Type: cross Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks. As these systems become more integrated into clinical decision-making, there is growing expectation that they demonstrate key dime...

📖 Read original article


100. EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector ​

Author: Javier Gonzalez-Ruiz, Carlos Rodriguez-Pardo, Alice Di Bella, Paolo Mastropietro, Jose Pablo Chavez-Avila, Massimo Tavoni
Published: 8/14/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.CY, cs.LG, cs.MA, cs.SY, eess.SY, q-fin.EC

arXiv:2608.12363v1 Announce Type: cross Abstract: European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. A notable example is Italy's 2026 Decreto Bollette package, which proposes to...

📖 Read original article


101. A Bayes-Markov Neuromorphic Model of Cortical Orientation Selectivity: A Computational Re-implementation and Quantitative Simulation Study ​

Author: Abolfazl Moslemi, Milad Sarabadani, Fatemeh Sefidian, Hossein Peyvandi
Published: 8/14/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG, cs.NE

arXiv:2608.12388v1 Announce Type: cross Abstract: The emergence of orientation selectivity in the primary visual cortex (V1) remains a central question in computational neuroscience. Shirazi's Bayes-Markov model proposed a probabilistic explanation for how orientation-selective inhibition can arise ...

📖 Read original article


102. Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models ​

Author: Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings. However, ex...

📖 Read original article


103. Black-Box Knowledge Transfer across Distinct Feature Sets ​

Author: Oh-Ran Kwon, Daeyoung Ham
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME

arXiv:2608.12403v1 Announce Type: cross Abstract: Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation. However, when the available input features differ from those the black box expects, direct use is infeasible. We introduce a method...

📖 Read original article


104. Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex ​

Author: Nils Leutenegger
Published: 8/14/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG

arXiv:2608.12408v1 Announce Type: cross Abstract: Representational similarity analysis (RSA) is increasingly used to ask which learning rules give convolutional networks brain-like representations. Because biologically plausible rules such as feedback alignment, predictive coding and STDP do not sca...

📖 Read original article


105. AI-Driven Multiscenario Interest Rate Forecasting: A Proof of Concept for Banking Asset Management ​

Author: Ekkehardt Bauer, Dirk Holl"ander, Linus Wolff, Christoph Ostermair, Kyrillus Aiad, Joachim Hasebrook
Published: 8/14/2026, 4:00:00 AM
Categories: q-fin.CP, cs.LG, q-fin.ST

arXiv:2608.12424v1 Announce Type: cross Abstract: This study focuses on developing an AI-supported prototype for multiperspective interest rate forecasting that combines classical econometric models with modern artificial intel-ligence methods. Tested in a major European bank, the system enables mor...

📖 Read original article


106. SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization ​

Author: Yuanyu Li, Jintao Xu, Zijiang Liu, Yongzhi Qi, Ningxuan Kang, Jianshen Zhang, Wei Qi, Chen Xie, Zuo-Jun Max Shen
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.OC

arXiv:2608.12443v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single bes...

📖 Read original article


107. Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\&CK-Aligned Triage as a Worked Instance ​

Author: Zhenpeng Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2608.12444v1 Announce Type: cross Abstract: An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. We show this is structural: any risk certificate is defined over a decision contract, the inpu...

📖 Read original article


108. Fast Length-Squared Sampling for Positive-Semidefinite Matrices ​

Author: Rajarshi Bhattacharjee, Ethan N. Epperly, Cameron Musco, Aaron Tian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DS, cs.LG

arXiv:2608.12503v1 Announce Type: cross Abstract: We describe a simple rejection-sampling-based algorithm to perform length-squared sampling on an $n \times n$ positive-semidefinite (psd) matrix: that is, to sample a column with probability proportional to its squared $\ell_2$-norm. The algorithm ru...

📖 Read original article


109. Analysis of Motor Signatures of Social Adaptation in Autism for Efficient Human-Centric Systems ​

Author: Lara Pereira, Teresa Sousa, Miguel Castelo-Branco, Jo~ao Ruivo Paulo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.HC, cs.LG, eess.SP

arXiv:2608.12548v1 Announce Type: cross Abstract: Dance imitation integrates motor planning, sensorimotor integration, and social cognition, offering a sensitive framework to characterize motor behavior in autism. In this work, we explore a computational analysis framework to identify potential biom...

📖 Read original article


110. CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence ​

Author: Michael Georgiades, Charalambia Varnava
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.12555v1 Announce Type: cross Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS...

📖 Read original article


111. DYSANOS Generative Dynamic Smooth Arbitrage-free Non-parametric Option Surfaces ​

Author: Hans Buehler, Blanka Horvath, Anastasis Kratsios
Published: 8/14/2026, 4:00:00 AM
Categories: q-fin.MF, cs.LG

arXiv:2608.12587v1 Announce Type: cross Abstract: This article presents with DYSANOS the first generative market model for smooth SANOS option surfaces for all strikes and expiries which are free of static arbitrage. Our model is designed to generate entire paths of daily spot and option prices for ...

📖 Read original article


112. DiG-bench: Discovery in Games ​

Author: Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, J"urgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths, Zeb Kurth-Nelson, James C. R. Whittington
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.12593v1 Announce Type: cross Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge...

📖 Read original article


113. What Makes a Peer? Valuation-Anchored Similarity in Private Markets ​

Author: Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta
Published: 8/14/2026, 4:00:00 AM
Categories: q-fin.ST, cs.AI, cs.LG

arXiv:2608.12594v1 Announce Type: cross Abstract: As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due dili...

📖 Read original article


114. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues ​

Author: Haoyuan Zhu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.12599v1 Announce Type: cross Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{beh...

📖 Read original article


115. From Visual Widgets to UI Code: Efficient Tool-Grounded Generation ​

Author: Houston H. Zhang, Tao Zhang, Li Gu, Linfeng Ye, Yuanhao Yu, Xinxin Zuo, Yang Wang, Zhixiang Chi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.12611v1 Announce Type: cross Abstract: Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefin...

📖 Read original article


116. Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences ​

Author: Cosmin Dragoiu, Nooshin Nabizadeh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SD, cs.LG

arXiv:2608.12615v1 Announce Type: cross Abstract: In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. Using dashcam imagery an...

📖 Read original article


117. Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection ​

Author: Florian Braun
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at datas...

📖 Read original article


118. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries ​

Author: Oguz Serdar, Cuneyt Mertayak
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.12654v1 Announce Type: cross Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerB...

📖 Read original article


119. Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping ​

Author: Yuan Zhuang, Sanaa Hobeichi, Peng Shi, Fei Huang
Published: 8/14/2026, 4:00:00 AM
Categories: stat.AP, cs.LG, stat.ML

arXiv:2608.12663v1 Announce Type: cross Abstract: Wildfire susceptibility mapping typically relies on physical variables assembled from multiple remote-sensing, climate, and geospatial products. AlphaEarth Foundations (AEF) provides analysis-ready geospatial embeddings that may reduce this dependenc...

📖 Read original article


120. A Local-Linearly Convergent Algorithm for Nonconvex Equality-Constrained Optimization ​

Author: Frank E. Curtis, Lingjun Guo, Daniel P. Robinson
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG, stat.ML

arXiv:2608.12665v1 Announce Type: cross Abstract: For solving nonconvex equality-constrained optimization problems, a recent Gradient-Eigenstep Algorithm by Goyens et al.~is an iteration-efficient approach, based on minimizing Fletcher's augmented Lagrangian function, for finding an approximate seco...

📖 Read original article


121. Designing AI Pipelines for Decision-Ready ITSM Intelligence ​

Author: Archan Dutta, Yash Dharmadhikari, Marat Valiullin, Rahul Guha, Alexander Liss
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.12670v1 Announce Type: cross Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and...

📖 Read original article


122. Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level ​

Author: Yicong Jiang, Feihu Huang
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.12704v1 Announce Type: cross Abstract: Multi-objective bilevel optimization has wide applications in the AI area such as automated learning and multi-task meta-learning. Although recently some works have been begun to study the multi-objective bilevel optimization, the proposed methods re...

📖 Read original article


123. ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization ​

Author: Lixing Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.12756v1 Announce Type: cross Abstract: Adaptive latent tokenization maps a fine-grained input to a shorter sequence of continuous representations associated with input-dependent spans. We introduce ReconSpan, which divides text into chunks that a backward decoder can reconstruct from a si...

📖 Read original article


124. Difference-of-Convex Regularization for Graph Learning by Differentiable Programming ​

Author: Liping Tao, Chee Wei Tan
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.12757v1 Announce Type: cross Abstract: Laplacian-regularized minimization is fundamental in signal processing and machine learning, but is limited by the dense and ill-conditioned nature of the graph Laplacian pseudoinverse. While the Laplacian itself is sparse, its pseudoinverse is dense...

📖 Read original article


125. CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers ​

Author: Ebenezer Tarubinga
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, eess.IV

arXiv:2608.12773v1 Announce Type: cross Abstract: Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident Res...

📖 Read original article


126. Thermodynamics of Learning: A Typed Four-Component Accounting of Memory, Fit, and Value ​

Author: Akihito Sudo
Published: 8/14/2026, 4:00:00 AM
Categories: cond-mat.stat-mech, cs.IT, cs.LG, math.IT

arXiv:2608.12791v1 Announce Type: cross Abstract: What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity. We develop a typed accounting for finite-state learning devices that separates four components: a training-side fit functional $\Phi...

📖 Read original article


127. Fine-tuned Normalizing Flows for ALICE Zero Degree Calorimeter Fast Simulation ​

Author: Emilia Majerz, Jacek Otwinowski, Witold Dzwinel, Jacek Kitowski
Published: 8/14/2026, 4:00:00 AM
Categories: physics.ins-det, cs.LG, hep-ex

arXiv:2608.12795v1 Announce Type: cross Abstract: Simulating the ALICE Zero Degree Calorimeter (ZDC) neutron detector responses at the LHC is computationally expensive, requiring complex Monte Carlo chains. We develop a generative surrogate, focusing on Normalizing Flows (NFs). Through transfer lear...

📖 Read original article


128. Distribution Steering via Sliced Optimal Transport Control ​

Author: Kaito Ito, Anqi Dong
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG, cs.SY, eess.SY, stat.ML

arXiv:2608.12828v1 Announce Type: cross Abstract: Distribution steering seeks feedback laws that drive the state law of a dynamical system between prescribed initial and terminal distributions. Optimal transport provides a natural geometric approach, but its implementation generally requires a trans...

📖 Read original article


129. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation ​

Author: Yuchen Zheng, Sihan Xu, Jingwen Yang, Xiangrui Cai, Haiwei Zhang, Xiaojie Yuan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2608.12845v1 Announce Type: cross Abstract: Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are...

📖 Read original article


130. Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics ​

Author: Dorottya Zelenyanszki, Zhe Hou, Kamanashis Biswas, Vallipuram Muthukkumarasamy
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2608.12864v1 Announce Type: cross Abstract: Public blockchain data enables large-scale DeFi-related analysis, but many existing approaches are application-specific, difficult to scale, or hard to interpret. This research proposes a scalable, application-agnostic framework for \emph{persistent ...

📖 Read original article


131. Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses ​

Author: Lei You
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.12935v1 Announce Type: cross Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfac...

📖 Read original article


132. Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization ​

Author: Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.12953v1 Announce Type: cross Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNI...

📖 Read original article


133. Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning ​

Author: Zijie Cheng, Yang Peng, Zhihua Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.12973v1 Announce Type: cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for bot...

📖 Read original article


134. From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion ​

Author: Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.13043v1 Announce Type: cross Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which ...

📖 Read original article


135. VALG: An Agentic System for ML Theory Research ​

Author: Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.OC, stat.ML

arXiv:2608.13060v1 Announce Type: cross Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem ...

📖 Read original article


136. Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) ​

Author: Sam Mao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.13063v1 Announce Type: cross Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, se...

📖 Read original article


137. Statistical Properties of Robust Learning under Distributional Shifts ​

Author: Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.13133v1 Announce Type: cross Abstract: Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data. Robust learning frameworks such as Distributionally Robust Optimization (DRO) and Robust Satisficing (RS) aim to ...

📖 Read original article


138. MergeOver: Post-Training Token Merging for Recursive Vision Transformers ​

Author: Junseo Kim, Uraz Odyurt, Amirreza Yousefzadeh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.13141v1 Announce Type: cross Abstract: Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware. While recursive...

📖 Read original article


139. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint ​

Author: Fnu Pramono, John Cai, Sourabh Kulkarni
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedural...

📖 Read original article


140. High-dimensional networks and mean squared error for possibly misspecified models ​

Author: Lourens Waldorp
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.13171v1 Announce Type: cross Abstract: To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a co...

📖 Read original article


141. Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich ​

Author: Han Dong, Jiaming Li, Yongqiang Gong, Ruixi Li, Yin Liu
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.OC, math.ST, stat.TH

arXiv:2608.13201v1 Announce Type: cross Abstract: We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). The core technical contribution is the Sinkhorn linearization -- the implicit-function sensiti...

📖 Read original article


142. Chance-constrained selection of sequential intervention strategies from counterfactual estimates ​

Author: Minkyoung Kim, Beakcheol Jang
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, stat.ML

arXiv:2608.13209v1 Announce Type: cross Abstract: Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew-hour budget. Choosing among them calls for the outcome and the cumulative cost each would produce, counterfactua...

📖 Read original article


143. Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test ​

Author: Saveliy Batruin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.13228v1 Announce Type: cross Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, res...

📖 Read original article


144. Foundations of Independent Component Analysis ​

Author: Patrick Forr'e
Published: 8/14/2026, 4:00:00 AM
Categories: math.ST, cs.LG, math.PR, stat.ML, stat.TH

arXiv:2608.13229v1 Announce Type: cross Abstract: We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop th...

📖 Read original article


145. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures ​

Author: Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG

arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when ...

📖 Read original article


146. Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services ​

Author: Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2608.13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but ...

📖 Read original article


147. Foundation models for movement data: Are they ready for prime-time? ​

Author: Alexander Br"auer, Benjamin Cauchi, Nils Strodthoff
Published: 8/14/2026, 4:00:00 AM
Categories: eess.SP, cs.LG

arXiv:2608.13316v1 Announce Type: cross Abstract: Foundation models (FMs) trained on large-scale accelerometer data have been proposed as general-purpose feature extractors for health monitoring, but systematic evidence of their advantages is lacking. We present the first comprehensive evaluation of...

📖 Read original article


148. Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning ​

Author: Yikai Xu, Zhao Chen, Jian Huang
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.13418v1 Announce Type: cross Abstract: Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of s...

📖 Read original article


149. LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles ​

Author: Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman, Md Rayhanur Rahman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.CR, cs.LG

arXiv:2608.13450v1 Announce Type: cross Abstract: Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming...

📖 Read original article


150. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification ​

Author: Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDI...

📖 Read original article


151. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval ​

Author: Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.13495v1 Announce Type: cross Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert...

📖 Read original article


152. Equivariant learning of a transferable three-dimensional classical density functional ​

Author: Bingqing Cheng
Published: 8/14/2026, 4:00:00 AM
Categories: cond-mat.stat-mech, cond-mat.soft, cs.LG, physics.chem-ph, physics.comp-ph

arXiv:2608.13506v1 Announce Type: cross Abstract: Liquids exhibit collective behavior that depends sensitively on thermodynamic conditions, interfaces and confinement, yet predicting each new state commonly requires a separate atomistic simulation. Classical density functional theory offers a reusab...

📖 Read original article


153. On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective ​

Author: Nestor R. Barraza, Gabriel Pena
Published: 8/14/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.TH

arXiv:2608.13510v1 Announce Type: cross Abstract: Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, wh...

📖 Read original article


154. TabSOM: A tabular-to-image encoding method based on self-organizing maps ​

Author: David Chushig-Muzo, Mar'ia 'Angeles Rodr'iguez de Cara, Eva Milara, Francisco J. Lara-Abelenda, Luis Zhinin-Vera, Diego H. Peluffo-Ord'o~nez
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.13513v1 Announce Type: cross Abstract: Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixe...

📖 Read original article


155. Bagging Robustly Learns VC Classes with Linear Sample Complexity ​

Author: Omar Montasser
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.DS, cs.LG

arXiv:2608.13514v1 Announce Type: cross Abstract: We revisit the problem of learning predictors robust to adversarial examples at test-time. We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over ...

📖 Read original article


156. Exponential quantum advantage for learning signals with a single qubit ​

Author: Ishaan Kannan, Sridhar Prabhu, Saeed A. Khan, Mandar M. Sohoni, Xingrui Song, Saswata Roy, Alen Senanian, Valla Fatemi, Peter L. McMahon, Jordan Cotler
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cs.IT, cs.LG, math.IT

arXiv:2608.13521v1 Announce Type: cross Abstract: Quantum technology has the potential to transform scientific discovery, but quantum advantages often require processing capabilities well beyond the reach of experimental platforms. We show that coupling a single controllable qubit to an otherwise co...

📖 Read original article


157. LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure ​

Author: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LIT...

📖 Read original article


158. Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA's Medical Device Clearance Policy ​

Author: Mohammad Zhalechian, Soroush Saghafian, Omar Robles
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.HC, math.OC, stat.ML

arXiv:2407.11823v4 Announce Type: replace Abstract: The United States Food and Drug Administration's (FDA's) 510(k) pathway allows manufacturers to gain medical device approval by demonstrating substantial equivalence to a legally marketed device. However, the inherent ambiguity of this regulatory p...

📖 Read original article


159. "Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning" ​

Author: Vyacheslav Kungurtsev, Leonardo Christov Moore, Gustav Sir, Martin Krutsky
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2501.05844v4 Announce Type: replace Abstract: Causal Learning has emerged as a major theme of research in statistics and machine learning in recent years, promising computational techniques to reveal ``true'' causality. In this paper, we critique the premise of causal learning by considering t...

📖 Read original article


160. Cueless EEG imagined speech for subject identification: dataset and benchmarks ​

Author: Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour, Hamid R. Rabiee
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2501.09700v2 Announce Type: replace Abstract: Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored the use of imagined speech with semantically meaningful words for subject identification, most have relied on...

📖 Read original article


161. Regularization can make diffusion models more efficient ​

Author: Mahsa Taheri, Johannes Lederer
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.ML, stat.TH

arXiv:2502.09151v4 Announce Type: replace Abstract: Diffusion models are one of the key architectures of generative AI. Their main drawback, however, is the computational costs. This study indicates that the concept of sparsity, well known especially in statistics, can provide a pathway to more effi...

📖 Read original article


162. Yes, Q-learning Helps Offline In-Context RL ​

Author: Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2502.17666v5 Announce Type: replace Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives wi...

📖 Read original article


163. Efficient Image Restoration with State-Dependent Forward Diffusion ​

Author: Ziwei Luo, Fredrik K. Gustafsson, Jens Sj"olund, Thomas B. Sch"on
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2505.16733v3 Announce Type: replace Abstract: This paper proposes to perform image restoration through a state-dependent mean-reverting forward diffusion (FoD) process. In contrast to traditional diffusion-based approaches that rely on a coupled forward-backward diffusion scheme, FoD directly ...

📖 Read original article


164. Trajectory First: A Curriculum for Discovering Diverse Policies ​

Author: Cornelius V. Braun, Sayantan Auddy, Marc Toussaint
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.RO

arXiv:2506.01568v4 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima. In this context, constrained diversity optimization has become a useful reinforcement learning (RL) framework for training a set ...

📖 Read original article


165. A Lyapunov Drift-Plus-Penalty Method Tailored for Reinforcement Learning with Queue Stability ​

Author: Wenhan Xu, Jiashuo Jiang, Lei Deng, Danny Hin-Kwok Tsang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2506.04291v2 Announce Type: replace Abstract: With the proliferation of Internet of Things (IoT) devices, the demand for addressing complex optimization challenges has intensified. The Lyapunov Drift-Plus-Penalty algorithm is a widely adopted approach for ensuring queue stability, and some res...

📖 Read original article


166. Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models ​

Author: Abdullah X
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2508.12220v2 Announce Type: replace Abstract: Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution controls while assigning r...

📖 Read original article


167. Performance-Carbon Trade-Offs across Architectural Biases in Shear Flow Forecasting ​

Author: Sophia N. Wilson, Jens Hesselbjerg Christensen, Raghavendra Selvan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2509.24517v3 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics), leading to large-scale models that require massive computational resources and result in considerable carbon footprin...

📖 Read original article


168. Bayesian Distributional Models of Executive Functioning ​

Author: Robert Kasumba, Zeyu Lu, Dom CP Marticorena, Mingyang Zhong, Paul Beggs, Anja Pahor, Geetha Ramani, Imani Goffney, Susanne M Jaeggi, Aaron R Seitz, Jacob R Gardner, Dennis L Barbour
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.HC

arXiv:2510.00387v4 Announce Type: replace Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian Distributional Active LEarning (DALE) perform in comparison to conventional Independent Maximum Like...

📖 Read original article


169. Online Correlation Clustering: Simultaneously Optimizing All $\ell_p$-norms ​

Author: Sami Davies, Benjamin Moseley, Heather Newman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.DM, cs.DS

arXiv:2510.15076v2 Announce Type: replace Abstract: The $\ell_p$-norm objectives for correlation clustering present a fundamental trade-off between minimizing total disagreements (the $\ell_1$-norm) and ensuring fairness to individual nodes (the $\ell_\infty$-norm). Surprisingly, in the offline sett...

📖 Read original article


170. Automated Design Optimization via Strategic Search with Large Language Models ​

Author: Anthony Carreon, Vansh Sharma, Venkat Raman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.MA

arXiv:2511.22651v2 Announce Type: replace Abstract: Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design parameters are difficult to define. Large language models (LLMs) offer a promising alternative by dynamically i...

📖 Read original article


171. Reduced Order Modeling for Tsunami Forecasting with Bayesian Hierarchical Pooling ​

Author: Shane X. Coffing, John Tipton, Arvind T. Mohan, Darren Engwirda
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph

arXiv:2512.19804v2 Announce Type: replace Abstract: Reduced-order models (ROMs) can represent spatiotemporal processes in significantly fewer dimensions and can often be solved many orders of magnitude faster than their governing partial differential equations (PDEs). For example, proper orthogonal ...

📖 Read original article


172. Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map ​

Author: Luis F. Rosario Freytes (University of Michigan)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.11618v3 Announce Type: replace Abstract: Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions answer different identity questions. We study the represented process available at a receiver: an actual facto...

📖 Read original article


173. Safe Exploration via Policy Priors ​

Author: Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2601.19612v4 Announce Type: replace Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated) environments. In this work, we tackle this challenge by utilizing suboptimal yet conservative policies (e.g.,...

📖 Read original article


174. SpinCastML an Open Decision-Making Application for Inverse Design of Electrospinning Manufacturing: A Machine Learning, Optimal Sampling and Inverse Monte Carlo Approach ​

Author: Elisa Roldan, Tasneem Sabir
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.09120v2 Announce Type: replace Abstract: Electrospinning is a powerful technique for producing micro to nanoscale fibers with application specific architectures. Small variations in solution or operating conditions can shift the jet regime, generating non Gaussian fiber diameter distribut...

📖 Read original article


175. Training and Benchmarking Code Generation for Physics-Inspired Animations ​

Author: Yanan Wang, Renxi Wang, Yongxin Wang, Xuezhi Liang, Fajri Koto, Timothy Baldwin, Xiaodan Liang, Haonan Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.10840v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving. However, their ability to generate executable code that visually depicts physical scenarios and their qua...

📖 Read original article


176. Physics-Informed Laplace Neural Operator for Solving Partial Differential Equations ​

Author: Heechang Kim, Qianying Cao, Hyomin Shin, Seungchul Lee, George Em Karniadakis, Minseok Choi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.12706v2 Announce Type: replace Abstract: Neural operators have emerged as fast surrogate solvers for parametric partial differential equations (PDEs). However, purely data-driven models often require extensive training data and can generalize poorly, especially in small-data regimes and u...

📖 Read original article


177. From Approximation Rates to Loss-Landscape Barrier Decay in Shallow ReLU Networks ​

Author: Saveliy Baturin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.17596v2 Announce Type: replace Abstract: We study pathwise connectivity of sublevel sets for one-hidden-layer ReLU networks with constrained first-layer weights and an $\ell_1$ penalty on the output layer. The data term is assumed convex and globally Lipschitz in the scalar logit. We firs...

📖 Read original article


178. A Prior-Aware Metric for Efficiently Distinguishing Memorization from Generalization in Large Language Models ​

Author: Trishita Tiwari, Ari Trachtenberg, G. Edward Suh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance. A central challenge in assessing this risk is distinguishing prefix-specific memorization of training data from ...

📖 Read original article


179. SEAR: Sample Efficient Action Chunking Reinforcement Learning ​

Author: C. F. Maximilian Nagy, Onur Celik, Emiliyan Gospodinov, Florian Seligmann, Weiran Liao, Aryan Kaushik, Gerhard Neumann
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2603.01891v2 Announce Type: replace Abstract: Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended action space at reduced decision frequency offsets these gains, leadin...

📖 Read original article


180. Distributed Online Submodular Maximization under Communication Delays: A Simultaneous Decision-Making Approach ​

Author: Zirui Xu, Vasileios Tzoumas
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.MA, cs.SY, eess.SY, math.OC

arXiv:2603.27803v2 Announce Type: replace Abstract: We provide a distributed online algorithm for multi-agent submodular maximization under communication delays. We are motivated by the future distributed information-gathering tasks in unknown and dynamic environments, where utility functions natura...

📖 Read original article


181. In-context superposition: human-like working memory interference in large language models ​

Author: Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.09670v3 Announce Type: replace Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments. This capacity, known as working memory, is fundamental to human reasoning. Yet, human working memory is strikingly limited, maintain...

📖 Read original article


182. Post-Hoc Uncertainty-Aware Explanations for Deployed Power Quality Disturbance Classifiers via Laplace Approximation ​

Author: Yinsong Chen, Samson S. Yu, Kashem M. Muttaqi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2604.13658v2 Announce Type: replace Abstract: Deep learning classifiers achieve high accuracy in power quality disturbance (PQD) recognition, but existing explanation methods return a single deterministic attribution map and provide no measure of its reliability. This paper develops a post-hoc...

📖 Read original article


183. OC-Distill: Ontology-aware Contrastive Learning with Cross-Modal Distillation for ICU Risk Prediction ​

Author: Zhongyuan Liang, Junhyung Jo, Hyang-Jung Lee, Sang Kyu Kim, Irene Y. Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2604.16878v4 Announce Type: replace Abstract: Early prediction of severe clinical deterioration and remaining length of stay can enable timely intervention and better resource allocation in high-acuity settings such as the ICU. This has driven the development of machine learning models that le...

📖 Read original article


184. The Optimal Sample Complexity of Multiclass and List Learning ​

Author: Chirag Pabbaraju
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2604.24749v2 Announce Type: replace Abstract: While the optimal sample complexity of binary classification in terms of the VC dimension is well-established, determining the optimal sample complexity of multiclass classification has remained open. The appropriate complexity parameter for multic...

📖 Read original article


185. Training Non-Differentiable Networks via Optimal Transport ​

Author: An T. Le
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.NE, cs.RO, math.OC

arXiv:2605.01928v2 Announce Type: replace Abstract: We optimize losses that jump: spiking thresholds, quantized layers, and discrete routing put jumps in the forward pass, where backpropagation does not apply. Finite differences fail: at a derivative-estimating radius, 99.5% of probe pairs on a quan...

📖 Read original article


186. SeBA: Semi-supervised few-shot learning via Separated-at-Birth Alignment for tabular data ​

Author: Kacper Jurek, Wojciech Batko, Marek 'Smieja, Marcin Przewi\k{e}'zlikowski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.08519v2 Announce Type: replace Abstract: Learning from scarce labeled data with a larger pool of unlabeled samples, known as semi-supervised few-shot learning (SS-FSL), remains critical for applications involving tabular data in domains like medicine, finance, and science. The existing SS...

📖 Read original article


187. SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models ​

Author: Chengjie Hong, Feixiang He, Yiheng Zeng, Lulu Kang, He Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.17985v2 Announce Type: replace Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential for reducing memory use and accelerating inference in large foundation models, it remains under-ex...

📖 Read original article


188. TabH2O: A Unified Foundation Model for Tabular Prediction ​

Author: Pascal Pfeiffer, Dmitry Gordeev, Mathias M"uller, Laura Fink, Joan Salv`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning. TabH2O builds on the TabICL architecture with several key modifications: (1) unified training, a si...

📖 Read original article


189. Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance ​

Author: Jing Chen, Shixiang Pan, Yujie Fan, Haocheng Ye, Haitao Xu, Wenqiang Xu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.18793v3 Announce Type: replace Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing methods face performance bottlenecks, typically yielding only incremental gains and often exhibiting ...

📖 Read original article


190. Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking ​

Author: Qinwu Xu, Zhuoheng Li, Jessie Salas
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2605.18852v2 Announce Type: replace Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be comparable to variability introdu...

📖 Read original article


191. INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification ​

Author: Seongjun Lee, Seokhyun Lee, Changhee Lee
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.20088v2 Announce Type: replace Abstract: Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent complexity of time-series classification (TSC) and to make model decision-making processes more transparent. Howev...

📖 Read original article


192. A Simple State Space Model Excels at Multivariate Time Series Classification ​

Author: Hassan Saadatmand, Geoffrey I. Webb, Hamid Rezatofighi, Mahsa Salehi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.27406v2 Announce Type: replace Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating strong performance through input-dependent state transitions, albeit at considerable complexity...

📖 Read original article


193. Annealed Softmax Greedy in Many-Armed Bayesian Bandits ​

Author: William Overman, Mohsen Bayati
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.31034v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher rewar...

📖 Read original article


194. Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems ​

Author: Jonathan Cola\c{c}o Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.00367v2 Announce Type: replace Abstract: Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to specify than scalar rewards, and they express certain goals that sc...

📖 Read original article


195. Constitutional On-Policy Safe Distillation ​

Author: Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.03089v3 Announce Type: replace Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervision. Prior work has shown that OPSD can collapse in verifiable reaso...

📖 Read original article


196. Do Transformers Need Three Projections? Systematic Study of QKV Variants ​

Author: Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.PF

arXiv:2606.04032v3 Announce Type: replace Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some ...

📖 Read original article


197. SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation ​

Author: Junghun Oh, Sungyong Baik, Kyoung Mu Lee
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.16454v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates with low-rank matrices. In this paper, we investigate the limitations of the LoRA parameterization from a geomet...

📖 Read original article


198. The Illusion of Improvement: Reject Inference Strategies in Credit Scoring ​

Author: Bruno Scarone, Ricardo Baeza-Yates
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CY

arXiv:2606.18479v2 Announce Type: replace Abstract: Reject inference methods are widely used to mitigate survival bias in credit scoring, yet their effectiveness remains poorly understood. We systematically evaluate several such methods and uncover a structural failure mode: in a natural retraining ...

📖 Read original article


199. Gradient-Free Warm-Start Library Recovery: an Amortized-Regret Separation ​

Author: Jianwei Lou (RailMind Systems, Neuss, Germany)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.NE

arXiv:2606.21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally. We give a provable account on recurring-regime streams. Given segmentation, a warm...

📖 Read original article


200. Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization ​

Author: Ryusei Yamada, Naoki Sato, Hideaki Iiduka
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.08104v2 Announce Type: replace Abstract: Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundame...

📖 Read original article


201. Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics ​

Author: Byung Gyu Chae
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.10923v4 Announce Type: replace Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning reorganizes the collective relaxation spectrum, the...

📖 Read original article


202. Scaling Time Series Classification via XAI-Driven Data Reduction ​

Author: Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15774v3 Announce Type: replace Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodolo...

📖 Read original article


203. Dimension-Calibrated Unexplained Mass: An Interpretable Drift Statistic for Contamination Monitoring in Data Streams ​

Author: Behnam Asadi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16811v4 Announce Type: replace Abstract: Drift detectors that work tend not to explain themselves, and drift detectors that explain themselves tend to fail in high dimension. We close that gap for Gaussian mixture models (GMMs): each fitted component is a named "regime," and the fraction ...

📖 Read original article


204. Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training ​

Author: Nuemaan Malik
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.19058v2 Announce Type: replace Abstract: Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE language model AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewA...

📖 Read original article


205. Wrong Design Intent Is Worse Than Never Conditioning: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion ​

Author: Yang Xiao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.GR

arXiv:2607.23191v4 Announce Type: replace Abstract: Fine-tuned code LLMs are routinely conditioned on a design-intent specification, but the correctness axis of such a signal -- a wrong intent rather than an absent one -- has not been tested, and the benefit of conditioning is usually scored with th...

📖 Read original article


206. SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups ​

Author: Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.27431v4 Announce Type: replace Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requires num...

📖 Read original article


207. Subtract, Transport, or Replay? Auditable Deletion from Language-Model Memory ​

Author: Vishwajith Ramesh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.27539v2 Announce Type: replace Abstract: Exact deletion from persistent language-model memory depends on whether a record's effect remains addressable after later computation. Native Kimi Delta Attention (KDA) gives a negative result for the tested receipt interface: the corpus-pooled raw...

📖 Read original article


208. Beckmann Transport Models: From Autonomous Flows to One-Step Maps ​

Author: Lee Cheuk-Kit, Florentin Coeurdoux, Yuyuan Chen, Sophia Tang, Peter Potaptchik, Yilun Du, Michael Samuel Albergo, Eric Vanden-Eijnden
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.01692v3 Announce Type: replace Abstract: We propose an instantiation of flow matching that relies on a time-independent velocity field (an \emph{autonomous flow}) to exactly map between two distributions, so long as the target is singular, i.e.\ supported on a lower-dimensional data manif...

📖 Read original article


209. Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality ​

Author: Shengzhi Deng, Chenqi Ye, Yanze Guo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2608.02575v2 Announce Type: replace Abstract: Digital learning systems consume concrete pseudorandom values rather than abstract random variables. These values enter the realized loss and its gradient during training. If a pseudorandom stream contains structure that is accessible to the model,...

📖 Read original article


210. Latent Fact-Checking: Detecting Misinformation through Activation Engineering ​

Author: Pedro T. Barcelos, Ot'avio Parraga, Marcelo M. Mussi, Lucas M. Fraga, Lucas S. Kupssinsk"u, Rodrigo C. Barros
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.06417v3 Announce Type: replace Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property o...

📖 Read original article


211. Which Decisions Low-Bit Quantization Breaks, and How to Predict Them ​

Author: Zekun Wu, Swati Dhiman, Adriano Koshiyama
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.06564v3 Announce Type: replace Abstract: Quantization is known to hurt below four bits, but nobody can say which of a model's decisions will change at a given bit-width. This matters most where a model acts rather than answers: a compressed agent stops calling its tools and, one bit lower...

📖 Read original article


212. Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks ​

Author: Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07335v2 Announce Type: replace Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm enables off-policy value learning without relying on experience replay buffe...

📖 Read original article


213. When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes ​

Author: Yu Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.PF

arXiv:2608.07911v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut expert traffic per t...

📖 Read original article


214. Neural Message Passing on Structural Interaction Graphs for Fully-Inductive Graph Neural Networks ​

Author: Omer Yom-Tov, Avigdor Gal
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.08567v2 Announce Type: replace Abstract: A central obstacle in building graph foundation models is the input heterogeneity in terms of feature space dimensionality, semantics, and structure. Such heterogeneity limits the capability of graph neural networks to generalize to new graphs with...

📖 Read original article


215. DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models ​

Author: Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.09233v2 Announce Type: replace Abstract: Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learnin...

📖 Read original article


216. From Recoverability to Functional Use: Certifying Temporal Reports in Time-Series Forecasting ​

Author: Qipeng Qian, Yuntao Qian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.10433v2 Announce Type: replace Abstract: Temporal reports are increasingly emitted alongside numerical forecasts and are often interpreted as statements about the computation producing those forecasts. We formalize the resulting certification problem as three distinct stages: \emph{recove...

📖 Read original article


217. Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning ​

Author: Daoyi Li, Yixian Zhang, Wenbo Ding, Yu Wang, Chao Yu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.10473v2 Announce Type: replace Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online fine-tuning: as the policy...

📖 Read original article


218. PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR ​

Author: Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.11368v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to a pointwise noti...

📖 Read original article


219. REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ​

Author: Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.11698v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation,...

📖 Read original article


220. Task- and dataset-specific information in protein language models ​

Author: Roman Joeres, Ilya Senatorov, Anastasia Kolchina, Dietrich Klakow, Olga V. Kalinina
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, q-bio.BM

arXiv:2608.12090v2 Announce Type: replace Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into l...

📖 Read original article


221. Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning ​

Author: Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.12108v2 Announce Type: replace Abstract: Federated learning (FL) enables collaborative model training across distributed clients while keeping data local. A central challenge is determining which client updates are beneficial for aggregation with respect to each client's target domain. Ex...

📖 Read original article


222. Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting ​

Author: Qifan Zhang, Yunhui Guo, Yu Xiang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2407.13911v5 Announce Type: replace-cross Abstract: Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone. However, the effect of backbone scale rema...

📖 Read original article


223. Langevin dynamics for high-dimensional optimization: the case of multi-spiked tensor PCA ​

Author: G'erard Ben Arous, C'edric Gerbelot, Vanessa Piccolo
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.PR, math.ST, stat.TH

arXiv:2408.06401v3 Announce Type: replace-cross Abstract: We study nonconvex optimization in high dimensions through Langevin dynamics, focusing on the multi-spiked tensor PCA problem. In this tensor estimation model, the goal is to recover a finite number of hidden signal vectors, or spikes, from n...

📖 Read original article


224. Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries ​

Author: Harshavardhan Battula, Jiacheng Liu, Jaideep Srivastava
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2411.16818v2 Announce Type: replace-cross Abstract: To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU) notes are fused with physiology for in-hospital mortality (IHM) prediction, and to determine how much o...

📖 Read original article


225. MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence ​

Author: Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Emily Mallaber, Harry R. Klein, Gufran Gungor, Matthew Galvin, Michael Deletto, Sabrina Y. Camp, Stephen C. Van Nostrand, James Provencher, Joyce Yu, Naeem Tahir, Jonathan Wischhusen, Olga Kozyreva, Taylor Ortiz, Hande Tuncer, Jad El Masri, Alys Malcolm, Tali Mazor, Ethan Cerami, Kenneth L. Kehl
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2412.17228v4 Announce Type: replace-cross Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. Open-source AI trial matching tools could democratize access to trial options. Methods: We create...

📖 Read original article


226. Optimizing Likelihoods via Mutual Information: Bridging Simulation-Based Inference and Bayesian Optimal Experimental Design ​

Author: Vincent D. Zaballa, Elliot E. Hui
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2502.08004v2 Announce Type: replace-cross Abstract: Simulation-based inference (SBI) is a method to perform inference on a variety of complex scientific models with challenging inference (inverse) problems. Bayesian Optimal Experimental Design (BOED) aims to efficiently use experimental resour...

📖 Read original article


227. Multiview Representation Learning via Distributed Joint Latent Space Structuring ​

Author: Milad Sefidgaran, Piotr Krasnowski, Abdellatif Zaidi
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.IT, cs.LG, math.IT

arXiv:2504.18455v2 Announce Type: replace-cross Abstract: We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view. The clients independently extract local representations from their views, which are the...

📖 Read original article


228. Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision ​

Author: Ahmet Bilican, M. Ak{\i}n Y{\i}lmaz, A. Murat Tekalp, R. G"okberk Cinbi\c{s}
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV, eess.SP

arXiv:2505.12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA achieve efficiency through low-rank updates, their discrete rank constraint limits...

📖 Read original article


229. Accelerated Markov Chain Monte Carlo Algorithms on Discrete States ​

Author: Bohan Zhou, Shu Liu, Xinzhe Zuo, Wuchen Li
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG, stat.CO

arXiv:2505.12599v3 Announce Type: replace-cross Abstract: We propose a class of discrete state sampling algorithms based on Nesterov's accelerated gradient method, which extends the classical Metropolis-Hastings (MH) algorithm. The evolution of the discrete states probability distribution governed b...

📖 Read original article


230. In Silico Study for Optimizing Intensity and Focality Electrode Configurations for Directional DBS Under Uncertainty Using Metaheuristic L1L1 Method ​

Author: Fernando Galaz Prieto, Antti Lassila, Maryam Samavaki, Sampsa Pursiainen
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2506.13452v3 Announce Type: replace-cross Abstract: Background and Objective: As Deep Brain Stimulation (DBS) advances toward directional leads and optimization-based current steering, selecting electrode contact configurations becomes complex. This study formulates configuration selection as ...

📖 Read original article


231. Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control ​

Author: Yujie Liu, Vincent Y. F. Tan, Yunbei Xu
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.IT, cs.LG, math.IT

arXiv:2506.18278v4 Announce Type: replace-cross Abstract: We introduce an original minimax framework for finite-time performance analysis in queueing control and propose a surprisingly simple Lyapunov-based scheduling policy with superior finite-time performance. The framework quantitatively charact...

📖 Read original article


232. DiffGRM: Diffusion-based Generative Recommendation Model ​

Author: Zhao Liu, Yichen Zhu, Yiqing Yang, Xiao Lv, Guoping Tang, Rui Huang, Qiang Luo, Ruiming Tang, Kun Gai, Guorui Zhou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2510.21805v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predicts the next item by autoregressively generating its SID conditioned on the user's history. However, two...

📖 Read original article


233. Functional Adjoint Sampler: Scalable Sampling on Infinite Dimensional Spaces ​

Author: Byoungwoo Park, Juho Lee, Guan-Horng Liu
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2511.06239v2 Announce Type: replace-cross Abstract: Learning-based methods for sampling from the Gibbs distribution in finite-dimensional spaces have progressed quickly, yet theory and algorithmic design for infinite-dimensional function spaces remain limited. This gap persists despite their s...

📖 Read original article


234. Self-Localizing MIMO Beam Mapping for Intelligent Open RAN with Continuously Evolving Channel Memory ​

Author: Wangqian Chen, Junting Chen, Shuguang Cui
Published: 8/14/2026, 4:00:00 AM
Categories: eess.SP, cs.LG, cs.SY, eess.SY

arXiv:2511.17007v2 Announce Type: replace-cross Abstract: Open and intelligent radio access networks (RANs) envisioned for 6G require accurate and reusable wireless channel knowledge for intelligent inference and control. However, full-dimensional channel state information (CSI) and accurate locatio...

📖 Read original article


235. Embedding networks with the random walk first return time distribution ​

Author: Vedanta Thapar, Renaud Lambiotte, George T. Cantwell
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SI, cs.LG

arXiv:2512.02694v3 Announce Type: replace-cross Abstract: We propose the first return time distribution (FRTD) of a random walk as an interpretable and mathematically grounded node embedding. The FRTD assigns a probability mass function to each node, allowing us to define a distance between any pair...

📖 Read original article


236. Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models ​

Author: Malte Hellmeier
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2512.13325v2 Announce Type: replace-cross Abstract: Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing control over data when it is being used to train such machine learning models or when distinguishing model...

📖 Read original article


237. RadarGen: Automotive Radar Point Cloud Generation from Cameras ​

Author: Tomer Borreda, Fangqiang Ding, Sanja Fidler, Shengyu Huang, Or Litany
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO

arXiv:2512.17897v2 Announce Type: replace-cross Abstract: We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen adapts efficient image-latent diffusion to the radar domain by representing radar measurements in bird's-...

📖 Read original article


238. SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models ​

Author: Turhan Can Kargin, Wojciech Jasi'nski, Adam Pardyl, Bartosz Zieli'nski, Marcin Przewi\k{e}'zlikowski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2601.11729v2 Announce Type: replace-cross Abstract: Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work incorporates ...

📖 Read original article


239. Manifold constrained steepest descent for smooth and closed-set optimization ​

Author: Kaiwei Yang, Lexiao Lai
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2601.21487v2 Announce Type: replace-cross Abstract: We study minimization of smooth functions over feasible sets that have smooth embedded-manifold structure throughout or only on selected regions, using linear minimization oracles (LMOs) to determine search directions under user-chosen norms....

📖 Read original article


240. Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise ​

Author: Puwei Lian, Yujun Cai, Songze Li, Bingkun Bao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2601.21628v2 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable progress in image generation, but their increasing deployment raises serious concerns about privacy and copyright. In particular, fine-tuned models are highly vulnerable, as they are often fine-tuned ...

📖 Read original article


241. Variance Reduction Based Experience Replay for Policy Optimization ​

Author: Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2602.05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization. However, classical experience replay treats all past observations unifor...

📖 Read original article


242. Stochastic Neural Networks for Quantum Devices ​

Author: Bodo Rosenhahn, Tobias J. Osborne, Christoph Hirche
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2602.22241v2 Announce Type: replace-cross Abstract: This work presents a formulation to express and optimize stochastic neural networks as quantum circuits in gate-based quantum computing. Motivated by a classical perceptron, stochastic artificial neurons are introduced and combined into a qua...

📖 Read original article


243. General Bayesian Policy Learning ​

Author: Masahiro Kato
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, econ.EM, math.ST, stat.ME, stat.TH

arXiv:2602.23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from a given set to maximize expected welfare. Typical examples include treatment choice and portfoli...

📖 Read original article


244. Minimax and Adaptive Covariance Matrix Estimation under Differential Privacy ​

Author: T. Tony Cai, Yicheng Li
Published: 8/14/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.TH

arXiv:2603.19703v2 Announce Type: replace-cross Abstract: Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-...

📖 Read original article


245. Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI ​

Author: Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk, Andy Nkansah
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA

arXiv:2603.25821v3 Announce Type: replace-cross Abstract: We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions. Unlike traditional medical benchmarks that rely on solving standardized test questions, the proposed ...

📖 Read original article


246. Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models ​

Author: Yiyang Zhang, Chaojian Yu, Ziming Hong, Yuanjie Shao, Qinmu Peng, Tongliang Liu, Xinge You
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2604.05809v2 Announce Type: replace-cross Abstract: This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs. Most existing backdoor attac...

📖 Read original article


247. Identifiability and Stability of Generative Drifting in the Companion-Elliptic Kernel Family ​

Author: HakGeun Lee, Hyonho Chun
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2604.24196v4 Announce Type: replace-cross Abstract: A drifting model is a one-step generator trained by moving each sample along a field of kernel-weighted attraction toward data samples and repulsion between model samples; training halts once this field vanishes. The soundness of this scheme ...

📖 Read original article


248. Functional-prior-based approaches to Bayesian PDE-constrained inversion using physics-informed neural networks ​

Author: Ryoichiro Agata, Tomohisa Okazaki
Published: 8/14/2026, 4:00:00 AM
Categories: physics.geo-ph, cs.LG, physics.comp-ph, stat.ML

arXiv:2605.07060v3 Announce Type: replace-cross Abstract: Physics-informed neural networks (PINNs) provide a mesh-free framework for solving PDE-constrained inverse problems, but their extension to Bayesian inversion still faces a fundamental difficulty: prior distributions are typically defined in ...

📖 Read original article


249. Identifiability and Estimation for Unlabeled Finite Mixtures under Marginal Independence ​

Author: Takafumi Kanamori, Yushi Hirose, Shohei Yamamoto
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2606.07914v2 Announce Type: replace-cross Abstract: We study component recovery and mixing-matrix estimation from unlabeled finite mixtures whose observable distributions share the same latent components but have unknown mixing weights. The main identifying signal is marginal independence: eac...

📖 Read original article


250. Liquidity-Based Audit of Algorithmic Trading Strategies ​

Author: Irene Aldridge
Published: 8/14/2026, 4:00:00 AM
Categories: econ.EM, cs.LG, q-fin.CP, q-fin.RM, stat.ML

arXiv:2606.29018v2 Announce Type: replace-cross Abstract: We show that net demand for liquidity by algo strategies is identifiable from its trade and price history alone, with no knowledge of its signal or optimization problem. An exact multi-period regret decomposition implies that the sign of this...

📖 Read original article


251. SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers ​

Author: Jianing Deng, Yuanzhe Li, Jialu Wang, Song Wang, Tianlong Chen, Huanrui Yang, Jingtong Hu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.03612v2 Announce Type: replace-cross Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. However, scaling them to long image sequences remains challenging, as the quadratic complexity of cross-view global attention quickly becomes the dom...

📖 Read original article


252. The Noise Premium in Adversarial Training for Kernel Regression ​

Author: Yiling Xie, Xiaoming Huo
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.27995v2 Announce Type: replace-cross Abstract: Adversarial training can improve the robustness of predictive models to bounded perturbations, often at the cost of statistical efficiency. We study this trade-off in kernel regression over a reproducing kernel Hilbert space (RKHS). It is sho...

📖 Read original article


253. LOCUS-DT: Localization via Observation-Conditioned Uncertainty Scoring with Digital Twins ​

Author: Haozhe Lei, Roberto Bomfin, Marwa Chafii, Sundeep Rangan
Published: 8/14/2026, 4:00:00 AM
Categories: eess.SP, cs.LG, cs.RO

arXiv:2608.00406v2 Announce Type: replace-cross Abstract: Accurate indoor localization is essential for emerging applications in robotic navigation and search and rescue. While classical methods typically focus on single-point estimates, complex indoor environments with heavy blockage and multipath ...

📖 Read original article


254. Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors ​

Author: Alexander Scheinker
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, physics.comp-ph, physics.plasm-ph

arXiv:2608.00675v2 Announce Type: replace-cross Abstract: Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it against. We train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a ...

📖 Read original article


255. Thermalizing Stochastic Programs ​

Author: Mirko Amico, Andra\v{z} Jelin\v{c}i\v{c}, Colin Oscar Nancarrow, Leo Tyrpak, David Roberts, Seth Morton, Dalton Sakthivadivel, Ashwin Gopal, Guillaume Verdon
Published: 8/14/2026, 4:00:00 AM
Categories: cs.ET, cs.LG

arXiv:2608.01615v2 Announce Type: replace-cross Abstract: We present a set of tools for mapping general stochastic programs to thermodynamic hardware designed for energy-efficient stochastic sampling. Given a target stochastic program expressed as a Directed Factor Graph (DFG) of stochastic channels...

📖 Read original article


256. Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis ​

Author: Rishov Paul, Frederick H. Epstein, Miaomiao Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.01677v2 Announce Type: replace-cross Abstract: Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. However, current techniques require either human-adjusted post-processing with suboptimal regional accuracy, or ...

📖 Read original article


257. Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study ​

Author: Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.03172v3 Announce Type: replace-cross Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But th...

📖 Read original article


258. Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load ​

Author: Thomas Bartz-Beielstein, Inalbek Akiev, Lalo Mohamad
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05018v3 Announce Type: replace-cross Abstract: Short-term load forecasting (STLF) plays a vital role in the electric power industry. It is relevant for critical infrastructure. STLF is no longer purely a performance and accuracy problem, because determinism, fail-safe handling, minimal-at...

📖 Read original article


259. SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System ​

Author: Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DC, cs.LG

arXiv:2608.05033v2 Announce Type: replace-cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy. For the same SpMM on the same matrix, cuSPARSE ...

📖 Read original article


260. Recursive Synthesis for Long-Horizon Terminal Tasks ​

Author: Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05466v3 Announce Type: replace-cross Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually ...

📖 Read original article


261. LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection ​

Author: Peter Lorenz, Anjith George, Marcel S'ebastien
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.09633v2 Announce Type: replace-cross Abstract: Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variati...

📖 Read original article


262. AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting ​

Author: Anna Allen, Wessel P. Bruinsma, Michael Maier-Gerber, Harrison Cook, Matthew Chantry, Richard E. Turner
Published: 8/14/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.LG

arXiv:2608.09959v2 Announce Type: replace-cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to physics-based NWP in forecasting tropical cyclone (TC) tracks, they tend to dramatically undere...

📖 Read original article


263. Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness ​

Author: Srijith Ravikumar
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.CL, cs.LG

arXiv:2608.10008v2 Announce Type: replace-cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits report a binary out-of-domain rate; none ask whether the model knew. We jointly audit hallucination rate (OOD@10) and verbalized-confide...

📖 Read original article


264. On the Importance of Geometric Nonlinearity and Temperature-Dependent Properties in Multi-Material Thermo-Mechanical Topology Optimization ​

Author: Shirin Hosseinmardi, Xiangyu Sun, Ramin Bostanabad
Published: 8/14/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.CE, cs.LG

arXiv:2608.10344v2 Announce Type: replace-cross Abstract: Thermo-mechanical compliant devices are commonly designed with small-strain linear elasticity and temperature-independent material properties, even though they might operate hundreds of kelvin above ambient where both assumptions are question...

📖 Read original article


265. Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information ​

Author: Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML

arXiv:2608.10766v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that...

📖 Read original article


266. CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation ​

Author: Haowei Lou, Hye-Young Paik, Dai Jia, Kai Li, Lina Yao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SD, cs.LG

arXiv:2608.11590v2 Announce Type: replace-cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, co...

📖 Read original article


267. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation ​

Author: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.11616v2 Announce Type: replace-cross Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We ...

📖 Read original article


268. Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp ​

Author: Wenzhi Gao, Zhaonan Qu, Yinyu Ye, Madeleine Udell
Published: 8/14/2026, 4:00:00 AM
Categories: math.OC, cs.LG, stat.ML

arXiv:2608.11760v2 Announce Type: replace-cross Abstract: We revisit the Sinkhorn-Knopp (SK) algorithm for the matrix scaling problem. Despite extensive literature on the global convergence of SK and its variants, its local linear convergence behavior remains less understood. We address this gap by ...

📖 Read original article