Skip to content

arXiv cs.LG - 2026-07-27 ​

230 items collected.


1. Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees ​

Author: Lei Yang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detec...

📖 Read original article


2. On the Depth Scalability of Logic Gate Networks ​

Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO

arXiv:2607.21633v1 Announce Type: new Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in de...

📖 Read original article


3. MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion ​

Author: Tasfia Nuzhat Ornee, Elias Hossain, Ivan Garibay, Niloofar Yousef
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21634v1 Announce Type: new Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence representation, implicitly treating structurally heterogeneous molecular components as equally diffic...

📖 Read original article


4. Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions ​

Author: Pin Qian, Su Wang, Yihang Chen, Qiaolin Yu, Xiaoyuan Wang, Zhitong Guo, Zhicheng Wang, Junxian You
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benc...

📖 Read original article


5. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​

Author: Jie Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21636v1 Announce Type: new Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and c...

📖 Read original article


6. Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning ​

Author: Julian G. Soltes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.21637v1 Announce Type: new Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and aggregate an optimal prior ...

📖 Read original article


7. Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations ​

Author: Jonathan Gallagher, Roberto Guglielmi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.21644v1 Announce Type: new Abstract: We present a goal-agnostic control framework for partial differential equations (PDEs) built around a joint-embedding predictive architecture (JEPA). The small 2D ViT encoder and action-conditioned latent dynamics are trained offline without a reward o...

📖 Read original article


8. Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not ​

Author: Kavya Bhand, Aadi Joshi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21645v1 Announce Type: new Abstract: Multi-horizon latent consistency is a common training knob in video predictors and world models, but practitioners rarely know what it does to transition geometry. We treat lambda, the weight on multi-step latent agreement, as a diagnostic control and ...

📖 Read original article


9. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning ​

Author: Timothy Tomashevskiy
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement learning methods typica...

📖 Read original article


10. A Drift Stable Quantum Federated Learning for Intelligent Services ​

Author: Shanika Iroshi Nanayakkara, Shiva Raj Pokhrel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21647v1 Announce Type: new Abstract: Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distributed ...

📖 Read original article


11. Shallower ReLU Network Representations via Exact Linear Algebra ​

Author: Kilian Rue{\ss}, Gennadiy Averkov, Florestan Brunck, Moritz Grillo, Christoph Hertrich, Georg Loho, Jack Stade, Moritz Stargalla, Matthew Sun, Martin Winter
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NE, math.CO

arXiv:2607.21651v1 Announce Type: new Abstract: We prove that the maximum of $n$ real numbers is exactly representable by a ReLU network with two hidden layers for every $n\le 10$. The constructions are obtained by reducing the problem to exact rational linear algebra: after a symmetry reduction, th...

📖 Read original article


12. Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning ​

Author: Jian Hu, Huiying Li, Hao Zhang, Binfeng Xu, Yifan Zhang, Shaokun Zhang, Hemil Desai, Michael Demoret, Pavlo Molchanov, Jan Kautz, Yi Dong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, cs.DC

arXiv:2607.21653v1 Announce Type: new Abstract: Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the...

📖 Read original article


13. Physically Constrained Federated Additive Models for O-RAN SLA-Risk Prediction ​

Author: Aubida A. Al-Hameed, Mohammed M. H. Qazzaz, Maryam Hafeez, Syed A. Zaidi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.21665v1 Announce Type: new Abstract: Proactive service assurance in O-RAN requires predicting per-slice SLA violations before they occur. The prediction model must be auditable by operators and must train across base stations without pooling per-slice KPIs, which are commercially sensitiv...

📖 Read original article


14. Neural Feature Governance: Extending Atom Prevalence ​

Author: Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait Fokou'e
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21671v1 Announce Type: new Abstract: Neural network compression and interpretability remain open challenges in modern deep learn- ing, where billion-parameter architectures deliver impressive accuracy at the cost of trans- parency, computational efficiency, and reliable uncertainty quanti...

📖 Read original article


15. Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration ​

Author: Vishnu Bindu Balachandran
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.CV, stat.ML

arXiv:2607.21673v1 Announce Type: new Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynamical law. Modelling bank impurity as a generalized P'olya urn, we prove almost-sure convergence to a ...

📖 Read original article


16. Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA ​

Author: Takato Yasuno
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21680v1 Announce Type: new Abstract: Bridge infrastructure deteriorates gradually, yet its root causes---salt intrusion, freezing, fatigue cracking, and others---remain invisible to the naked eye. Expert diagnosis relies on tacit knowledge built over years of practice. We address the chal...

📖 Read original article


17. CARNet Cycle-Conditioned Core Aggregation and Redistribution for Multivariate Time Series Forecasting ​

Author: Awsaf Tausif Adib, Md. Shahria Sarker Shuvo, Md. Estehaar Ahmed Emon, Mustafa Kamal, Fuad Rahman, Shafin Rahman, Nabeel Mohammed
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21681v1 Announce Type: new Abstract: Accurately modeling cross-variate dependencies remains a key challenge in multivariate time series forecasting, particularly in the presence of strong periodic patterns. Many existing approaches rely on attention-based mechanisms that incur quadratic c...

📖 Read original article


18. Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets ​

Author: Jim Allchin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that attention reveals which context...

📖 Read original article


19. An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning ​

Author: Maximilian Dax, Theo Heimel, Gilles Louppe
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, astro-ph.CO, astro-ph.GA, hep-ex, hep-ph, stat.ML

arXiv:2607.21702v1 Announce Type: new Abstract: Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects. We provide an overview of the Bayesia...

📖 Read original article


20. A Defense of the Quadratic Model ​

Author: Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML

arXiv:2607.21716v1 Announce Type: new Abstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimizatio...

📖 Read original article


21. RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory ​

Author: Zahra Yousefijamarani, Alaa Alameldeen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AR

arXiv:2607.21731v1 Announce Type: new Abstract: Transformers are widely used across many domains, including natural language processing, computer vision, web search, and DNA sequence analysis. Given their broad applicability, improving the performance of transformer models is critical. However, the ...

📖 Read original article


22. Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection ​

Author: Debarshi Kundu, Swaroop Ghosh, Vasant Honavar
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21752v1 Announce Type: new Abstract: Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive approaches---including SBM-Transformer, Dynamic Mask Attention, and NSA...

📖 Read original article


23. Smart predict-then-robustly-optimize ​

Author: Aakil Caunhye, Xuefei Lu, Belen Martin-Barragan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.21773v1 Announce Type: new Abstract: In this paper, we propose and study a robust variant of the smart predict-then-optimize approach that accounts for prediction shifts due to disturbance in the covariate feature space. While traditional integrated-learning-and-optimization models assume...

📖 Read original article


24. Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection ​

Author: Othmane Harraq, Tamer Aldwairi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CR, cs.CV, cs.MM

arXiv:2607.21776v1 Announce Type: new Abstract: Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify. Unlike face-swap ma- nipulation, TF ...

📖 Read original article


25. Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness ​

Author: Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21806v1 Announce Type: new Abstract: Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice. There is a growing recognition of the need to evaluate the causal impact of deploying t...

📖 Read original article


26. Data eccentricity, asymptotics of Gaussian RBF reproducing kernel Hilbert space, and kernel PCA ​

Author: Sergio A. Alvarez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21823v1 Announce Type: new Abstract: We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric proper...

📖 Read original article


27. A Graph-Based Control Interface for Traffic Signals on Heterogeneous Road Networks ​

Author: Bertil Braun
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.21831v1 Announce Type: new Abstract: We present a traffic-signal control interface in which a shared graph neural network assigns scores to individual traffic movements. Each junction converts these scores into its own variable-sized set of legal signal phases using a deterministic incide...

📖 Read original article


28. Searching the Space of Feed-Forward Neural-Network Weight-Update Rules with Fixed Depth Symbolic Regression ​

Author: Charles Brum, Edward Finkelstein
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21855v1 Announce Type: new Abstract: We investigate whether symbolic regression can discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks. Candidate update rules are represented as fixed-depth symboli...

📖 Read original article


29. LeAct: Learning to Reason from Expert Actions ​

Author: Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem p...

📖 Read original article


30. Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study ​

Author: Kaihua Ding
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2607.21866v1 Announce Type: new Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 students each ra...

📖 Read original article


31. Variance-Reduced Q-Learning over Static and Time-Varying Networks ​

Author: Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath Jr
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.21876v1 Announce Type: new Abstract: We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value f...

📖 Read original article


32. Remedying Coarsening-Based GNN Training under Heterophily via Adaptive Complementary Enhancement ​

Author: Guoming Li, Jian Yang, Xukun Wang, Zixiao Wang, Shangsong Liang, Yifan Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, eess.SP, math.NA

arXiv:2607.21885v1 Announce Type: new Abstract: Coarsening-based training for graph neural networks (GNNs), i.e.\ training on coarsened graphs rather than the original large ones, has become a promising direction for scaling GNNs to massive graphs. However, prior work has been evaluated almost exclu...

📖 Read original article


33. MissHyper: Restoring Clinical Synchronicity in Missingness-Guided Hypergraph Forecasting ​

Author: Mingyi Ma, Qingxiong Tan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21922v1 Announce Type: new Abstract: Clinical irregular multivariate time series are shaped not only by physiological dynamics but also by the measurement process that determines when and what to observe. In event-centric models, however, co-timestamp structure can be flattened too early:...

📖 Read original article


34. RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention ​

Author: Anderson R. Santos
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this constraint as a model-a...

📖 Read original article


35. LatentFlow: Visual Analytics for Latent Space Analysis in Molecular Graph Neural Networks ​

Author: Shiyi Liu, Jiaqing Chen, Nicholas Hadler, Rostyslav Hnatyshyn, Michael W. Mahoney, Talita Perciano, John F. Hartwig, Gunther H. Weber, Ross Maciejewski
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.HC

arXiv:2607.21941v1 Announce Type: new Abstract: Chemists and materials scientists increasingly use machine learning models, such as graph neural networks (GNNs), to predict properties of molecules and the outcomes of their reactions. Beyond predictive performance, understanding how these models orga...

📖 Read original article


36. Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge ​

Author: Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21946v1 Announce Type: new Abstract: This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image. Visual ph...

📖 Read original article


37. MA-DAR: Manifold-Aligned Dynamic Adaptive Routing for Continual Temporal Knowledge Graph Reasoning ​

Author: Xiangjun Shi, Chong Mu, Jinchuan Zhang, Lizong Zhang, Yuefeng He, Shang Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21949v1 Announce Type: new Abstract: Continual temporal knowledge graph (TKG) reasoning aims to continuously incorporate newly emerging facts while preserving previously acquired knowledge. Replay-based continual learning has achieved promising performance by revisiting historical represe...

📖 Read original article


38. Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ​

Author: Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.21971v1 Announce Type: new Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflectio...

📖 Read original article


39. On the Convergence of Stochastic Low-Rank Adaptation ​

Author: Ru Wang, Chengchang Liu, John C. S. Lui
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.21975v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \m...

📖 Read original article


40. From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning ​

Author: Jiaxin Deng, Junbiao Pang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21999v1 Announce Type: new Abstract: Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape. Conventional re-sampling addresses the former without con...

📖 Read original article


41. Energy Manifold Natural Gradient Descent: Riemannian Optimization for Neural PDE Solvers ​

Author: Zhangyong Liang, Huanhuan Gao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain. We introduce \EMNGDfull{}, a manifold optimization ...

📖 Read original article


42. Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits ​

Author: Yuta Natsubori, Masataka Ushiku, Yuta Saito
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22012v1 Announce Type: new Abstract: Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handl...

📖 Read original article


Author: Xiafeng Man
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common styli...

📖 Read original article


44. CEL: Comprehensive Counterfactual Explanations Library and Benchmark ​

Author: Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minim...

📖 Read original article


45. A Leakage-Free Stacked Ensemble Method for Multiclass Classification ​

Author: S. P. Sharmila, Aruna Tiwari
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22081v1 Announce Type: new Abstract: Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as...

📖 Read original article


46. Pretraining EHR Foundation Models with Patient-Aware Sampling ​

Author: Joshua Placidi, Yuxuan Liu, Jinpei Han, Marek Rei, A. Aldo Faisal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22114v1 Announce Type: new Abstract: Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into a single token stream and windows are sampled from that stream. In EHR ...

📖 Read original article


47. TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex ​

Author: Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22143v1 Announce Type: new Abstract: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of molec...

📖 Read original article


48. Unbiased Open World Regularization for Fair Self-Supervised Learning ​

Author: L{'e}o Nicollier (CB, ATT), Marc Pic (ATT), Pablo Mus{'e} (CB, IFUMI), Enric Meinhardt-Llopis (CB), Gabriele Facciolo (CB)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset. These techniques rely on regularization, which prevents representation co...

📖 Read original article


49. From Score Approximation to Distribution Approximation in Score-Based Diffusion Models ​

Author: Lan V. Truong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.IT, math.IT, stat.ML

arXiv:2607.22199v1 Announce Type: new Abstract: Score-based diffusion models have achieved remarkable empirical success in generative modeling, yet their approximation-theoretic foundations remain incomplete. In particular, although classical universal approximation theorems guarantee that neural ne...

📖 Read original article


50. Latent PDE mapping for efficient physics-informed learning across geometries with limited data ​

Author: Ingvild Askim Adde, Mary M. Maleckar, Gabriel Balaban
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph

arXiv:2607.22215v1 Announce Type: new Abstract: In this study, we introduce latent PDE mapping, a broadly applicable physics-informed learning technique designed to enable efficient geometric generalization with sparse training data. Latent PDE mapping pulls back geometry-specific PDE residuals and ...

📖 Read original article


51. Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression ​

Author: Hirotaka Sugawara, Yujin Taguchi, Kei Minagawa, Yusuke Hiki, Takashi Morikura, Akira Funahashi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22238v1 Announce Type: new Abstract: Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables by balancing exploration and exploitation. BO is often used under a limited evaluation budget, such...

📖 Read original article


52. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning ​

Author: Wei Zhang, Xinwu Liu, Yihang Cheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22251v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method for large language models, but its performance depends strongly on how a fixed rank budget is distributed across Transformer modules. Existing adaptive-rank methods usua...

📖 Read original article


53. Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed Recognition ​

Author: Yi-Hang Zhu, Rajeev Raman, Shiqi Su, Jianyuan Sun, Xinyu Yang, Nan Xing, Huiyu Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2607.22258v1 Announce Type: new Abstract: Deep learning models using traditional softmax classifiers have achieved remarkable success in various classification tasks. However, their performance degrades significantly on imbalanced datasets. Although Balanced Softmax is widely adopted as a stat...

📖 Read original article


54. Autoregressive EHR Foundation Models with Multimodal Inputs ​

Author: Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way. We prese...

📖 Read original article


55. An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection ​

Author: Romain Hermary, Nesryne Mejri, Djamila Aouada
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2607.22286v1 Announce Type: new Abstract: Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. Although metrics such as AUROC, AUPR, F1-score, and MCC are widely used, their values convey different meanings depend...

📖 Read original article


56. Efficient Recommendations via Graph Coarsening and Label Propagation ​

Author: Alessandro Sbandi, Federico Siciliano, Fabrizio Silvestri
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22287v1 Announce Type: new Abstract: Graph-based recommendations are widely adopted in real-world industrial applications. However, graphs in these systems often reach a massive scale, posing notable scalability and efficiency challenges. This requires techniques that can effectively bala...

📖 Read original article


57. Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning ​

Author: Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez, Cole Looney, Xiaoliang Wu, Alexandra Livia Georgescu, Stefano Goria
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SD

arXiv:2607.22304v1 Announce Type: new Abstract: Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice c...

📖 Read original article


58. Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs ​

Author: Zhangzhi Xiong, Minzhang Li, Haotian Yu, Sixian Shen, Kexin Zhang, Mingrui Li, Jie Zheng, Kewei Tu, Jingyi Yu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22314v1 Announce Type: new Abstract: Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based fil...

📖 Read original article


59. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization ​

Author: Hao Wang, Kun Yuan, Wenlin Zhong, Minglei Zhang, Han Xiao, Ming Sun, Honggang Qi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.22334v1 Announce Type: new Abstract: Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while...

📖 Read original article


60. Beyond Binary Rooftop Mapping: A Four-Class Deep Learning Framework for Green Roof Potential Assessment from Open Swiss Geospatial Data ​

Author: Htet Yamin Ko Ko
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22342v1 Announce Type: new Abstract: The development of effective urban climate adaptation strategies requires comprehensive spatial information on rooftops and buildings, since such information underpins the assessment of ecosystem services provided by green infrastructure, particularly ...

📖 Read original article


61. IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data ​

Author: Masashi Sode, Gianmarco Pinton
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.med-ph

arXiv:2607.22351v1 Announce Type: new Abstract: The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated soun...

📖 Read original article


62. Integrated Order Dispatching and Routing for Last-Mile Pickup via Deep Reinforcement Learning ​

Author: Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22356v1 Announce Type: new Abstract: In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms. This challenge is fundamentally driven by two key and tightly coupled decision-making processes:...

📖 Read original article


63. Indexing: the Beginning and the End ​

Author: Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22361v1 Announce Type: new Abstract: We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits ...

📖 Read original article


64. Interior interpretability with attention rollout: contraction and propagation profiles in Transformers ​

Author: Umberto Biccari, Qian Huang, Enrique Zuazua
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretabil...

📖 Read original article


65. Local-Global Geometric Insights for Graph Neural Networks via Entropic Curvature ​

Author: Rachid Caich, Yassine Abbahaddou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SI, stat.ML

arXiv:2607.22381v1 Announce Type: new Abstract: Curvature notions on graphs, particularly Ollivier-Ricci and Forman, have emerged as powerful tools for addressing fundamental issues in Graph Neural Networks (GNNs) such as oversmoothing and oversquashing, but rely almost exclusively on local edge-lev...

📖 Read original article


66. LunarFM: A Shared Multimodal Representation of the Moon's Surface ​

Author: Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, Abigail Calzada-Diaz, Sylvester Kaczmarek, Ra'ul Ramos-Poll'an
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22408v1 Announce Type: new Abstract: The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast ...

📖 Read original article


67. On the Identifiability of Controlled World Models ​

Author: Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework fo...

📖 Read original article


68. Hyperball May Not Be a Free Lunch ​

Author: Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Startin...

📖 Read original article


69. Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining ​

Author: V'ictor Rinc'on Yepes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22458v1 Announce Type: new Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from ...

📖 Read original article


70. Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates ​

Author: Anjian Li, Ryne Beeson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22467v1 Announce Type: new Abstract: Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve. We therefore study a $k$-neighborhood data collection strategy ...

📖 Read original article


71. Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent ​

Author: Peng Zhao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.ML, stat.TH

arXiv:2607.22474v1 Announce Type: new Abstract: In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structur...

📖 Read original article


72. \k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating ​

Author: Jianghui Wang, Silong Yong, Francesco Orabona, Marco Canini, Katia P. Sycara, Yaqi Xie
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22489v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardl...

📖 Read original article


73. Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting ​

Author: Aliaksei Kaliutau
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22491v1 Announce Type: new Abstract: Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, and it...

📖 Read original article


74. Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes ​

Author: Athanasios Papastathopoulos-Katsaros, Steven T. Lee, Lin Yao, Ajay Thomas, Junseok Park, Matthew J. McGinley, Zhandong Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, eess.SP

arXiv:2607.22508v1 Announce Type: new Abstract: Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks. Predefined features carry a strong bias, since they fix in advance what cou...

📖 Read original article


75. Dysphagia Risk Stratification in Head and Neck Cancer via Two-Stage PRO-Clinical Stacking ​

Author: Siyuan Zhao, Eric Ababio Anyimadu, Zachary G. Brumm, Yue Ma, Clifton David Fuller, Xinhua Zhang, G. Elisabeta Marai, Guadalupe Canahuate
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.22514v1 Announce Type: new Abstract: Dysphagia is a debilitating late effect of head and neck cancer (HNC) treatment, yet timely identification of at-risk patients remains challenging in survivorship care. Definitive assessment relies on videofluoroscopic imaging, as captured by the Dynam...

📖 Read original article


76. An Explainable FFT-Based Spatial-Frequency Fusion Framework for Deepfake Detection ​

Author: Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.17441v1 Announce Type: cross Abstract: Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust. Recent studies show that combining spatial and frequency features leads to stronger detection results than using i...

📖 Read original article


77. Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification ​

Author: Guillermo Rubi~nos Rodr'iguez, Mart'in Ottavianelli, Mateo Alonso, Gonzalo Bl'azquez Gil, Boris-Stephan Rauchmann, Pablo D'iez-Valle, Sergio Altares-L'opez
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2607.21186v1 Announce Type: cross Abstract: Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, while the scalabil...

📖 Read original article


78. Spectral Flow Certificates for Depth-Aware Long-Range Propagation in Graph Neural Networks ​

Author: Ranjan Veerabhadraswamy, Ajith Jubilson Emerson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.21607v1 Announce Type: cross Abstract: Graph Neural Networks propagate information through local message passing, but the graph topologies themselves can silently prevent any amount of training from solving long-range tasks. When we deploy GNNs on new graphs, there is currently no inexpen...

📖 Read original article


79. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text ​

Author: Miaobo Hu, Xiaobo Guo, Shuhao Hu, Bokun Wang, Rui Chen, Xin Wang, Daren Zha, Jun Xiao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.21610v1 Announce Type: cross Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pi...

📖 Read original article


80. Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures ​

Author: Simon Dennis, Kevin Shabahang, Hao Guo, Rivaan Patil
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.21612v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following, style transfer, and factual adaptation. We show that for procedural knowledge--the ability to follo...

📖 Read original article


81. FrED: External Data Influence Estimation via Domain Knowledge Graph Grounding ​

Author: Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.21615v1 Announce Type: cross Abstract: The rapid deployment of generative AI has amplified the critical need for Training Data Attribution to ensure transparency and accountability. However, current parametric approaches require computationally prohibitive access to model weights, while s...

📖 Read original article


82. Local Synaptic Rules Can Implement a SIGReg Gradient Without Backpropagation ​

Author: Martin Andrews
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV, cs.LG

arXiv:2607.21622v1 Announce Type: cross Abstract: We prove that two canonical local synaptic learning rules, the potentiation arm of spike-timing-dependent plasticity (STDP$^+$) and homeostatic plasticity (instantiated here via flashlight granule-cell-like neurons), together can implement the exact ...

📖 Read original article


83. FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs ​

Author: Kahou Tam, Wei Niu, Yu Bao, Xiaomin Ouyang, Chengzhong Xu, Li Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG

arXiv:2607.21624v1 Announce Type: cross Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs du...

📖 Read original article


84. Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems ​

Author: Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.21627v1 Announce Type: cross Abstract: End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure mode in which modules preserve or improve end-task performance while ...

📖 Read original article


85. Computer Vision Based Neurology Brain Activity Rejection Architecture and Implementation ​

Author: Zag ElSayed, Nathan Suer, Grace Westerkamp, Jack Yanchen Liu, Makoto Miyakoshi, Craig Erickson, Ernest Pedapati
Published: 7/27/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG, physics.data-an, q-bio.QM

arXiv:2607.21654v1 Announce Type: cross Abstract: The electroencephalogram (EEG) is a valuable and widely applied tool for investigating brain disorders and behavioral changes. It offers a minimally restrictive and non-invasive method. However, challenges in using EEG for cognitive development studi...

📖 Read original article


86. Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives ​

Author: Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, Nicola A. Morley, Robert D. J. Oliver, Alexandra J. Ramadan, Delvin Ce Zhang, Katerina A. Christofidou, Haiping Lu
Published: 7/27/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.LG

arXiv:2607.21660v1 Announce Type: cross Abstract: Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompass...

📖 Read original article


87. Ordered Action Tokens for Visuomotor Policy Learning ​

Author: Chaoqi Liu, Yue Zhao, Haonan Chen, Xiaoshen Han, Jiawei Gao, Ehsan Adeli, Yilun Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long toke...

📖 Read original article


88. Explainable quantum-compressed machine learning for complex fluid flows ​

Author: Xiao Xue, Maida Wang, Mingyang Gao, Minh Chung, Peter V. Coveney
Published: 7/27/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.LG, quant-ph

arXiv:2607.21688v1 Announce Type: cross Abstract: Machine-learning surrogates of physical systems face a paradox: explainable models facing the challenge of expressivity to capture complex nonlinear flows, whereas expressive deep surrogates match high-fidelity simulations only through massive parame...

📖 Read original article


89. Prior laundering: learned priors with inherited, undetectable overconfidence ​

Author: Ali Siahkoohi, Sina Alemohammad
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.21721v1 Announce Type: cross Abstract: Learned generative priors are increasingly used for ill-posed Bayesian inverse problems, their posterior uncertainty treated as earned from data. But training one requires truths, scarce in seismic and medical imaging, so the recourse is an archive o...

📖 Read original article


90. Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery ​

Author: Khalid El-Darymli, Christoph H. Gierull, Katerina Biron, Weimin Huang
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2607.21745v1 Announce Type: cross Abstract: Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study introduces a deep sigma-point process (DSPP) model for predicting RCS in synthetic aperture radar (SAR) imagery using ...

📖 Read original article


91. Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting ​

Author: Denis Mayr Lima Martins, Gottfried Vossen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.LG

arXiv:2607.21756v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in database-backed applications to classify tuples, filter records using semantic predicates, extract structured attributes, and enrich query results. Yet the prompt that start these computations are...

📖 Read original article


92. Encoding orders and trees in real-valued functions ​

Author: G Conant, C Terry
Published: 7/27/2026, 4:00:00 AM
Categories: math.CO, cs.LG, math.LO

arXiv:2607.21761v1 Announce Type: cross Abstract: We prove function-theoretic analogues of a quantitative result of Hodges on extracting the order property from a sufficiently large 2-tree coded in a binary relation. Similar analogues for functions were previously obtained by Daskalakis and Golowich...

📖 Read original article


93. Reliability-Aware Bayesian Optimization of 1310 nm PCSELs with FDTD Verification ​

Author: Jinglin Yu, Feiyang Wu, Longying Wen, Chongxian Yuan, Renjie Li, Zhaoyu Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: physics.optics, cs.LG, physics.app-ph

arXiv:2607.21772v1 Announce Type: cross Abstract: Near 1310 nm photonic-crystal surface-emitting lasers (PCSELs) are attractive narrow-beam sources for optical communication and sensing, but their final design refinement is costly. Small geometry changes simultaneously shift the band-edge resonance,...

📖 Read original article


94. From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia ​

Author: Rodrigo Vargas Sainz, Christian Ber'on Curti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.21785v1 Announce Type: cross Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local predictive systems to...

📖 Read original article


95. Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling ​

Author: William Gaudelier, Albert Cohen, Dumitru Potop Butucaru
Published: 7/27/2026, 4:00:00 AM
Categories: cs.PL, cs.LG

arXiv:2607.21797v1 Announce Type: cross Abstract: Previous work has shown that the simple dataflow primitives of the Lustre language allow the natural, semantically unambiguous, and compact representation of machine learning (ML) applications, including models featuring complex conditional execution...

📖 Read original article


96. Adversarial Prompts for Acceptance Collapse in Speculative Decoding ​

Author: Run Wang, Chaoyi Zhou, Xi Liu, Yi Zhu, Amir Salarpour, Pedram MohajerAnsari, Zhi-Qi Cheng, Feng Luo, Siyu Huang, Mert D. Pes'e
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.CL, cs.LG

arXiv:2607.21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operati...

📖 Read original article


97. Natural Invariant Measures for Chaotic Game Dynamics: Finding Order in Chaos ​

Author: Jakub Bielawski, Thiparat Chotibut, Fryderyk Falniowski, Micha{\l} Misiurewicz, Georgios Piliouras
Published: 7/27/2026, 4:00:00 AM
Categories: math.DS, cs.LG, econ.TH

arXiv:2607.21805v1 Announce Type: cross Abstract: We study the long-term behavior of the Multiplicative Weights Update (MWU) algorithm in game settings where learning dynamics frequently fail to converge to Nash equilibria and instead exhibit Li-Yorke chaos. While such chaos precludes the prediction...

📖 Read original article


98. Longitudinal Random Forests for Sparse and Irregular Response Trajectories ​

Author: Yangsheng Wang, Xiaotian Dai, Haoda Fu, Guifang Fu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, stat.ML

arXiv:2607.21817v1 Announce Type: cross Abstract: Longitudinal studies often collect data at sparse, irregular, and unequally spaced time points. Such heterogeneity is often driven by subject-specific covariates, yet existing methods have been restricted to a scalar endpoint value, completely neglec...

📖 Read original article


99. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors ​

Author: Daniyal Kabir Dar, Arun Ross
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR, cs.LG

arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold w...

📖 Read original article


100. How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests ​

Author: Iren Mazloomzadeh, Mohammad Mehdi Morovati, Foutse Khomh
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.LG

arXiv:2607.21832v1 Announce Type: cross Abstract: Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows. While developers increasingly ben...

📖 Read original article


101. Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification ​

Author: Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2607.21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data. This makes them especially attractive for auditing models deployed...

📖 Read original article


102. Quantifying Political Partisanship for Cross-Platform Analyses ​

Author: Fathima Ameen, Christopher G. Healey
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SI, cs.LG

arXiv:2607.21842v1 Announce Type: cross Abstract: Research on political polarization on social media depends on the ability to reliably measure partisanship in user-generated content. However, existing approaches are typically tailored to platform-specific properties, such as structural affordances ...

📖 Read original article


103. Simulation-Based Empirical Bayes ​

Author: Xinwei Shen, Diana Cai, Cheng Zhang, David M. Blei
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.21843v1 Announce Type: cross Abstract: Empirical Bayes (EB) performs simultaneous inference across many related latent variables. Classical EB assumes that the likelihood p(x | z) is tractable. In many scientific applications, however, the likelihood is available only through a simulator....

📖 Read original article


104. Distributional Determinantal Point Process for Repulsive Clustering of Distributions ​

Author: Khai Nguyen, Yang Ni, Elizabeth Juarez-Colunga, Peter Mueller
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, stat.AP, stat.CO, stat.ML

arXiv:2607.21847v1 Announce Type: cross Abstract: We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (S...

📖 Read original article


105. Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 Refinement ​

Author: Mohammadreza Narimani, Vikram Anand, Parastoo Farajpoor
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, eess.IV

arXiv:2607.21881v1 Announce Type: cross Abstract: Agricultural field maps are often proprietary, incomplete, or outdated, yet they provide the spatial framework for crop monitoring, production accounting, and land-conversion analysis. This study presents a reproducible workflow for mapping farmland ...

📖 Read original article


106. Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes ​

Author: Lu Luo, Dandan Mo, Chengdong Xu, Ting Li, Jinhan Xie, Huiqiong Li, Niansheng Tang
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.21958v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential. Statistical watermarking has emerged as a promising solution, yet most...

📖 Read original article


107. Unified Static-Dynamic Pruning for Efficient LLM Inference ​

Author: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.AR, cs.LG

arXiv:2607.21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offers a prom...

📖 Read original article


108. QC-PHAST Search: Classical--Quantum Query Benchmarks for Finite-Pool Rare-Regime Discovery ​

Author: Harsh Milind Tirhekar, Chandrajit Bajaj
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.ET, cs.LG

arXiv:2607.21995v1 Announce Type: cross Abstract: Rare-regime discovery in parameterized dynamical systems is an active-search problem: find one verified parameter at which a scientifically defined qualitative threshold is crossed, even when acceptable candidates are rare, nonconvex, or fragmented. ...

📖 Read original article


109. Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation ​

Author: M M Asif Ferdous
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.CL, cs.LG

arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw accuracy, because it d...

📖 Read original article


110. Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era ​

Author: Yu Wang, Hongyu Yang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations. In this work, we revisit this assumption in the foundation-model ...

📖 Read original article


111. MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection ​

Author: Phurich Saengthong, Takahiro Shinozaki
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.LG

arXiv:2607.22086v1 Announce Type: cross Abstract: Autoencoder-based anomalous sound detection is attractive for machine condition monitoring because it can be trained using only normal recordings and yields an interpretable anomaly score from reconstruction error. Most prior work uses spectrogram au...

📖 Read original article


112. Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models ​

Author: Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22098v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps...

📖 Read original article


113. One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments ​

Author: Jan Ole von Hartz, Abhinav Valada, Joschka Boedecker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames. However, existing approaches typically assume these frames to be strictly exogenous. Th...

📖 Read original article


114. CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography ​

Author: Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo'nski, Tomasz Figatowski, Natalia Zieli'nska
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols. In this work we demonstrate a new benchmark for the assessment of deep learning model...

📖 Read original article


115. Industrial Tokenization for LLM-Based Health Intelligence: A Federated Architecture for Industrial Evidence Integration ​

Author: Deshui Li, Xiao-Ming Yuan, Zishun Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22153v1 Announce Type: cross Abstract: Industrial health management increasingly relies on heterogeneous information sources, including condition monitoring systems, supervisory control and data acquisition systems, maintenance records, inspection results, and prognostic models. Although ...

📖 Read original article


116. DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents ​

Author: Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.LG

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a runnin...

📖 Read original article


117. Bowel Obstruction Detection and Localization on Abdominal CT with Deep Learning ​

Author: Moritz Vandenhirtz, Andrea Agostini, Dana Belde, M'elanie Roschewitz, Ismaiel Chikh Bakri, Tilo Niemann, Andr'e Euler, Julia E Vogt
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.22173v1 Announce Type: cross Abstract: Bowel obstruction is a common and potentially life-threatening gastrointestinal condition. In the face of rising diagnostic workloads, the automated diagnosis of bowel obstruction on CT scans supports radiologists by accelerating detection and improv...

📖 Read original article


118. Trajectory-Regularized Stochastic Optimal Control via KL Divergence ​

Author: Mintae Kim, Koushil Sreenath
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY

arXiv:2607.22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theo...

📖 Read original article


119. Deep Convolutional Large-Margin $\ell_p$-SVDD for Visual Anomaly Detection ​

Author: Alireza Dastmalchi Saei, Shervin Rahimzadeh Arashloo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.22212v1 Announce Type: cross Abstract: Visual anomaly detection requires adaptive representations and reliable decision boundaries, particularly when anomalous training samples are scarce and class distributions are highly imbalanced. Classical kernel-based methods yield principled geomet...

📖 Read original article


120. Convergence analysis of a family of Zermelo-type iterations for the Bradley--Terry model ​

Author: Ruijian Han, Ding Lu, Yiming Xu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA, stat.CO

arXiv:2607.22221v1 Announce Type: cross Abstract: Zermelo's algorithm is a classical method for computing the maximum likelihood estimator in the Bradley--Terry (BT) model, but its convergence can be slow in practice. To accelerate computation, Newman introduced a family of Zermelo-type fixed-point ...

📖 Read original article


121. Variational Low-rank Tensor Decomposition for Multisubject Spatiotemporal Data Analysis ​

Author: Laura M. Montaldo, Ricardo A. Borsoi, Sebastian Miron, Tulay Adali
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, eess.SP

arXiv:2607.22262v1 Announce Type: cross Abstract: Modeling shared and subject-specific structure in multisubject spatiotemporal data remains challenging, particularly in neuroimaging, where both spatial and temporal patterns exhibit rich variability across subjects. Existing matrix and tensor decomp...

📖 Read original article


122. Explicit Iteration Complexity of Exact Data-Driven Inverse Optimization for Integer Linear Programs ​

Author: Akira Kitaoka
Published: 7/27/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG, stat.ML

arXiv:2607.22263v1 Announce Type: cross Abstract: A data-driven inverse optimization problem (DDIOP) is the problem of estimating the objective-function parameters (weights) that explain observed optimal-solution data, and it arises in many applications, including integer linear programming (ILP). I...

📖 Read original article


123. General Value Functions for Remaining Useful Life and Failure-Mode Prediction ​

Author: Hao Yan, Ali Sarabi, Qing Zou, Boyang Xu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP

arXiv:2607.22268v1 Announce Type: cross Abstract: Remaining useful life (RUL) prediction and failure-mode classification are central tasks in predictive maintenance. Many data-driven pipelines use fixed-window supervised learning with complete terminal labels; such routes do not naturally encode the...

📖 Read original article


124. Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting ​

Author: Wan Zhang, Qinjie Lin, Chan Lee, Weijian Li, Han Liu, Kai Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.22299v1 Announce Type: cross Abstract: Forecasting multiple time-series with high-dimensional covariates presents a core challenge: unifying common temporal patterns while retaining meaningful series-specific information. We introduce Hopformer (Homogeneity-Pursuit Transformer), a two-sta...

📖 Read original article


125. Learning Bidirectional Causal Interactions with Heteroscedastic Neural Networks ​

Author: Masahiro Tanaka
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME

arXiv:2607.22313v1 Announce Type: cross Abstract: Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence. This paper proposes SEM-DNN, a heterosced...

📖 Read original article


126. Agentic Root Cause Analysis through Evidence-Grounded Reasoning ​

Author: Amaury Wei, Olga Fink
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22385v1 Announce Type: cross Abstract: Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existi...

📖 Read original article


127. HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding ​

Author: Chao Fang, Jun Yin, Man Shi, Marian Verhelst
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG

arXiv:2607.22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that ex...

📖 Read original article


128. Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation Models ​

Author: Sergey Musienko
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.LG, q-bio.NC

arXiv:2607.22397v1 Announce Type: cross Abstract: Frozen EEG encoders proliferate; per-model fine-tune defaults do not scale. We present Nimbus Personalizer: one contract encode to Bayesian head to BrainState (optional affine mid-tier) that sits on heterogeneous frozen trunks without a new personali...

📖 Read original article


129. Learning Ergodic Dynamical Systems from a Finite Trajectory ​

Author: Oleksii Kachaiev, Silvia Villa, Lorenzo Rosasco
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.22399v1 Announce Type: cross Abstract: We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system. More precisely, we study discrete-time autonomous stochastic systems defining time-homogeneous Markov processes. We first focus on estimati...

📖 Read original article


130. Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition ​

Author: Austin Rockman
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.IR, cs.LG

arXiv:2607.22413v1 Announce Type: cross Abstract: Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision. We present Reflect...

📖 Read original article


131. Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability ​

Author: Ahmed M. Abuzuraiq, Philippe Pasquier
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.MM

arXiv:2607.22428v1 Announce Type: cross Abstract: Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque ...

📖 Read original article


132. Graph-Based Correlation Matrix Generation: A Convex Optimization Approach ​

Author: Ali Fakhar (UGA), K{'e}vin Polisano (UGA), Ir{`e}ne Gannaz (G-SCOP_GROG, G-SCOP, Grenoble INP, UGA), Sophie Achard (STATIFY, LJK, UGA)
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.22436v1 Announce Type: cross Abstract: This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures. We propose a novel convex optimization framework in which an initial matrix is projected onto an elliptope under ...

📖 Read original article


133. TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI ​

Author: Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2607.22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic app...

📖 Read original article


134. Singular value soft-thresholding via the polar decomposition ​

Author: Stephen Becker
Published: 7/27/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, math.OC

arXiv:2607.22484v1 Announce Type: cross Abstract: Singular value soft-thresholding can be computed via a reduction to the matrix polar decomposition, which allows one to exploit GPU-friendly algorithms for computing the polar decomposition. Empirically, there is a significant speed-up on GPUs compar...

📖 Read original article


135. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference ​

Author: Jiyuan Tan, Vasilis Syrgkanis
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, econ.EM

arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers...

📖 Read original article


136. Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support ​

Author: Peiyong Wang, Udaya Parampalli, Casey R. Myers
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2607.22516v1 Announce Type: cross Abstract: A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through sp...

📖 Read original article


137. PinEqualizer: Full Funnel Content Exploration and Debiasing System at Pinterest ​

Author: Olafur Gudmundsson, Bo Zhao, Huayi Liao, Anna Kiyantseva, Sai Xiao, Heath Vinicombe, Mostafa Keikha, Luke DeLuccia, Zihao Chen, Junpeng Hou, Weijie Jiang, Bhawna Juneja, Andreanne Lemay, Wei-Ting Lin, Keyvan Moghadam, Jiaxing Qu, Zhiqing Rao, Zhihua Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.IR, cs.LG

arXiv:2607.22518v1 Announce Type: cross Abstract: In this paper, we propose a new solution for addressing the content cold-start problem in industry-scale search and recommender systems. Compared to prior approaches, we have made the following new contributions: 1) our solution spans the entire mult...

📖 Read original article


138. Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning ​

Author: Seyeon Kim, Joonhun Lee, Namhoon Cho, Sungjun Han, Wooseop Hwang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.PR, stat.ML

arXiv:2408.02295v4 Announce Type: replace Abstract: Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residuals induced by bootstrapping and exploration. We introduce a state-co...

📖 Read original article


139. Online Pricing and Allocation with Demand Learning and Fulfillment Cost ​

Author: Jianyu Xu, Xuan Wang, Yu-Xiang Wang, Jiashuo Jiang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML

arXiv:2501.18049v3 Announce Type: replace Abstract: We study online learning for a seller that jointly chooses per-period inventory positions and a uniform price, then fulfills realized demand through a downstream allocation. The main difficulty is not only demand learning: the price shifts demand a...

📖 Read original article


140. Carpe Diem: Critical Learning Period-Aware Contract-Based Incentives for Federated Learning ​

Author: Thanh Linh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Quoc-Viet Pham
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.GT

arXiv:2503.07869v4 Announce Type: replace Abstract: Critical learning periods (CLPs) in federated learning (FL) refer to early stages during which low-quality contributions (e.g., sparse training data availability) can permanently impair the performance of the global model. However, existing incenti...

📖 Read original article


141. Spatially-Enhanced Temporal Fusion Transformer: Interpretable Multi-Output Prediction for Parametric Dynamical Systems with Time-Varying Inputs ​

Author: Shuwen Sun, Lihong Feng, Peter Benner
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2505.00473v2 Announce Type: replace Abstract: We explore the promising performance of a transformer model in predicting outputs of parametric dynamical systems with external time-varying input signals. The outputs of such systems vary not only with physical parameters but also with external ti...

📖 Read original article


142. Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models ​

Author: Roseline Polle, Agnes Norbury, Alexandra Livia Georgescu, Nicholas Cummins, Stefano Goria
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2505.23378v3 Announce Type: replace Abstract: Speaker-dependent modelling can substantially improve performance in speech-based health monitoring applications. While mixed-effect models are commonly used for such speaker adaptation, they require computationally expensive retraining for each ne...

📖 Read original article


143. Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning ​

Author: Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2507.01551v3 Announce Type: replace Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, a...

📖 Read original article


144. A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data ​

Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2509.10517v3 Announce Type: replace Abstract: Machine learning can predict in-hospital mortality, but data privacy and the statistical heterogeneity of clinical data hamper its use. Federated Learning (FL) is privacy-preserving, yet its behavior under non-IID and imbalanced conditions needs sc...

📖 Read original article


145. LiMuon: Light and Fast Muon Optimizer for Large Models ​

Author: Feihu Huang, Yuning Luo, Songcan Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2509.14562v5 Announce Type: replace Abstract: Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention. More recently, the useful Muon optimizer is specifically designed for matrix-structured parameters of large model...

📖 Read original article


146. HD3C: Efficient Medical Data Classification for Edge Devices ​

Author: Jianglan Wei, Zhenyu Zhang, Pengcheng Wang, Mingjie Zeng, Zhigang Zeng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2509.14617v4 Announce Type: replace Abstract: Efficient medical data classification is essential for modern disease screening, particularly in resource-constrained environments where power budgets and computing capabilities are limited. We present HD3C, a lightweight classification framework d...

📖 Read original article


147. SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis ​

Author: Marie Brockschmidt, Maresa Schr"oder, Stefan Feuerriegel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.22352v3 Announce Type: replace Abstract: Survival analysis is a cornerstone of clinical research by modeling time-to-event outcomes such as metastasis, disease relapse, or patient death. Unlike standard tabular data, survival data often come with incomplete event information due to dropou...

📖 Read original article


148. Safe In-Context Reinforcement Learning ​

Author: Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2509.25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history. While ...

📖 Read original article


149. Correlating Cross-Iteration Noise for DP-SGD using Model Curvature ​

Author: Xin Gu, Yingtai Xiao, Guanlin He, Jiamu Bai, Daniel Kifer, Kiwan Maeng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2510.05416v3 Announce Type: replace Abstract: Differentially private stochastic gradient descent (DP-SGD) offers the promise of training deep learning models while mitigating many privacy risks. However, there is currently a large accuracy gap between DP-SGD and normal SGD training. This has r...

📖 Read original article


150. Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization ​

Author: Jinwoo Baek
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2510.21770v2 Announce Type: replace Abstract: Low-precision execution can induce substantial forward discrepancies in Transformers even for fixed weights and input, yet these discrepancies are usually monitored only at the output and lack a layer-wise theoretical account. We develop a first-or...

📖 Read original article


151. CorVS+: Correspondence-Driven Association of Video Trajectories and Sensors for Identity-Aware Person Localization in Warehouses ​

Author: Kazuma Kano, Yuki Mori, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, cs.RO

arXiv:2510.26369v2 Announce Type: replace Abstract: Logistics warehouses have struggled with labor shortages, but the inbound processes remain particularly human-powered. Worker location data is a key to higher productivity in such cases. Fixed cameras are a promising tool for localization, as they ...

📖 Read original article


152. AdamNX: An Adam improvement algorithm based on a novel exponential decay mechanism for the second-order moment estimate ​

Author: Meng Zhu, Quan Xiao, Weidong Min
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2511.13465v5 Announce Type: replace Abstract: This paper studies the exponential decay mechanism of the second-moment estimate in Adam. We propose AdamNX and a time-varying second-moment decay rate that gradually weakens the correction applied to the update scale. Under the assumptions used in...

📖 Read original article


153. gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points ​

Author: Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.PR

arXiv:2512.06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists. This is because the vast majority of existin...

📖 Read original article


154. Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning ​

Author: Pritthijit Nath, Sebastian Schemm, Henry Moss, Peter Haynes, Emily Shuckburgh, Mark J. Webb
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.ao-ph

arXiv:2601.04268v3 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that are weakly constrained and tuned offline, contributing to persistent biases that limit their ability...

📖 Read original article


155. Eluder dimension: localise it! ​

Author: Alireza Bakhtiari, Alex Ayoub, Samuel Robertson, David Janz, Csaba Szepesv'ari
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2601.09825v3 Announce Type: replace Abstract: We establish a lower bound on the eluder dimension of generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds. To address this, we introduce a localisation method for the elud...

📖 Read original article


156. Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach ​

Author: Samaresh Kumar Singh, Joyjit Roy, Chirag Agrawal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.ET

arXiv:2601.17303v2 Announce Type: replace Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce latency bottlenecks that sophisticated attackers can exploit to compromise an entire manufacturing ec...

📖 Read original article


157. Automatic Stability and Recovery for Neural Network Training ​

Author: Barak Or
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.17483v2 Announce Type: replace Abstract: Training modern neural networks is increasingly fragile, with rare but severe destabilizing updates often causing irreversible divergence or silent performance degradation. Existing optimization methods primarily rely on preventive mechanisms embed...

📖 Read original article


158. Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems ​

Author: Samaresh Kumar Singh, Joyjit Roy
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.SE

arXiv:2602.04120v4 Announce Type: replace Abstract: Though Explainable AI (XAI) has made significant advancements, its inclusion in edge and IoT systems is typically ad-hoc and inefficient. Most current methods are "coupled" in such a way that they generate explanations simultaneously with model inf...

📖 Read original article


159. Layer-wise LoRA fine-tuning: a similarity metric approach ​

Author: Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.05988v2 Announce Type: replace Abstract: Pre-training Large Language Models (LLMs) on web-scale datasets becomes fundamental for advancing general-purpose AI. In contrast, enhancing their predictive performance on downstream tasks typically involves adapting their knowledge through fine-t...

📖 Read original article


160. Heavy-Tailed Principal Component Analysis ​

Author: Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2603.11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise. While numero...

📖 Read original article


161. Hierarchical Latent Structure Learning through Online Inference ​

Author: Ines Aitsahalia, Kiyohito Iigaya
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, q-bio.NC

arXiv:2603.19139v2 Announce Type: replace Abstract: Learning systems must balance generalization across experiences with discrimination of task-relevant details. Effective learning therefore requires representations that support both. Online latent-cause models support incremental inference but assu...

📖 Read original article


162. LLM-Extracted Covariates for Clinical Causal Inference: Rethinking Integration Strategies ​

Author: Lei Liu, Jialin Chen, Kathy Macropol
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2604.16763v3 Announce Type: replace Abstract: Causal inference from electronic health records (EHR) is fundamentally limited by unmeasured confounding: critical clinical states such as frailty, goals of care, and mental status are documented in free-text notes but absent from structured data. ...

📖 Read original article


163. Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair ​

Author: Yuelin Hu, Zhenbo Yu, Zhengxue Cheng, Wei Liu, Li Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.22407v2 Announce Type: replace Abstract: Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition has a hidden failure mode. In a high-overlap, non-adaptive 8-domain co...

📖 Read original article


164. Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics ​

Author: Chao Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2605.11017v2 Announce Type: replace Abstract: Behavioral curve modeling -- fitting parametric functions to engagement-versus-exposure data -- is standard practice in recommendation, advertising, and clinical dosing. We show that aggregation introduces a systematic distortion: Simpson's paradox...

📖 Read original article


165. DriftXpress: Faster Drifting Models via Projected RKHS Fields ​

Author: Ali Falahati, Elliot Creager, Gautam Kamath, Shubhankar Mohapatra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.12183v2 Announce Type: replace Abstract: Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference. The premise is to replace the iterative denoising process in diffusion models with a single evaluation of a...

📖 Read original article


166. Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability ​

Author: Taewoon Kim, Vincent Fran\c{c}ois-Lavet, Michael Cochez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.22142v3 Announce Type: replace Abstract: Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not explicitly model short-term-to-long-term transfer of symbolic observations. We study this transfer process in a...

📖 Read original article


167. Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO ​

Author: Yiran Xu, Yiming Ren, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong Wang, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, Ruihang Chu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.30789v3 Announce Type: replace Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, w...

📖 Read original article


168. Pretraining Recurrent Networks without Recurrence ​

Author: Akarsh Kumar, Phillip Isola
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.06479v2 Announce Type: replace Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorly: it is sequential in time, limiting parallelism, and suffers from ...

📖 Read original article


169. PostDeg: Placement Beats Parameterization in LayerNorm GNNs ​

Author: Yash Tomar, Aryav Das
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.14022v2 Announce Type: replace Abstract: LayerNorm-based GNNs routinely erase the topology signals (degree, centrality, $k$-core) that node-selection policies should depend on, but the literature has not located where in the residual block the erasure happens. We answer that question: a p...

📖 Read original article


170. An Empirical Study of OpenPangu Quantization on Ascend NPUs ​

Author: Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.21257v3 Announce Type: replace Abstract: OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a contr...

📖 Read original article


171. A General Framework for Learning Algebraic Properties from Cayley Graphs using Graph Neural Networks ​

Author: Tal Weissblat
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.GR

arXiv:2606.26212v2 Announce Type: replace Abstract: In this work, we present a general Graph Neural Network (GNN) framework for learning algebraic properties of finite groups from their Cayley graph representations. The framework provides a unified computational pipeline consisting of a common graph...

📖 Read original article


172. Geometry-Conditioned Fourier Neural Operators for Cubic Nonlinear Schrodinger Dynamics on Periodic Domains ​

Author: Emmanuel E. Oguadimma, Victory C. Obieke, Xueying Yu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.AP, math.NA

arXiv:2606.27459v2 Announce Type: replace Abstract: We consider the cubic nonlinear Schr"odinger (NLS) equation on two-dimensional flat tori with varying aspect ratios. In this formulation, the choice of aspect ratio governs the Fourier resonance structure, so rational and irrational geometries can...

📖 Read original article


173. A Linear Matching Bandit Approach to Online Multi-Human Multi-Robot Teaming ​

Author: Yaohui Guo, X. Jessie Yang, Cong Shi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.29221v2 Announce Type: replace Abstract: We address the problem of online multi-human multi-robot matching through the lens of a linear matching bandit framework, where a learner assigns robots with unknown features from a fixed pool to distinct sets of human agents over multiple rounds. ...

📖 Read original article


174. A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function ​

Author: Roberto Rossi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML

arXiv:2607.03664v3 Announce Type: replace Abstract: The Gaussian Error Linear Unit is usually motivated as the expected output of an input-dependent Bernoulli gate. This work gives an alternative interpretation: GELU is the expected output of a hard linear gate with a Gaussian random threshold. This...

📖 Read original article


175. Dissociating the Internal Representations of Sycophancy in LLMs ​

Author: Anthony Baez, Sheer Karny, Pat Pataranutaporn
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.07003v2 Announce Type: replace Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raisi...

📖 Read original article


176. The Computational Basis of Confidence in Large Language Models ​

Author: Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov, Petar Veli\v{c}kovi'c, Nathaniel Daw
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12447v2 Announce Type: replace Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is cali...

📖 Read original article


177. RF Spectrogram Anomaly Detection with Quantum Kitchen Sinks: Architecture, Representation, and Hardware Validation ​

Author: Abdallah Aaraba, Alexis Vieloszynski, Remon Polus, Soumaya Cherkaoui, Ola Ahmad
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.13897v3 Announce Type: replace Abstract: The broadcast nature of wireless channels exposes radio-frequency (RF) networks to anomalous and malicious transmissions, making anomaly detection a fundamental requirement for secure spectrum management. Quantum Kitchen Sinks (QKS) offer a lightwe...

📖 Read original article


178. Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers ​

Author: James O' Neill, Fergal Reid
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.15456v2 Announce Type: replace Abstract: Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/...

📖 Read original article


179. Interpretable Anomaly and Drift Detection with Gaussian Mixture Models ​

Author: Behnam Asadi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16811v2 Announce Type: replace Abstract: We revisit Gaussian Mixture Models (GMMs) as a lightweight, interpretable tool for anomaly detection and, in particular, for detecting distributional drift in data streams. We make three practical choices explicit and evaluate them on seven public ...

📖 Read original article


180. FlashPDE: A Drop-In Fused Triton Operator Library for Neural PDE Solvers ​

Author: Peiyu Zang, Bosen Xie, Ruoxiang Xu, Yongqiang Cai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.MS

arXiv:2607.18020v3 Announce Type: replace Abstract: Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memory overhead and inefficient execution of grid-based PDE op...

📖 Read original article


181. Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster ​

Author: Kushal Chakrabarti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account -- more data, retrieval, or scale -- misses an auto-regressive risk residual that increases with scale: the ...

📖 Read original article


182. Now We Know? A Systematic Comparison of TerraMind and THOR ​

Author: Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.18504v2 Announce Type: replace Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific arte...

📖 Read original article


183. Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning ​

Author: Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.18722v3 Announce Type: replace Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but the resulting staleness is an inevitable byproduct, compounded jointly by policy lag, engine delays, and mixture-of-experts routing. Fro...

📖 Read original article


184. H$^2$SD: Hybrid Hindsight Self-Distillation ​

Author: Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Xiaocheng Feng, Bing Qin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.18955v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teache...

📖 Read original article


185. Post-Training in Time Series Foundation Models: A Unifying Framework ​

Author: Shifeng Xie, Ambroise Odonnat, Zehao Xiao, Lei Zan, Malik Tiomoko, Lujia Pan, Themis Palpanas, Boris N. Oreshkin, Chenghao Liu, Keli Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.20002v2 Announce Type: replace Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention to handle domain ...

📖 Read original article


186. AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries ​

Author: Mengda Xing (CRIL, UA), Jean-Marie Lagniez (CRIL, UA), Alejandro Franco (LRCS)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.20577v2 Announce Type: replace Abstract: Physics-based simulations are essential for understanding the electrode-scale discharge behavior of lithium-ion batteries (LIBs) but suffer from prohibitive computational costs. To address this, we introduce a novel deep learning surrogate pipeline...

📖 Read original article


187. GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries ​

Author: Miguel P. Bento, Jo~ao F. Seabra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.20757v2 Announce Type: replace Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in-training by introducing a LogSumExp term to the loss that breaks the symmetries, thus selecting a...

📖 Read original article


188. Robust Asynchronous Q-Learning under Reward and State Corruption via Batching ​

Author: Sreejeet Maity, Aritra Mitra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.20822v2 Announce Type: replace Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observ...

📖 Read original article


189. Test-Time Scaling via Error Localization ​

Author: Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta, Deepanway Ghosal, Aravindan Raghuveer
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.21453v2 Announce Type: replace Abstract: Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn ...

📖 Read original article


190. The Role of Pseudo-labels in Self-training Linear Classifiers on High-dimensional Gaussian Mixture Data ​

Author: Takashi Takahashi
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.LG, math.ST, stat.TH

arXiv:2205.07739v4 Announce Type: replace-cross Abstract: Self-training (ST) is a simple yet effective semi-supervised learning method. However, why and how ST improves generalization performance by using potentially erroneous pseudo-labels is still not well understood. To deepen the understanding o...

📖 Read original article


191. Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection ​

Author: Xinjie Cui, Yuezun Li, Delong Zhu, Jiaran Zhou, Junyu Dong, Siwei Lyu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.CR, cs.LG

arXiv:2411.19715v4 Announce Type: replace-cross Abstract: We describe Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is non-trivial as forgery-relate...

📖 Read original article


192. Ask for More Than Bayes Optimal: A Theory of Indecisions for Selective Hypothesis Testing ​

Author: Mohamed Ndaoud, Peter Radchenko, Bradley Rava
Published: 7/27/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.ME, stat.ML, stat.TH

arXiv:2412.12807v4 Announce Type: replace-cross Abstract: Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high. Given a target accuracy, our goal is to minimize the numbe...

📖 Read original article


193. Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation ​

Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT

arXiv:2501.18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal ge...

📖 Read original article


194. PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework ​

Author: Abhineet Agarwal, Fange Xiao, Rebecca Barter, Omer Ronen, Boyu Fan, Bin Yu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.ME, stat.TH

arXiv:2505.08784v3 Announce Type: replace-cross Abstract: As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. In this paper we introduce PCS-UQ, a framework based on the Predictability, Computability, and Stability (PCS) principle...

📖 Read original article


195. Statistical mechanics of extensive-width Bayesian neural networks near interpolation ​

Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT

arXiv:2505.24849v2 Announce Type: replace-cross Abstract: For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models and committee...

📖 Read original article


196. Fast State-Augmented Learning for Wireless Resource Allocation with Dual Variable Regression ​

Author: Yigit Berkay Uslu, Navid NaderiAlizadeh, Mark Eisen, Alejandro Ribeiro
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SP, cs.LG

arXiv:2506.18748v2 Announce Type: replace-cross Abstract: We consider resource allocation problems in multi-user wireless networks, where the goal is to optimize a network-wide utility function subject to constraints on the ergodic average performance of users. We demonstrate how a state-augmented g...

📖 Read original article


197. A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx ​

Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SE

arXiv:2508.10017v2 Announce Type: replace-cross Abstract: Federated Learning (FL) presents a groundbreaking approach for collaborative health research, allowing model training on decentralized data while safeguarding patient privacy. FL offers formal security guarantees when combined with Differenti...

📖 Read original article


198. Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators ​

Author: Sven Dummer, Tjeerd Jan Heeringa, Jos'e A. Iglesias
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.AI, cs.LG, stat.ML

arXiv:2509.26371v3 Announce Type: replace-cross Abstract: Recently, there has been growing interest in characterizing the function spaces underlying neural networks. While shallow and deep scalar-valued neural networks have been linked to scalar-valued reproducing kernel Banach spaces (RKBS), $\math...

📖 Read original article


199. Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation ​

Author: Eduardo Fernandes Montesuma, Yassir Bendou, Mike Gartrell
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2510.04602v4 Announce Type: replace-cross Abstract: Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not because as they assume access to the complete set of samples f...

📖 Read original article


200. Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation ​

Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT

arXiv:2510.24616v5 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks. A long-standing question remained on its capacity to tackle deep learning models capturing rich feature learning effects, thus going beyond the na...

📖 Read original article


201. Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning ​

Author: Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, d...

📖 Read original article


202. Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense ​

Author: Samaresh Kumar Singh, Joyjit Roy, Sriharsha Anand Pushkala
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC, cs.LG

arXiv:2512.23849v2 Announce Type: replace-cross Abstract: Sophisticated attackers can evade detection-based security by using encryption, stealth tactics, and low-rate attack patterns. This challenge is particularly acute in Internet of Things (IoT) and edge environments, where limited resources mak...

📖 Read original article


203. Atlas 2 -- Foundation models for clinical deployment ​

Author: Maximilian Alber, Timo Milbich, Alexandra Carpen-Amarie, Stephan Tietz, Jonas Dippel, Lukas Muttenthaler, Beatriz Perez Cancer, Alessandro Benetti, Panos Korfiatis, Elias Eulig, J'er^ome L"uscher, Jiasen Wu, Sayed Abid Hashimi, Gabriel Dernbach, Simon Schallenberg, Neelay Shah, Moritz Kr"ugener, Aniruddh Jammoria, Jake Matras, Patrick Duffy, Matt Redlon, Philipp Jurmeister, David Horst, Lukas Ruff, Klaus-Robert M"uller, Frederick Klauschen, Andrew Norgan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2601.05148v2 Announce Type: replace-cross Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this r...

📖 Read original article


204. On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization ​

Author: Sharan Sahu, Cameron J. Hogan, Martin T. Wells
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.OC

arXiv:2601.12238v5 Announce Type: replace-cross Abstract: In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-...

📖 Read original article


205. Cross-reality location privacy protection in 6G-enabled vehicular metaverses: an LLM-enhanced hybrid generative diffusion model-based approach ​

Author: Xiaofeng Luo, Jiayi He, Jiawen Kang, Ruichen Zhang, Zhaoshui He, Ekram Hossain, Dong In Kim
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NI, cs.CR, cs.HC, cs.LG

arXiv:2601.12311v2 Announce Type: replace-cross Abstract: The emergence of 6G-enabled vehicular metaverses enables Autonomous Vehicles (AVs) to operate across physical and virtual spaces through space-air-ground-sea integrated networks. The AVs can deploy AI agents powered by large AI models as pers...

📖 Read original article


206. Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition ​

Author: Rui-Juan Jing, Yuegang Zhao, Changbo Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SC, cs.LG

arXiv:2601.13731v2 Announce Type: replace-cross Abstract: Symbolic computation, powered by modern computer algebra systems, has important applications in mathematical reasoning through exact deep computations. The efficiency of symbolic computation is largely constrained by such deep computations in...

📖 Read original article


207. Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning ​

Author: Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
Published: 7/27/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG

arXiv:2601.17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observ...

📖 Read original article


208. LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics ​

Author: Anselm Strohmaier, Samira B"odefeld, Oliver Straser, Frank Reinhold
Published: 7/27/2026, 4:00:00 AM
Categories: math.HO, cs.LG

arXiv:2601.18685v4 Announce Type: replace-cross Abstract: The capabilities of generative AI in mathematics education are rapidly evolving, posing significant challenges for research to keep pace. Research syntheses remain scarce and risk being outdated by the time of publication. To address this iss...

📖 Read original article


209. The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis ​

Author: Soroosh Tayebi Arasteh, Mina Farajiamiri, Mahshad Lotfinia, Behrus Hinrichs-Puladi, Jonas Bienzeisler, Mohamed Alhaskir, Mirabela Rusu, Christiane Kuhl, Sven Nebelung, Daniel Truhn
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2601.19618v2 Announce Type: replace-cross Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is the strongest known remedy. Practice increasingly favors large generic self-supervised encoders...

📖 Read original article


210. Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases ​

Author: Vid Kocijan, Jinu Sunil, Jan Eric Lenssen, Viman Deb, Xinwei Xe, Federico Reyes Gomez, Matthias Fey, Jure Leskovec
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG

arXiv:2602.09572v3 Announce Type: replace-cross Abstract: The purpose of predictive modeling on relational data is to predict future or missing values in a relational database, for example, future purchases of a user, risk of readmission of the patient, or the likelihood that a financial transaction...

📖 Read original article


211. Statistical Early Stopping for Reasoning Models ​

Author: Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Edgar Dobriban
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML

arXiv:2602.13935v3 Announce Type: replace-cross Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically pr...

📖 Read original article


212. The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games ​

Author: Nikolaos Al. Papadopoulos, Ismael Tito Freire, Marti Sanchez-Fibla, Konstantinos E. Psannis
Published: 7/27/2026, 4:00:00 AM
Categories: cs.MA, cs.GT, cs.LG

arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization. Conventional outcome-based fairness measures can assign similar aggregate scores to temporally disti...

📖 Read original article


213. Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance ​

Author: Hangyi Zhao
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.GT, cs.LG

arXiv:2603.06851v3 Announce Type: replace-cross Abstract: We study contextual bilateral trade under full feedback when, conditionally on the context, trader valuations have bounded density but infinite variance. We first extend the self-bounding property of Bachoc et al. (ICML 2025) from bounded to ...

📖 Read original article


214. Minimum Norm Interpolation via the Local Theory of Banach Spaces: The Role of $2$-Uniform Convexity ​

Author: Gil Kur, Pierre Bizeul
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.LG, math.MG, math.PR, math.ST, stat.TH

arXiv:2603.28956v2 Announce Type: replace-cross Abstract: The minimum-norm interpolator (MNI) framework has recently attracted considerable attention as a tool for understanding generalization in overparameterized models, such as neural networks. In this work, we study the MNI under a $2$-uniform co...

📖 Read original article


215. Parameterized Quantum Circuits as Feature Maps: Representation Quality and Readout Effects in Multispectral Land-Cover Classification ​

Author: Ralntion Komini, Aikaterini Mandilara, Georgios Maragkopoulos, Dimitris Syvridis
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2604.26675v2 Announce Type: replace-cross Abstract: We investigate variational quantum classifiers (VQCs) for land-cover classification from multispectral satellite imagery, adopting a feature-map perspective in which the quantum circuit defines a nonlinear data embedding while the readout det...

📖 Read original article


216. Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs ​

Author: Naomi Esposito, Anthony Tricarico, Luisa Porzio, Ali Aghazadeh Ardebili, Massimo Stella
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC, cs.LG, cs.SI

arXiv:2604.27618v2 Announce Type: replace-cross Abstract: Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason a...

📖 Read original article


217. SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation ​

Author: Jingxi Qiu, Zeyu Han, Cheng Huang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.IR, cs.LG

arXiv:2605.03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer. We study evidence sufficiency verification for selective RAG ans...

📖 Read original article


218. Scalable Gaussian process inference via neural feature maps ​

Author: Anthony Stephenson
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2605.10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels. We show that the learned feature map can be interpreted as an optimal low-rank approximation to a Gram matrix d...

📖 Read original article


219. Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform ​

Author: Oliver Hennh"ofer, Maximilian Kirsch, Christine Preisach
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.CO

arXiv:2605.13642v2 Announce Type: replace-cross Abstract: Most anomaly detection systems output scores rather than calibrated decisions, leaving practitioners to choose thresholds heuristically and without clear statistical interpretation. Conformal anomaly detection addresses this limitation by con...

📖 Read original article


220. Approximation and learning of anisotropic and mixed smooth functions by deep ReLU neural networks ​

Author: Yunfei Yang, Jun Fan
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA

arXiv:2605.31152v2 Announce Type: replace-cross Abstract: This paper studies how efficiently deep ReLU neural networks can approximate and learn smooth functions. When the error is measured in $L^p([0,1]^d)$ norm and the approximator is a network with width $W$ and depth $L$, recent works have prove...

📖 Read original article


221. Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017 ​

Author: Zach Moczkodan (Royal Military College of Canada, Kingston, Canada), Hany Ragab (Royal Military College of Canada, Kingston, Canada)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2606.11098v2 Announce Type: replace-cross Abstract: Recent deep learning approaches for network intrusion detection increasingly incorporate temporal architectures such as recurrent networks and Transformers, often reporting near-perfect performance on CIC-IDS2017. However, many existing studi...

📖 Read original article


222. Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks ​

Author: Greg Ongie, Rahul Parhi
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.LG, math.OC, stat.ML

arXiv:2606.14954v4 Announce Type: replace-cross Abstract: We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its representation cost and native function space, prove existenc...

📖 Read original article


223. Local Multimodal Music Alignment from Global Supervision ​

Author: Irmak Bukey, Zachary Novack, Jongmin Jung, Dasaem Jeong, Chris Donahue
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.LG, cs.MM

arXiv:2607.10023v2 Announce Type: replace-cross Abstract: Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain-in pract...

📖 Read original article


224. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs ​

Author: Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.14186v5 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, eac...

📖 Read original article


225. Operator-Informed Gaussian Processes for Complex Helmholtz Wavefields: From Synthetic Benchmarks to In Vivo Brain Elastography ​

Author: Boyuan Deng, Kshitiz Upadhyay, Michael Shields
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA, physics.med-ph

arXiv:2607.14193v2 Announce Type: replace-cross Abstract: The Helmholtz equation governs time-harmonic wave propagation, and in dissipative media a complex modulus renders its squared wavenumber $\kappa^2$ complex. Inferring such fields from sparse, noisy data calls for solvers that also quantify th...

📖 Read original article


226. It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability ​

Author: Carson Rodrigues
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.16292v3 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a downstream ...

📖 Read original article


227. Entanglement geometry separates circuit cutting, classical hardness, and trainability ​

Author: Maria Gragera Garces, Sabina Dr\u{a}goi, Lirand"e Pira
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.DC, cs.ET, cs.LG

arXiv:2607.17872v2 Announce Type: replace-cross Abstract: Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability. We show that these properties are strongly constrained...

📖 Read original article


228. Multi-Mask Diffusion Language Models for Few-Step Generation ​

Author: Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2607.19686v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal e...

📖 Read original article


229. DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making ​

Author: Raffi Khatchadourian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.20491v2 Announce Type: replace-cross Abstract: A financial AI agent can repeat a decision while changing the tools, order, or recorded arguments and results used to reach it. Outcome-only evaluation misses this variation, even when it matters for replay and change control. DFAH-Bench oper...

📖 Read original article


230. The Geometry of Personality: Activation Steering with Jungian Cognitive Functions ​

Author: Liu Zai (University of Glasgow), Yumeng Wang (Leiden University), Junchen Fu (University of Glasgow), Joemon M. Jose (University of Glasgow)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.20803v2 Announce Type: replace-cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled ...

📖 Read original article