arXiv cs.LG - 2026-07-27 ​
230 items collected.
1. Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees ​
Author: Lei Yang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detec...
2. On the Depth Scalability of Logic Gate Networks ​
Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO
arXiv:2607.21633v1 Announce Type: new Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in de...
3. MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion ​
Author: Tasfia Nuzhat Ornee, Elias Hossain, Ivan Garibay, Niloofar Yousef
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21634v1 Announce Type: new Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence representation, implicitly treating structurally heterogeneous molecular components as equally diffic...
4. Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions ​
Author: Pin Qian, Su Wang, Yihang Chen, Qiaolin Yu, Xiaoyuan Wang, Zhitong Guo, Zhicheng Wang, Junxian You
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benc...
5. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​
Author: Jie Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21636v1 Announce Type: new Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and c...
6. Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning ​
Author: Julian G. Soltes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.21637v1 Announce Type: new Abstract: This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and aggregate an optimal prior ...
7. Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations ​
Author: Jonathan Gallagher, Roberto Guglielmi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.21644v1 Announce Type: new Abstract: We present a goal-agnostic control framework for partial differential equations (PDEs) built around a joint-embedding predictive architecture (JEPA). The small 2D ViT encoder and action-conditioned latent dynamics are trained offline without a reward o...
8. Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not ​
Author: Kavya Bhand, Aadi Joshi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21645v1 Announce Type: new Abstract: Multi-horizon latent consistency is a common training knob in video predictors and world models, but practitioners rarely know what it does to transition geometry. We treat lambda, the weight on multi-step latent agreement, as a diagnostic control and ...
9. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning ​
Author: Timothy Tomashevskiy
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement learning methods typica...
10. A Drift Stable Quantum Federated Learning for Intelligent Services ​
Author: Shanika Iroshi Nanayakkara, Shiva Raj Pokhrel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21647v1 Announce Type: new Abstract: Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distributed ...
11. Shallower ReLU Network Representations via Exact Linear Algebra ​
Author: Kilian Rue{\ss}, Gennadiy Averkov, Florestan Brunck, Moritz Grillo, Christoph Hertrich, Georg Loho, Jack Stade, Moritz Stargalla, Matthew Sun, Martin Winter
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NE, math.CO
arXiv:2607.21651v1 Announce Type: new Abstract: We prove that the maximum of $n$ real numbers is exactly representable by a ReLU network with two hidden layers for every $n\le 10$. The constructions are obtained by reducing the problem to exact rational linear algebra: after a symmetry reduction, th...
12. Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning ​
Author: Jian Hu, Huiying Li, Hao Zhang, Binfeng Xu, Yifan Zhang, Shaokun Zhang, Hemil Desai, Michael Demoret, Pavlo Molchanov, Jan Kautz, Yi Dong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, cs.DC
arXiv:2607.21653v1 Announce Type: new Abstract: Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the...
13. Physically Constrained Federated Additive Models for O-RAN SLA-Risk Prediction ​
Author: Aubida A. Al-Hameed, Mohammed M. H. Qazzaz, Maryam Hafeez, Syed A. Zaidi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.21665v1 Announce Type: new Abstract: Proactive service assurance in O-RAN requires predicting per-slice SLA violations before they occur. The prediction model must be auditable by operators and must train across base stations without pooling per-slice KPIs, which are commercially sensitiv...
14. Neural Feature Governance: Extending Atom Prevalence ​
Author: Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait Fokou'e
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21671v1 Announce Type: new Abstract: Neural network compression and interpretability remain open challenges in modern deep learn- ing, where billion-parameter architectures deliver impressive accuracy at the cost of trans- parency, computational efficiency, and reliable uncertainty quanti...
15. Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration ​
Author: Vishnu Bindu Balachandran
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.CV, stat.ML
arXiv:2607.21673v1 Announce Type: new Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynamical law. Modelling bank impurity as a generalized P'olya urn, we prove almost-sure convergence to a ...
16. Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA ​
Author: Takato Yasuno
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21680v1 Announce Type: new Abstract: Bridge infrastructure deteriorates gradually, yet its root causes---salt intrusion, freezing, fatigue cracking, and others---remain invisible to the naked eye. Expert diagnosis relies on tacit knowledge built over years of practice. We address the chal...
17. CARNet Cycle-Conditioned Core Aggregation and Redistribution for Multivariate Time Series Forecasting ​
Author: Awsaf Tausif Adib, Md. Shahria Sarker Shuvo, Md. Estehaar Ahmed Emon, Mustafa Kamal, Fuad Rahman, Shafin Rahman, Nabeel Mohammed
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21681v1 Announce Type: new Abstract: Accurately modeling cross-variate dependencies remains a key challenge in multivariate time series forecasting, particularly in the presence of strong periodic patterns. Many existing approaches rely on attention-based mechanisms that incur quadratic c...
18. Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets ​
Author: Jim Allchin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that attention reveals which context...
19. An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning ​
Author: Maximilian Dax, Theo Heimel, Gilles Louppe
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, astro-ph.CO, astro-ph.GA, hep-ex, hep-ph, stat.ML
arXiv:2607.21702v1 Announce Type: new Abstract: Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects. We provide an overview of the Bayesia...
20. A Defense of the Quadratic Model ​
Author: Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML
arXiv:2607.21716v1 Announce Type: new Abstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimizatio...
21. RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory ​
Author: Zahra Yousefijamarani, Alaa Alameldeen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AR
arXiv:2607.21731v1 Announce Type: new Abstract: Transformers are widely used across many domains, including natural language processing, computer vision, web search, and DNA sequence analysis. Given their broad applicability, improving the performance of transformer models is critical. However, the ...
22. Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection ​
Author: Debarshi Kundu, Swaroop Ghosh, Vasant Honavar
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21752v1 Announce Type: new Abstract: Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive approaches---including SBM-Transformer, Dynamic Mask Attention, and NSA...
23. Smart predict-then-robustly-optimize ​
Author: Aakil Caunhye, Xuefei Lu, Belen Martin-Barragan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.21773v1 Announce Type: new Abstract: In this paper, we propose and study a robust variant of the smart predict-then-optimize approach that accounts for prediction shifts due to disturbance in the covariate feature space. While traditional integrated-learning-and-optimization models assume...
24. Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection ​
Author: Othmane Harraq, Tamer Aldwairi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CR, cs.CV, cs.MM
arXiv:2607.21776v1 Announce Type: new Abstract: Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify. Unlike face-swap ma- nipulation, TF ...
25. Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness ​
Author: Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21806v1 Announce Type: new Abstract: Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice. There is a growing recognition of the need to evaluate the causal impact of deploying t...
26. Data eccentricity, asymptotics of Gaussian RBF reproducing kernel Hilbert space, and kernel PCA ​
Author: Sergio A. Alvarez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21823v1 Announce Type: new Abstract: We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric proper...
27. A Graph-Based Control Interface for Traffic Signals on Heterogeneous Road Networks ​
Author: Bertil Braun
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.21831v1 Announce Type: new Abstract: We present a traffic-signal control interface in which a shared graph neural network assigns scores to individual traffic movements. Each junction converts these scores into its own variable-sized set of legal signal phases using a deterministic incide...
28. Searching the Space of Feed-Forward Neural-Network Weight-Update Rules with Fixed Depth Symbolic Regression ​
Author: Charles Brum, Edward Finkelstein
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21855v1 Announce Type: new Abstract: We investigate whether symbolic regression can discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks. Candidate update rules are represented as fixed-depth symboli...
29. LeAct: Learning to Reason from Expert Actions ​
Author: Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem p...
30. Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study ​
Author: Kaihua Ding
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.21866v1 Announce Type: new Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 students each ra...
31. Variance-Reduced Q-Learning over Static and Time-Varying Networks ​
Author: Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath Jr
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.21876v1 Announce Type: new Abstract: We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value f...
32. Remedying Coarsening-Based GNN Training under Heterophily via Adaptive Complementary Enhancement ​
Author: Guoming Li, Jian Yang, Xukun Wang, Zixiao Wang, Shangsong Liang, Yifan Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, eess.SP, math.NA
arXiv:2607.21885v1 Announce Type: new Abstract: Coarsening-based training for graph neural networks (GNNs), i.e.\ training on coarsened graphs rather than the original large ones, has become a promising direction for scaling GNNs to massive graphs. However, prior work has been evaluated almost exclu...
33. MissHyper: Restoring Clinical Synchronicity in Missingness-Guided Hypergraph Forecasting ​
Author: Mingyi Ma, Qingxiong Tan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21922v1 Announce Type: new Abstract: Clinical irregular multivariate time series are shaped not only by physiological dynamics but also by the measurement process that determines when and what to observe. In event-centric models, however, co-timestamp structure can be flattened too early:...
34. RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention ​
Author: Anderson R. Santos
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this constraint as a model-a...
35. LatentFlow: Visual Analytics for Latent Space Analysis in Molecular Graph Neural Networks ​
Author: Shiyi Liu, Jiaqing Chen, Nicholas Hadler, Rostyslav Hnatyshyn, Michael W. Mahoney, Talita Perciano, John F. Hartwig, Gunther H. Weber, Ross Maciejewski
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.HC
arXiv:2607.21941v1 Announce Type: new Abstract: Chemists and materials scientists increasingly use machine learning models, such as graph neural networks (GNNs), to predict properties of molecules and the outcomes of their reactions. Beyond predictive performance, understanding how these models orga...
36. Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge ​
Author: Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21946v1 Announce Type: new Abstract: This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image. Visual ph...
37. MA-DAR: Manifold-Aligned Dynamic Adaptive Routing for Continual Temporal Knowledge Graph Reasoning ​
Author: Xiangjun Shi, Chong Mu, Jinchuan Zhang, Lizong Zhang, Yuefeng He, Shang Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21949v1 Announce Type: new Abstract: Continual temporal knowledge graph (TKG) reasoning aims to continuously incorporate newly emerging facts while preserving previously acquired knowledge. Replay-based continual learning has achieved promising performance by revisiting historical represe...
38. Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ​
Author: Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21971v1 Announce Type: new Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflectio...
39. On the Convergence of Stochastic Low-Rank Adaptation ​
Author: Ru Wang, Chengchang Liu, John C. S. Lui
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.21975v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \m...
40. From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning ​
Author: Jiaxin Deng, Junbiao Pang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21999v1 Announce Type: new Abstract: Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape. Conventional re-sampling addresses the former without con...
41. Energy Manifold Natural Gradient Descent: Riemannian Optimization for Neural PDE Solvers ​
Author: Zhangyong Liang, Huanhuan Gao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain. We introduce \EMNGDfull{}, a manifold optimization ...
42. Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits ​
Author: Yuta Natsubori, Masataka Ushiku, Yuta Saito
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22012v1 Announce Type: new Abstract: Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handl...
43. DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection ​
Author: Xiafeng Man
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common styli...
44. CEL: Comprehensive Counterfactual Explanations Library and Benchmark ​
Author: Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minim...
45. A Leakage-Free Stacked Ensemble Method for Multiclass Classification ​
Author: S. P. Sharmila, Aruna Tiwari
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22081v1 Announce Type: new Abstract: Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as...
46. Pretraining EHR Foundation Models with Patient-Aware Sampling ​
Author: Joshua Placidi, Yuxuan Liu, Jinpei Han, Marek Rei, A. Aldo Faisal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22114v1 Announce Type: new Abstract: Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into a single token stream and windows are sampled from that stream. In EHR ...
47. TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex ​
Author: Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22143v1 Announce Type: new Abstract: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of molec...
48. Unbiased Open World Regularization for Fair Self-Supervised Learning ​
Author: L{'e}o Nicollier (CB, ATT), Marc Pic (ATT), Pablo Mus{'e} (CB, IFUMI), Enric Meinhardt-Llopis (CB), Gabriele Facciolo (CB)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset. These techniques rely on regularization, which prevents representation co...
49. From Score Approximation to Distribution Approximation in Score-Based Diffusion Models ​
Author: Lan V. Truong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.IT, math.IT, stat.ML
arXiv:2607.22199v1 Announce Type: new Abstract: Score-based diffusion models have achieved remarkable empirical success in generative modeling, yet their approximation-theoretic foundations remain incomplete. In particular, although classical universal approximation theorems guarantee that neural ne...
50. Latent PDE mapping for efficient physics-informed learning across geometries with limited data ​
Author: Ingvild Askim Adde, Mary M. Maleckar, Gabriel Balaban
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph
arXiv:2607.22215v1 Announce Type: new Abstract: In this study, we introduce latent PDE mapping, a broadly applicable physics-informed learning technique designed to enable efficient geometric generalization with sparse training data. Latent PDE mapping pulls back geometry-specific PDE residuals and ...
51. Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression ​
Author: Hirotaka Sugawara, Yujin Taguchi, Kei Minagawa, Yusuke Hiki, Takashi Morikura, Akira Funahashi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22238v1 Announce Type: new Abstract: Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables by balancing exploration and exploitation. BO is often used under a limited evaluation budget, such...
52. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning ​
Author: Wei Zhang, Xinwu Liu, Yihang Cheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22251v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method for large language models, but its performance depends strongly on how a fixed rank budget is distributed across Transformer modules. Existing adaptive-rank methods usua...
53. Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed Recognition ​
Author: Yi-Hang Zhu, Rajeev Raman, Shiqi Su, Jianyuan Sun, Xinyu Yang, Nan Xing, Huiyu Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.22258v1 Announce Type: new Abstract: Deep learning models using traditional softmax classifiers have achieved remarkable success in various classification tasks. However, their performance degrades significantly on imbalanced datasets. Although Balanced Softmax is widely adopted as a stat...
54. Autoregressive EHR Foundation Models with Multimodal Inputs ​
Author: Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way. We prese...
55. An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection ​
Author: Romain Hermary, Nesryne Mejri, Djamila Aouada
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.22286v1 Announce Type: new Abstract: Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. Although metrics such as AUROC, AUPR, F1-score, and MCC are widely used, their values convey different meanings depend...
56. Efficient Recommendations via Graph Coarsening and Label Propagation ​
Author: Alessandro Sbandi, Federico Siciliano, Fabrizio Silvestri
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22287v1 Announce Type: new Abstract: Graph-based recommendations are widely adopted in real-world industrial applications. However, graphs in these systems often reach a massive scale, posing notable scalability and efficiency challenges. This requires techniques that can effectively bala...
57. Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning ​
Author: Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez, Cole Looney, Xiaoliang Wu, Alexandra Livia Georgescu, Stefano Goria
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SD
arXiv:2607.22304v1 Announce Type: new Abstract: Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice c...
58. Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs ​
Author: Zhangzhi Xiong, Minzhang Li, Haotian Yu, Sixian Shen, Kexin Zhang, Mingrui Li, Jie Zheng, Kewei Tu, Jingyi Yu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22314v1 Announce Type: new Abstract: Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based fil...
59. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization ​
Author: Hao Wang, Kun Yuan, Wenlin Zhong, Minglei Zhang, Han Xiao, Ming Sun, Honggang Qi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.22334v1 Announce Type: new Abstract: Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while...
60. Beyond Binary Rooftop Mapping: A Four-Class Deep Learning Framework for Green Roof Potential Assessment from Open Swiss Geospatial Data ​
Author: Htet Yamin Ko Ko
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22342v1 Announce Type: new Abstract: The development of effective urban climate adaptation strategies requires comprehensive spatial information on rooftops and buildings, since such information underpins the assessment of ecosystem services provided by green infrastructure, particularly ...
61. IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data ​
Author: Masashi Sode, Gianmarco Pinton
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.med-ph
arXiv:2607.22351v1 Announce Type: new Abstract: The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated soun...
62. Integrated Order Dispatching and Routing for Last-Mile Pickup via Deep Reinforcement Learning ​
Author: Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22356v1 Announce Type: new Abstract: In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms. This challenge is fundamentally driven by two key and tightly coupled decision-making processes:...
63. Indexing: the Beginning and the End ​
Author: Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22361v1 Announce Type: new Abstract: We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits ...
64. Interior interpretability with attention rollout: contraction and propagation profiles in Transformers ​
Author: Umberto Biccari, Qian Huang, Enrique Zuazua
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretabil...
65. Local-Global Geometric Insights for Graph Neural Networks via Entropic Curvature ​
Author: Rachid Caich, Yassine Abbahaddou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SI, stat.ML
arXiv:2607.22381v1 Announce Type: new Abstract: Curvature notions on graphs, particularly Ollivier-Ricci and Forman, have emerged as powerful tools for addressing fundamental issues in Graph Neural Networks (GNNs) such as oversmoothing and oversquashing, but rely almost exclusively on local edge-lev...
66. LunarFM: A Shared Multimodal Representation of the Moon's Surface ​
Author: Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, Abigail Calzada-Diaz, Sylvester Kaczmarek, Ra'ul Ramos-Poll'an
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22408v1 Announce Type: new Abstract: The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast ...
67. On the Identifiability of Controlled World Models ​
Author: Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework fo...
68. Hyperball May Not Be a Free Lunch ​
Author: Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22444v1 Announce Type: new Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Startin...
69. Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining ​
Author: V'ictor Rinc'on Yepes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22458v1 Announce Type: new Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from ...
70. Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates ​
Author: Anjian Li, Ryne Beeson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22467v1 Announce Type: new Abstract: Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve. We therefore study a $k$-neighborhood data collection strategy ...
71. Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent ​
Author: Peng Zhao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.ML, stat.TH
arXiv:2607.22474v1 Announce Type: new Abstract: In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structur...
72. \k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating ​
Author: Jianghui Wang, Silong Yong, Francesco Orabona, Marco Canini, Katia P. Sycara, Yaqi Xie
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22489v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardl...
73. Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting ​
Author: Aliaksei Kaliutau
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22491v1 Announce Type: new Abstract: Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, and it...
74. Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes ​
Author: Athanasios Papastathopoulos-Katsaros, Steven T. Lee, Lin Yao, Ajay Thomas, Junseok Park, Matthew J. McGinley, Zhandong Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, eess.SP
arXiv:2607.22508v1 Announce Type: new Abstract: Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks. Predefined features carry a strong bias, since they fix in advance what cou...
75. Dysphagia Risk Stratification in Head and Neck Cancer via Two-Stage PRO-Clinical Stacking ​
Author: Siyuan Zhao, Eric Ababio Anyimadu, Zachary G. Brumm, Yue Ma, Clifton David Fuller, Xinhua Zhang, G. Elisabeta Marai, Guadalupe Canahuate
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22514v1 Announce Type: new Abstract: Dysphagia is a debilitating late effect of head and neck cancer (HNC) treatment, yet timely identification of at-risk patients remains challenging in survivorship care. Definitive assessment relies on videofluoroscopic imaging, as captured by the Dynam...
76. An Explainable FFT-Based Spatial-Frequency Fusion Framework for Deepfake Detection ​
Author: Pamela Kirui, Cho Hyuk, Qingzhong Liu, Haodi Jiang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.17441v1 Announce Type: cross Abstract: Deepfake generation has raised growing concerns regarding digital media authenticity, misinformation, identity fraud, and public trust. Recent studies show that combining spatial and frequency features leads to stronger detection results than using i...
77. Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification ​
Author: Guillermo Rubi~nos Rodr'iguez, Mart'in Ottavianelli, Mateo Alonso, Gonzalo Bl'azquez Gil, Boris-Stephan Rauchmann, Pablo D'iez-Valle, Sergio Altares-L'opez
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.21186v1 Announce Type: cross Abstract: Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, while the scalabil...
78. Spectral Flow Certificates for Depth-Aware Long-Range Propagation in Graph Neural Networks ​
Author: Ranjan Veerabhadraswamy, Ajith Jubilson Emerson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21607v1 Announce Type: cross Abstract: Graph Neural Networks propagate information through local message passing, but the graph topologies themselves can silently prevent any amount of training from solving long-range tasks. When we deploy GNNs on new graphs, there is currently no inexpen...
79. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text ​
Author: Miaobo Hu, Xiaobo Guo, Shuhao Hu, Bokun Wang, Rui Chen, Xin Wang, Daren Zha, Jun Xiao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21610v1 Announce Type: cross Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pi...
80. Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures ​
Author: Simon Dennis, Kevin Shabahang, Hao Guo, Rivaan Patil
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21612v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following, style transfer, and factual adaptation. We show that for procedural knowledge--the ability to follo...
81. FrED: External Data Influence Estimation via Domain Knowledge Graph Grounding ​
Author: Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21615v1 Announce Type: cross Abstract: The rapid deployment of generative AI has amplified the critical need for Training Data Attribution to ensure transparency and accountability. However, current parametric approaches require computationally prohibitive access to model weights, while s...
82. Local Synaptic Rules Can Implement a SIGReg Gradient Without Backpropagation ​
Author: Martin Andrews
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV, cs.LG
arXiv:2607.21622v1 Announce Type: cross Abstract: We prove that two canonical local synaptic learning rules, the potentiation arm of spike-timing-dependent plasticity (STDP$^+$) and homeostatic plasticity (instantiated here via flashlight granule-cell-like neurons), together can implement the exact ...
83. FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs ​
Author: Kahou Tam, Wei Niu, Yu Bao, Xiaomin Ouyang, Chengzhong Xu, Li Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG
arXiv:2607.21624v1 Announce Type: cross Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs du...
84. Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems ​
Author: Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21627v1 Announce Type: cross Abstract: End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure mode in which modules preserve or improve end-task performance while ...
85. Computer Vision Based Neurology Brain Activity Rejection Architecture and Implementation ​
Author: Zag ElSayed, Nathan Suer, Grace Westerkamp, Jack Yanchen Liu, Makoto Miyakoshi, Craig Erickson, Ernest Pedapati
Published: 7/27/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG, physics.data-an, q-bio.QM
arXiv:2607.21654v1 Announce Type: cross Abstract: The electroencephalogram (EEG) is a valuable and widely applied tool for investigating brain disorders and behavioral changes. It offers a minimally restrictive and non-invasive method. However, challenges in using EEG for cognitive development studi...
86. Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives ​
Author: Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, Nicola A. Morley, Robert D. J. Oliver, Alexandra J. Ramadan, Delvin Ce Zhang, Katerina A. Christofidou, Haiping Lu
Published: 7/27/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.LG
arXiv:2607.21660v1 Announce Type: cross Abstract: Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompass...
87. Ordered Action Tokens for Visuomotor Policy Learning ​
Author: Chaoqi Liu, Yue Zhao, Haonan Chen, Xiaoshen Han, Jiawei Gao, Ehsan Adeli, Yilun Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long toke...
88. Explainable quantum-compressed machine learning for complex fluid flows ​
Author: Xiao Xue, Maida Wang, Mingyang Gao, Minh Chung, Peter V. Coveney
Published: 7/27/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.LG, quant-ph
arXiv:2607.21688v1 Announce Type: cross Abstract: Machine-learning surrogates of physical systems face a paradox: explainable models facing the challenge of expressivity to capture complex nonlinear flows, whereas expressive deep surrogates match high-fidelity simulations only through massive parame...
89. Prior laundering: learned priors with inherited, undetectable overconfidence ​
Author: Ali Siahkoohi, Sina Alemohammad
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.21721v1 Announce Type: cross Abstract: Learned generative priors are increasingly used for ill-posed Bayesian inverse problems, their posterior uncertainty treated as earned from data. But training one requires truths, scarce in seismic and medical imaging, so the recourse is an archive o...
90. Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery ​
Author: Khalid El-Darymli, Christoph H. Gierull, Katerina Biron, Weimin Huang
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2607.21745v1 Announce Type: cross Abstract: Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study introduces a deep sigma-point process (DSPP) model for predicting RCS in synthetic aperture radar (SAR) imagery using ...
91. Prompt as a Data Type: In-Database LLM Prompt Management and Rewriting ​
Author: Denis Mayr Lima Martins, Gottfried Vossen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.LG
arXiv:2607.21756v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in database-backed applications to classify tuples, filter records using semantic predicates, extract structured attributes, and enrich query results. Yet the prompt that start these computations are...
92. Encoding orders and trees in real-valued functions ​
Author: G Conant, C Terry
Published: 7/27/2026, 4:00:00 AM
Categories: math.CO, cs.LG, math.LO
arXiv:2607.21761v1 Announce Type: cross Abstract: We prove function-theoretic analogues of a quantitative result of Hodges on extracting the order property from a sufficiently large 2-tree coded in a binary relation. Similar analogues for functions were previously obtained by Daskalakis and Golowich...
93. Reliability-Aware Bayesian Optimization of 1310 nm PCSELs with FDTD Verification ​
Author: Jinglin Yu, Feiyang Wu, Longying Wen, Chongxian Yuan, Renjie Li, Zhaoyu Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: physics.optics, cs.LG, physics.app-ph
arXiv:2607.21772v1 Announce Type: cross Abstract: Near 1310 nm photonic-crystal surface-emitting lasers (PCSELs) are attractive narrow-beam sources for optical communication and sensing, but their final design refinement is costly. Small geometry changes simultaneously shift the band-edge resonance,...
94. From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia ​
Author: Rodrigo Vargas Sainz, Christian Ber'on Curti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.21785v1 Announce Type: cross Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local predictive systems to...
95. Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling ​
Author: William Gaudelier, Albert Cohen, Dumitru Potop Butucaru
Published: 7/27/2026, 4:00:00 AM
Categories: cs.PL, cs.LG
arXiv:2607.21797v1 Announce Type: cross Abstract: Previous work has shown that the simple dataflow primitives of the Lustre language allow the natural, semantically unambiguous, and compact representation of machine learning (ML) applications, including models featuring complex conditional execution...
96. Adversarial Prompts for Acceptance Collapse in Speculative Decoding ​
Author: Run Wang, Chaoyi Zhou, Xi Liu, Yi Zhu, Amir Salarpour, Pedram MohajerAnsari, Zhi-Qi Cheng, Feng Luo, Siyu Huang, Mert D. Pes'e
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.CL, cs.LG
arXiv:2607.21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operati...
97. Natural Invariant Measures for Chaotic Game Dynamics: Finding Order in Chaos ​
Author: Jakub Bielawski, Thiparat Chotibut, Fryderyk Falniowski, Micha{\l} Misiurewicz, Georgios Piliouras
Published: 7/27/2026, 4:00:00 AM
Categories: math.DS, cs.LG, econ.TH
arXiv:2607.21805v1 Announce Type: cross Abstract: We study the long-term behavior of the Multiplicative Weights Update (MWU) algorithm in game settings where learning dynamics frequently fail to converge to Nash equilibria and instead exhibit Li-Yorke chaos. While such chaos precludes the prediction...
98. Longitudinal Random Forests for Sparse and Irregular Response Trajectories ​
Author: Yangsheng Wang, Xiaotian Dai, Haoda Fu, Guifang Fu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, stat.ML
arXiv:2607.21817v1 Announce Type: cross Abstract: Longitudinal studies often collect data at sparse, irregular, and unequally spaced time points. Such heterogeneity is often driven by subject-specific covariates, yet existing methods have been restricted to a scalar endpoint value, completely neglec...
99. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors ​
Author: Daniyal Kabir Dar, Arun Ross
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR, cs.LG
arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold w...
100. How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests ​
Author: Iren Mazloomzadeh, Mohammad Mehdi Morovati, Foutse Khomh
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.LG
arXiv:2607.21832v1 Announce Type: cross Abstract: Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows. While developers increasingly ben...
101. Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification ​
Author: Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova, Akira Takahashi, Antigoni Polychroniadou, Nicolas Papernot
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.LG
arXiv:2607.21839v1 Announce Type: cross Abstract: Privacy-preserving machine learning auditing protocols allow auditors to assess models for properties such as accuracy or fairness, without revealing their internals or training data. This makes them especially attractive for auditing models deployed...
102. Quantifying Political Partisanship for Cross-Platform Analyses ​
Author: Fathima Ameen, Christopher G. Healey
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SI, cs.LG
arXiv:2607.21842v1 Announce Type: cross Abstract: Research on political polarization on social media depends on the ability to reliably measure partisanship in user-generated content. However, existing approaches are typically tailored to platform-specific properties, such as structural affordances ...
103. Simulation-Based Empirical Bayes ​
Author: Xinwei Shen, Diana Cai, Cheng Zhang, David M. Blei
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.21843v1 Announce Type: cross Abstract: Empirical Bayes (EB) performs simultaneous inference across many related latent variables. Classical EB assumes that the likelihood p(x | z) is tractable. In many scientific applications, however, the likelihood is available only through a simulator....
104. Distributional Determinantal Point Process for Repulsive Clustering of Distributions ​
Author: Khai Nguyen, Yang Ni, Elizabeth Juarez-Colunga, Peter Mueller
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, stat.AP, stat.CO, stat.ML
arXiv:2607.21847v1 Announce Type: cross Abstract: We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (S...
105. Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 Refinement ​
Author: Mohammadreza Narimani, Vikram Anand, Parastoo Farajpoor
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, eess.IV
arXiv:2607.21881v1 Announce Type: cross Abstract: Agricultural field maps are often proprietary, incomplete, or outdated, yet they provide the spatial framework for crop monitoring, production accounting, and land-conversion analysis. This study presents a reproducible workflow for mapping farmland ...
106. Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes ​
Author: Lu Luo, Dandan Mo, Chengdong Xu, Ting Li, Jinhan Xie, Huiqiong Li, Niansheng Tang
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.21958v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential. Statistical watermarking has emerged as a promising solution, yet most...
107. Unified Static-Dynamic Pruning for Efficient LLM Inference ​
Author: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.AR, cs.LG
arXiv:2607.21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offers a prom...
108. QC-PHAST Search: Classical--Quantum Query Benchmarks for Finite-Pool Rare-Regime Discovery ​
Author: Harsh Milind Tirhekar, Chandrajit Bajaj
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.ET, cs.LG
arXiv:2607.21995v1 Announce Type: cross Abstract: Rare-regime discovery in parameterized dynamical systems is an active-search problem: find one verified parameter at which a scientifically defined qualitative threshold is crossed, even when acceptable candidates are rare, nonconvex, or fragmented. ...
109. Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation ​
Author: M M Asif Ferdous
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.CL, cs.LG
arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor lighting. In such settings, a reliable uncertainty signal matters more than raw accuracy, because it d...
110. Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era ​
Author: Yu Wang, Hongyu Yang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22068v1 Announce Type: cross Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combining complementary representations. In this work, we revisit this assumption in the foundation-model ...
111. MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection ​
Author: Phurich Saengthong, Takahiro Shinozaki
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.LG
arXiv:2607.22086v1 Announce Type: cross Abstract: Autoencoder-based anomalous sound detection is attractive for machine condition monitoring because it can be trained using only normal recordings and yields an interpretable anomaly score from reconstruction error. Most prior work uses spectrogram au...
112. Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models ​
Author: Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22098v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps...
113. One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments ​
Author: Jan Ole von Hartz, Abhinav Valada, Joschka Boedecker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames. However, existing approaches typically assume these frames to be strictly exogenous. Th...
114. CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography ​
Author: Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo'nski, Tomasz Figatowski, Natalia Zieli'nska
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols. In this work we demonstrate a new benchmark for the assessment of deep learning model...
115. Industrial Tokenization for LLM-Based Health Intelligence: A Federated Architecture for Industrial Evidence Integration ​
Author: Deshui Li, Xiao-Ming Yuan, Zishun Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22153v1 Announce Type: cross Abstract: Industrial health management increasingly relies on heterogeneous information sources, including condition monitoring systems, supervisory control and data acquisition systems, maintenance records, inspection results, and prognostic models. Although ...
116. DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents ​
Author: Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.LG
arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a runnin...
117. Bowel Obstruction Detection and Localization on Abdominal CT with Deep Learning ​
Author: Moritz Vandenhirtz, Andrea Agostini, Dana Belde, M'elanie Roschewitz, Ismaiel Chikh Bakri, Tilo Niemann, Andr'e Euler, Julia E Vogt
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22173v1 Announce Type: cross Abstract: Bowel obstruction is a common and potentially life-threatening gastrointestinal condition. In the face of rising diagnostic workloads, the automated diagnosis of bowel obstruction on CT scans supports radiologists by accelerating detection and improv...
118. Trajectory-Regularized Stochastic Optimal Control via KL Divergence ​
Author: Mintae Kim, Koushil Sreenath
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY
arXiv:2607.22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theo...
119. Deep Convolutional Large-Margin $\ell_p$-SVDD for Visual Anomaly Detection ​
Author: Alireza Dastmalchi Saei, Shervin Rahimzadeh Arashloo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22212v1 Announce Type: cross Abstract: Visual anomaly detection requires adaptive representations and reliable decision boundaries, particularly when anomalous training samples are scarce and class distributions are highly imbalanced. Classical kernel-based methods yield principled geomet...
120. Convergence analysis of a family of Zermelo-type iterations for the Bradley--Terry model ​
Author: Ruijian Han, Ding Lu, Yiming Xu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA, stat.CO
arXiv:2607.22221v1 Announce Type: cross Abstract: Zermelo's algorithm is a classical method for computing the maximum likelihood estimator in the Bradley--Terry (BT) model, but its convergence can be slow in practice. To accelerate computation, Newman introduced a family of Zermelo-type fixed-point ...
121. Variational Low-rank Tensor Decomposition for Multisubject Spatiotemporal Data Analysis ​
Author: Laura M. Montaldo, Ricardo A. Borsoi, Sebastian Miron, Tulay Adali
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, eess.SP
arXiv:2607.22262v1 Announce Type: cross Abstract: Modeling shared and subject-specific structure in multisubject spatiotemporal data remains challenging, particularly in neuroimaging, where both spatial and temporal patterns exhibit rich variability across subjects. Existing matrix and tensor decomp...
122. Explicit Iteration Complexity of Exact Data-Driven Inverse Optimization for Integer Linear Programs ​
Author: Akira Kitaoka
Published: 7/27/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG, stat.ML
arXiv:2607.22263v1 Announce Type: cross Abstract: A data-driven inverse optimization problem (DDIOP) is the problem of estimating the objective-function parameters (weights) that explain observed optimal-solution data, and it arises in many applications, including integer linear programming (ILP). I...
123. General Value Functions for Remaining Useful Life and Failure-Mode Prediction ​
Author: Hao Yan, Ali Sarabi, Qing Zou, Boyang Xu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP
arXiv:2607.22268v1 Announce Type: cross Abstract: Remaining useful life (RUL) prediction and failure-mode classification are central tasks in predictive maintenance. Many data-driven pipelines use fixed-window supervised learning with complete terminal labels; such routes do not naturally encode the...
124. Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting ​
Author: Wan Zhang, Qinjie Lin, Chan Lee, Weijian Li, Han Liu, Kai Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.22299v1 Announce Type: cross Abstract: Forecasting multiple time-series with high-dimensional covariates presents a core challenge: unifying common temporal patterns while retaining meaningful series-specific information. We introduce Hopformer (Homogeneity-Pursuit Transformer), a two-sta...
125. Learning Bidirectional Causal Interactions with Heteroscedastic Neural Networks ​
Author: Masahiro Tanaka
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2607.22313v1 Announce Type: cross Abstract: Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence. This paper proposes SEM-DNN, a heterosced...
126. Agentic Root Cause Analysis through Evidence-Grounded Reasoning ​
Author: Amaury Wei, Olga Fink
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22385v1 Announce Type: cross Abstract: Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existi...
127. HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding ​
Author: Chao Fang, Jun Yin, Man Shi, Marian Verhelst
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG
arXiv:2607.22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that ex...
128. Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation Models ​
Author: Sergey Musienko
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.LG, q-bio.NC
arXiv:2607.22397v1 Announce Type: cross Abstract: Frozen EEG encoders proliferate; per-model fine-tune defaults do not scale. We present Nimbus Personalizer: one contract encode to Bayesian head to BrainState (optional affine mid-tier) that sits on heterogeneous frozen trunks without a new personali...
129. Learning Ergodic Dynamical Systems from a Finite Trajectory ​
Author: Oleksii Kachaiev, Silvia Villa, Lorenzo Rosasco
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.22399v1 Announce Type: cross Abstract: We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system. More precisely, we study discrete-time autonomous stochastic systems defining time-homogeneous Markov processes. We first focus on estimati...
130. Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition ​
Author: Austin Rockman
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.IR, cs.LG
arXiv:2607.22413v1 Announce Type: cross Abstract: Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision. We present Reflect...
131. Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability ​
Author: Ahmed M. Abuzuraiq, Philippe Pasquier
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.MM
arXiv:2607.22428v1 Announce Type: cross Abstract: Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque ...
132. Graph-Based Correlation Matrix Generation: A Convex Optimization Approach ​
Author: Ali Fakhar (UGA), K{'e}vin Polisano (UGA), Ir{`e}ne Gannaz (G-SCOP_GROG, G-SCOP, Grenoble INP, UGA), Sophie Achard (STATIFY, LJK, UGA)
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.22436v1 Announce Type: cross Abstract: This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures. We propose a novel convex optimization framework in which an initial matrix is projected onto an elliptope under ...
133. TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI ​
Author: Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2607.22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic app...
134. Singular value soft-thresholding via the polar decomposition ​
Author: Stephen Becker
Published: 7/27/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, math.OC
arXiv:2607.22484v1 Announce Type: cross Abstract: Singular value soft-thresholding can be computed via a reduction to the matrix polar decomposition, which allows one to exploit GPU-friendly algorithms for computing the polar decomposition. Empirically, there is a significant speed-up on GPUs compar...
135. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference ​
Author: Jiyuan Tan, Vasilis Syrgkanis
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, econ.EM
arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers...
136. Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support ​
Author: Peiyong Wang, Udaya Parampalli, Casey R. Myers
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.22516v1 Announce Type: cross Abstract: A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through sp...
137. PinEqualizer: Full Funnel Content Exploration and Debiasing System at Pinterest ​
Author: Olafur Gudmundsson, Bo Zhao, Huayi Liao, Anna Kiyantseva, Sai Xiao, Heath Vinicombe, Mostafa Keikha, Luke DeLuccia, Zihao Chen, Junpeng Hou, Weijie Jiang, Bhawna Juneja, Andreanne Lemay, Wei-Ting Lin, Keyvan Moghadam, Jiaxing Qu, Zhiqing Rao, Zhihua Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.IR, cs.LG
arXiv:2607.22518v1 Announce Type: cross Abstract: In this paper, we propose a new solution for addressing the content cold-start problem in industry-scale search and recommender systems. Compared to prior approaches, we have made the following new contributions: 1) our solution spans the entire mult...
138. Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning ​
Author: Seyeon Kim, Joonhun Lee, Namhoon Cho, Sungjun Han, Wooseop Hwang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.PR, stat.ML
arXiv:2408.02295v4 Announce Type: replace Abstract: Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residuals induced by bootstrapping and exploration. We introduce a state-co...
139. Online Pricing and Allocation with Demand Learning and Fulfillment Cost ​
Author: Jianyu Xu, Xuan Wang, Yu-Xiang Wang, Jiashuo Jiang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML
arXiv:2501.18049v3 Announce Type: replace Abstract: We study online learning for a seller that jointly chooses per-period inventory positions and a uniform price, then fulfills realized demand through a downstream allocation. The main difficulty is not only demand learning: the price shifts demand a...
140. Carpe Diem: Critical Learning Period-Aware Contract-Based Incentives for Federated Learning ​
Author: Thanh Linh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Quoc-Viet Pham
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.GT
arXiv:2503.07869v4 Announce Type: replace Abstract: Critical learning periods (CLPs) in federated learning (FL) refer to early stages during which low-quality contributions (e.g., sparse training data availability) can permanently impair the performance of the global model. However, existing incenti...
141. Spatially-Enhanced Temporal Fusion Transformer: Interpretable Multi-Output Prediction for Parametric Dynamical Systems with Time-Varying Inputs ​
Author: Shuwen Sun, Lihong Feng, Peter Benner
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA
arXiv:2505.00473v2 Announce Type: replace Abstract: We explore the promising performance of a transformer model in predicting outputs of parametric dynamical systems with external time-varying input signals. The outputs of such systems vary not only with physical parameters but also with external ti...
142. Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models ​
Author: Roseline Polle, Agnes Norbury, Alexandra Livia Georgescu, Nicholas Cummins, Stefano Goria
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2505.23378v3 Announce Type: replace Abstract: Speaker-dependent modelling can substantially improve performance in speech-based health monitoring applications. While mixed-effect models are commonly used for such speaker adaptation, they require computationally expensive retraining for each ne...
143. Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning ​
Author: Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2507.01551v3 Announce Type: replace Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, a...
144. A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data ​
Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2509.10517v3 Announce Type: replace Abstract: Machine learning can predict in-hospital mortality, but data privacy and the statistical heterogeneity of clinical data hamper its use. Federated Learning (FL) is privacy-preserving, yet its behavior under non-IID and imbalanced conditions needs sc...
145. LiMuon: Light and Fast Muon Optimizer for Large Models ​
Author: Feihu Huang, Yuning Luo, Songcan Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2509.14562v5 Announce Type: replace Abstract: Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention. More recently, the useful Muon optimizer is specifically designed for matrix-structured parameters of large model...
146. HD3C: Efficient Medical Data Classification for Edge Devices ​
Author: Jianglan Wei, Zhenyu Zhang, Pengcheng Wang, Mingjie Zeng, Zhigang Zeng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2509.14617v4 Announce Type: replace Abstract: Efficient medical data classification is essential for modern disease screening, particularly in resource-constrained environments where power budgets and computing capabilities are limited. We present HD3C, a lightweight classification framework d...
147. SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis ​
Author: Marie Brockschmidt, Maresa Schr"oder, Stefan Feuerriegel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2509.22352v3 Announce Type: replace Abstract: Survival analysis is a cornerstone of clinical research by modeling time-to-event outcomes such as metastasis, disease relapse, or patient death. Unlike standard tabular data, survival data often come with incomplete event information due to dropou...
148. Safe In-Context Reinforcement Learning ​
Author: Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2509.25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history. While ...
149. Correlating Cross-Iteration Noise for DP-SGD using Model Curvature ​
Author: Xin Gu, Yingtai Xiao, Guanlin He, Jiamu Bai, Daniel Kifer, Kiwan Maeng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2510.05416v3 Announce Type: replace Abstract: Differentially private stochastic gradient descent (DP-SGD) offers the promise of training deep learning models while mitigating many privacy risks. However, there is currently a large accuracy gap between DP-SGD and normal SGD training. This has r...
150. Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization ​
Author: Jinwoo Baek
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA
arXiv:2510.21770v2 Announce Type: replace Abstract: Low-precision execution can induce substantial forward discrepancies in Transformers even for fixed weights and input, yet these discrepancies are usually monitored only at the output and lack a layer-wise theoretical account. We develop a first-or...
151. CorVS+: Correspondence-Driven Association of Video Trajectories and Sensors for Identity-Aware Person Localization in Warehouses ​
Author: Kazuma Kano, Yuki Mori, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, cs.RO
arXiv:2510.26369v2 Announce Type: replace Abstract: Logistics warehouses have struggled with labor shortages, but the inbound processes remain particularly human-powered. Worker location data is a key to higher productivity in such cases. Fixed cameras are a promising tool for localization, as they ...
152. AdamNX: An Adam improvement algorithm based on a novel exponential decay mechanism for the second-order moment estimate ​
Author: Meng Zhu, Quan Xiao, Weidong Min
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2511.13465v5 Announce Type: replace Abstract: This paper studies the exponential decay mechanism of the second-moment estimate in Adam. We propose AdamNX and a time-varying second-moment decay rate that gradually weakens the correction applied to the update scale. Under the assumptions used in...
153. gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points ​
Author: Marcus M. Noack, Mark D. Risser, Hengrui Luo, Vardaan Tekriwal, Ronald J. Pandolfi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.PR
arXiv:2512.06143v2 Announce Type: replace Abstract: Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists. This is because the vast majority of existin...
154. Replacing Tunable Parameters in Weather and Climate Models with State-Dependent Functions using Reinforcement Learning ​
Author: Pritthijit Nath, Sebastian Schemm, Henry Moss, Peter Haynes, Emily Shuckburgh, Mark J. Webb
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, physics.ao-ph
arXiv:2601.04268v3 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that are weakly constrained and tuned offline, contributing to persistent biases that limit their ability...
155. Eluder dimension: localise it! ​
Author: Alireza Bakhtiari, Alex Ayoub, Samuel Robertson, David Janz, Csaba Szepesv'ari
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2601.09825v3 Announce Type: replace Abstract: We establish a lower bound on the eluder dimension of generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds. To address this, we introduce a localisation method for the elud...
156. Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach ​
Author: Samaresh Kumar Singh, Joyjit Roy, Chirag Agrawal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.ET
arXiv:2601.17303v2 Announce Type: replace Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce latency bottlenecks that sophisticated attackers can exploit to compromise an entire manufacturing ec...
157. Automatic Stability and Recovery for Neural Network Training ​
Author: Barak Or
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.17483v2 Announce Type: replace Abstract: Training modern neural networks is increasingly fragile, with rare but severe destabilizing updates often causing irreversible divergence or silent performance degradation. Existing optimization methods primarily rely on preventive mechanisms embed...
158. Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems ​
Author: Samaresh Kumar Singh, Joyjit Roy
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.SE
arXiv:2602.04120v4 Announce Type: replace Abstract: Though Explainable AI (XAI) has made significant advancements, its inclusion in edge and IoT systems is typically ad-hoc and inefficient. Most current methods are "coupled" in such a way that they generate explanations simultaneously with model inf...
159. Layer-wise LoRA fine-tuning: a similarity metric approach ​
Author: Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2602.05988v2 Announce Type: replace Abstract: Pre-training Large Language Models (LLMs) on web-scale datasets becomes fundamental for advancing general-purpose AI. In contrast, enhancing their predictive performance on downstream tasks typically involves adapting their knowledge through fine-t...
160. Heavy-Tailed Principal Component Analysis ​
Author: Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2603.11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise. While numero...
161. Hierarchical Latent Structure Learning through Online Inference ​
Author: Ines Aitsahalia, Kiyohito Iigaya
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, q-bio.NC
arXiv:2603.19139v2 Announce Type: replace Abstract: Learning systems must balance generalization across experiences with discrimination of task-relevant details. Effective learning therefore requires representations that support both. Online latent-cause models support incremental inference but assu...
162. LLM-Extracted Covariates for Clinical Causal Inference: Rethinking Integration Strategies ​
Author: Lei Liu, Jialin Chen, Kathy Macropol
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2604.16763v3 Announce Type: replace Abstract: Causal inference from electronic health records (EHR) is fundamentally limited by unmeasured confounding: critical clinical states such as frailty, goals of care, and mental status are documented in free-text notes but absent from structured data. ...
163. Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair ​
Author: Yuelin Hu, Zhenbo Yu, Zhengxue Cheng, Wei Liu, Li Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.22407v2 Announce Type: replace Abstract: Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition has a hidden failure mode. In a high-overlap, non-adaptive 8-domain co...
164. Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics ​
Author: Chao Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2605.11017v2 Announce Type: replace Abstract: Behavioral curve modeling -- fitting parametric functions to engagement-versus-exposure data -- is standard practice in recommendation, advertising, and clinical dosing. We show that aggregation introduces a systematic distortion: Simpson's paradox...
165. DriftXpress: Faster Drifting Models via Projected RKHS Fields ​
Author: Ali Falahati, Elliot Creager, Gautam Kamath, Shubhankar Mohapatra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.12183v2 Announce Type: replace Abstract: Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference. The premise is to replace the iterative denoising process in diffusion models with a single evaluation of a...
166. Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability ​
Author: Taewoon Kim, Vincent Fran\c{c}ois-Lavet, Michael Cochez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.22142v3 Announce Type: replace Abstract: Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not explicitly model short-term-to-long-term transfer of symbolic observations. We study this transfer process in a...
167. Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO ​
Author: Yiran Xu, Yiming Ren, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong Wang, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, Ruihang Chu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.30789v3 Announce Type: replace Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, w...
168. Pretraining Recurrent Networks without Recurrence ​
Author: Akarsh Kumar, Phillip Isola
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.06479v2 Announce Type: replace Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorly: it is sequential in time, limiting parallelism, and suffers from ...
169. PostDeg: Placement Beats Parameterization in LayerNorm GNNs ​
Author: Yash Tomar, Aryav Das
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2606.14022v2 Announce Type: replace Abstract: LayerNorm-based GNNs routinely erase the topology signals (degree, centrality, $k$-core) that node-selection policies should depend on, but the literature has not located where in the residual block the erasure happens. We answer that question: a p...
170. An Empirical Study of OpenPangu Quantization on Ascend NPUs ​
Author: Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.21257v3 Announce Type: replace Abstract: OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a contr...
171. A General Framework for Learning Algebraic Properties from Cayley Graphs using Graph Neural Networks ​
Author: Tal Weissblat
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.GR
arXiv:2606.26212v2 Announce Type: replace Abstract: In this work, we present a general Graph Neural Network (GNN) framework for learning algebraic properties of finite groups from their Cayley graph representations. The framework provides a unified computational pipeline consisting of a common graph...
172. Geometry-Conditioned Fourier Neural Operators for Cubic Nonlinear Schrodinger Dynamics on Periodic Domains ​
Author: Emmanuel E. Oguadimma, Victory C. Obieke, Xueying Yu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.AP, math.NA
arXiv:2606.27459v2 Announce Type: replace Abstract: We consider the cubic nonlinear Schr"odinger (NLS) equation on two-dimensional flat tori with varying aspect ratios. In this formulation, the choice of aspect ratio governs the Fourier resonance structure, so rational and irrational geometries can...
173. A Linear Matching Bandit Approach to Online Multi-Human Multi-Robot Teaming ​
Author: Yaohui Guo, X. Jessie Yang, Cong Shi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2606.29221v2 Announce Type: replace Abstract: We address the problem of online multi-human multi-robot matching through the lens of a linear matching bandit framework, where a learner assigns robots with unknown features from a fixed pool to distinct sets of human agents over multiple rounds. ...
174. A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function ​
Author: Roberto Rossi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML
arXiv:2607.03664v3 Announce Type: replace Abstract: The Gaussian Error Linear Unit is usually motivated as the expected output of an input-dependent Bernoulli gate. This work gives an alternative interpretation: GELU is the expected output of a hard linear gate with a Gaussian random threshold. This...
175. Dissociating the Internal Representations of Sycophancy in LLMs ​
Author: Anthony Baez, Sheer Karny, Pat Pataranutaporn
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.07003v2 Announce Type: replace Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raisi...
176. The Computational Basis of Confidence in Large Language Models ​
Author: Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov, Petar Veli\v{c}kovi'c, Nathaniel Daw
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.12447v2 Announce Type: replace Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is cali...
177. RF Spectrogram Anomaly Detection with Quantum Kitchen Sinks: Architecture, Representation, and Hardware Validation ​
Author: Abdallah Aaraba, Alexis Vieloszynski, Remon Polus, Soumaya Cherkaoui, Ola Ahmad
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.13897v3 Announce Type: replace Abstract: The broadcast nature of wireless channels exposes radio-frequency (RF) networks to anomalous and malicious transmissions, making anomaly detection a fundamental requirement for secure spectrum management. Quantum Kitchen Sinks (QKS) offer a lightwe...
178. Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers ​
Author: James O' Neill, Fergal Reid
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.15456v2 Announce Type: replace Abstract: Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/...
179. Interpretable Anomaly and Drift Detection with Gaussian Mixture Models ​
Author: Behnam Asadi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.16811v2 Announce Type: replace Abstract: We revisit Gaussian Mixture Models (GMMs) as a lightweight, interpretable tool for anomaly detection and, in particular, for detecting distributional drift in data streams. We make three practical choices explicit and evaluate them on seven public ...
180. FlashPDE: A Drop-In Fused Triton Operator Library for Neural PDE Solvers ​
Author: Peiyu Zang, Bosen Xie, Ruoxiang Xu, Yongqiang Cai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.MS
arXiv:2607.18020v3 Announce Type: replace Abstract: Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memory overhead and inefficient execution of grid-based PDE op...
181. Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster ​
Author: Kushal Chakrabarti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account -- more data, retrieval, or scale -- misses an auto-regressive risk residual that increases with scale: the ...
182. Now We Know? A Systematic Comparison of TerraMind and THOR ​
Author: Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.18504v2 Announce Type: replace Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specific arte...
183. Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning ​
Author: Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.18722v3 Announce Type: replace Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but the resulting staleness is an inevitable byproduct, compounded jointly by policy lag, engine delays, and mixture-of-experts routing. Fro...
184. H$^2$SD: Hybrid Hindsight Self-Distillation ​
Author: Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Xiaocheng Feng, Bing Qin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.18955v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teache...
185. Post-Training in Time Series Foundation Models: A Unifying Framework ​
Author: Shifeng Xie, Ambroise Odonnat, Zehao Xiao, Lei Zan, Malik Tiomoko, Lujia Pan, Themis Palpanas, Boris N. Oreshkin, Chenghao Liu, Keli Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.20002v2 Announce Type: replace Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention to handle domain ...
186. AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries ​
Author: Mengda Xing (CRIL, UA), Jean-Marie Lagniez (CRIL, UA), Alejandro Franco (LRCS)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.20577v2 Announce Type: replace Abstract: Physics-based simulations are essential for understanding the electrode-scale discharge behavior of lithium-ion batteries (LIBs) but suffer from prohibitive computational costs. To address this, we introduce a novel deep learning surrogate pipeline...
187. GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries ​
Author: Miguel P. Bento, Jo~ao F. Seabra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.20757v2 Announce Type: replace Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in-training by introducing a LogSumExp term to the loss that breaks the symmetries, thus selecting a...
188. Robust Asynchronous Q-Learning under Reward and State Corruption via Batching ​
Author: Sreejeet Maity, Aritra Mitra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.20822v2 Announce Type: replace Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both the reward and state observ...
189. Test-Time Scaling via Error Localization ​
Author: Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta, Deepanway Ghosal, Aravindan Raghuveer
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.21453v2 Announce Type: replace Abstract: Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn ...
190. The Role of Pseudo-labels in Self-training Linear Classifiers on High-dimensional Gaussian Mixture Data ​
Author: Takashi Takahashi
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.LG, math.ST, stat.TH
arXiv:2205.07739v4 Announce Type: replace-cross Abstract: Self-training (ST) is a simple yet effective semi-supervised learning method. However, why and how ST improves generalization performance by using potentially erroneous pseudo-labels is still not well understood. To deepen the understanding o...
191. Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection ​
Author: Xinjie Cui, Yuezun Li, Delong Zhu, Jiaran Zhou, Junyu Dong, Siwei Lyu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.CR, cs.LG
arXiv:2411.19715v4 Announce Type: replace-cross Abstract: We describe Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is non-trivial as forgery-relate...
192. Ask for More Than Bayes Optimal: A Theory of Indecisions for Selective Hypothesis Testing ​
Author: Mohamed Ndaoud, Peter Radchenko, Bradley Rava
Published: 7/27/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.ME, stat.ML, stat.TH
arXiv:2412.12807v4 Announce Type: replace-cross Abstract: Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high. Given a target accuracy, our goal is to minimize the numbe...
193. Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation ​
Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT
arXiv:2501.18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal ge...
194. PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework ​
Author: Abhineet Agarwal, Fange Xiao, Rebecca Barter, Omer Ronen, Boyu Fan, Bin Yu
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.ME, stat.TH
arXiv:2505.08784v3 Announce Type: replace-cross Abstract: As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. In this paper we introduce PCS-UQ, a framework based on the Predictability, Computability, and Stability (PCS) principle...
195. Statistical mechanics of extensive-width Bayesian neural networks near interpolation ​
Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT
arXiv:2505.24849v2 Announce Type: replace-cross Abstract: For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models and committee...
196. Fast State-Augmented Learning for Wireless Resource Allocation with Dual Variable Regression ​
Author: Yigit Berkay Uslu, Navid NaderiAlizadeh, Mark Eisen, Alejandro Ribeiro
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SP, cs.LG
arXiv:2506.18748v2 Announce Type: replace-cross Abstract: We consider resource allocation problems in multi-user wireless networks, where the goal is to optimize a network-wide utility function subject to constraints on the ergodic average performance of users. We demonstrate how a state-augmented g...
197. A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx ​
Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SE
arXiv:2508.10017v2 Announce Type: replace-cross Abstract: Federated Learning (FL) presents a groundbreaking approach for collaborative health research, allowing model training on decentralized data while safeguarding patient privacy. FL offers formal security guarantees when combined with Differenti...
198. Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators ​
Author: Sven Dummer, Tjeerd Jan Heeringa, Jos'e A. Iglesias
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.AI, cs.LG, stat.ML
arXiv:2509.26371v3 Announce Type: replace-cross Abstract: Recently, there has been growing interest in characterizing the function spaces underlying neural networks. While shallow and deep scalar-valued neural networks have been linked to scalar-valued reproducing kernel Banach spaces (RKBS), $\math...
199. Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation ​
Author: Eduardo Fernandes Montesuma, Yassir Bendou, Mike Gartrell
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2510.04602v4 Announce Type: replace-cross Abstract: Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not because as they assume access to the complete set of samples f...
200. Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation ​
Author: Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cond-mat.dis-nn, cond-mat.stat-mech, cs.IT, cs.LG, math.IT
arXiv:2510.24616v5 Announce Type: replace-cross Abstract: For four decades statistical physics has been providing a framework to analyse neural networks. A long-standing question remained on its capacity to tackle deep learning models capturing rich feature learning effects, thus going beyond the na...
201. Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning ​
Author: Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2511.02130v2 Announce Type: replace-cross Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, d...
202. Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense ​
Author: Samaresh Kumar Singh, Joyjit Roy, Sriharsha Anand Pushkala
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC, cs.LG
arXiv:2512.23849v2 Announce Type: replace-cross Abstract: Sophisticated attackers can evade detection-based security by using encryption, stealth tactics, and low-rate attack patterns. This challenge is particularly acute in Internet of Things (IoT) and edge environments, where limited resources mak...
203. Atlas 2 -- Foundation models for clinical deployment ​
Author: Maximilian Alber, Timo Milbich, Alexandra Carpen-Amarie, Stephan Tietz, Jonas Dippel, Lukas Muttenthaler, Beatriz Perez Cancer, Alessandro Benetti, Panos Korfiatis, Elias Eulig, J'er^ome L"uscher, Jiasen Wu, Sayed Abid Hashimi, Gabriel Dernbach, Simon Schallenberg, Neelay Shah, Moritz Kr"ugener, Aniruddh Jammoria, Jake Matras, Patrick Duffy, Matt Redlon, Philipp Jurmeister, David Horst, Lukas Ruff, Klaus-Robert M"uller, Frederick Klauschen, Andrew Norgan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2601.05148v2 Announce Type: replace-cross Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this r...
204. On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization ​
Author: Sharan Sahu, Cameron J. Hogan, Martin T. Wells
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.OC
arXiv:2601.12238v5 Announce Type: replace-cross Abstract: In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. Our finite-...
205. Cross-reality location privacy protection in 6G-enabled vehicular metaverses: an LLM-enhanced hybrid generative diffusion model-based approach ​
Author: Xiaofeng Luo, Jiayi He, Jiawen Kang, Ruichen Zhang, Zhaoshui He, Ekram Hossain, Dong In Kim
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NI, cs.CR, cs.HC, cs.LG
arXiv:2601.12311v2 Announce Type: replace-cross Abstract: The emergence of 6G-enabled vehicular metaverses enables Autonomous Vehicles (AVs) to operate across physical and virtual spaces through space-air-ground-sea integrated networks. The AVs can deploy AI agents powered by large AI models as pers...
206. Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition ​
Author: Rui-Juan Jing, Yuegang Zhao, Changbo Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SC, cs.LG
arXiv:2601.13731v2 Announce Type: replace-cross Abstract: Symbolic computation, powered by modern computer algebra systems, has important applications in mathematical reasoning through exact deep computations. The efficiency of symbolic computation is largely constrained by such deep computations in...
207. Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning ​
Author: Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
Published: 7/27/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG
arXiv:2601.17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observ...
208. LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics ​
Author: Anselm Strohmaier, Samira B"odefeld, Oliver Straser, Frank Reinhold
Published: 7/27/2026, 4:00:00 AM
Categories: math.HO, cs.LG
arXiv:2601.18685v4 Announce Type: replace-cross Abstract: The capabilities of generative AI in mathematics education are rapidly evolving, posing significant challenges for research to keep pace. Research syntheses remain scarce and risk being outdated by the time of publication. To address this iss...
209. The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis ​
Author: Soroosh Tayebi Arasteh, Mina Farajiamiri, Mahshad Lotfinia, Behrus Hinrichs-Puladi, Jonas Bienzeisler, Mohamed Alhaskir, Mirabela Rusu, Christiane Kuhl, Sven Nebelung, Daniel Truhn
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2601.19618v2 Announce Type: replace-cross Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is the strongest known remedy. Practice increasingly favors large generic self-supervised encoders...
210. Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases ​
Author: Vid Kocijan, Jinu Sunil, Jan Eric Lenssen, Viman Deb, Xinwei Xe, Federico Reyes Gomez, Matthias Fey, Jure Leskovec
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG
arXiv:2602.09572v3 Announce Type: replace-cross Abstract: The purpose of predictive modeling on relational data is to predict future or missing values in a relational database, for example, future purchases of a user, risk of readmission of the patient, or the likelihood that a financial transaction...
211. Statistical Early Stopping for Reasoning Models ​
Author: Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Edgar Dobriban
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2602.13935v3 Announce Type: replace-cross Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically pr...
212. The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games ​
Author: Nikolaos Al. Papadopoulos, Ismael Tito Freire, Marti Sanchez-Fibla, Konstantinos E. Psannis
Published: 7/27/2026, 4:00:00 AM
Categories: cs.MA, cs.GT, cs.LG
arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization. Conventional outcome-based fairness measures can assign similar aggregate scores to temporally disti...
213. Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance ​
Author: Hangyi Zhao
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.GT, cs.LG
arXiv:2603.06851v3 Announce Type: replace-cross Abstract: We study contextual bilateral trade under full feedback when, conditionally on the context, trader valuations have bounded density but infinite variance. We first extend the self-bounding property of Bachoc et al. (ICML 2025) from bounded to ...
214. Minimum Norm Interpolation via the Local Theory of Banach Spaces: The Role of $2$-Uniform Convexity ​
Author: Gil Kur, Pierre Bizeul
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.LG, math.MG, math.PR, math.ST, stat.TH
arXiv:2603.28956v2 Announce Type: replace-cross Abstract: The minimum-norm interpolator (MNI) framework has recently attracted considerable attention as a tool for understanding generalization in overparameterized models, such as neural networks. In this work, we study the MNI under a $2$-uniform co...
215. Parameterized Quantum Circuits as Feature Maps: Representation Quality and Readout Effects in Multispectral Land-Cover Classification ​
Author: Ralntion Komini, Aikaterini Mandilara, Georgios Maragkopoulos, Dimitris Syvridis
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.LG
arXiv:2604.26675v2 Announce Type: replace-cross Abstract: We investigate variational quantum classifiers (VQCs) for land-cover classification from multispectral satellite imagery, adopting a feature-map perspective in which the quantum circuit defines a nonlinear data embedding while the readout det...
216. Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs ​
Author: Naomi Esposito, Anthony Tricarico, Luisa Porzio, Ali Aghazadeh Ardebili, Massimo Stella
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC, cs.LG, cs.SI
arXiv:2604.27618v2 Announce Type: replace-cross Abstract: Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason a...
217. SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation ​
Author: Jingxi Qiu, Zeyu Han, Cheng Huang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.IR, cs.LG
arXiv:2605.03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer. We study evidence sufficiency verification for selective RAG ans...
218. Scalable Gaussian process inference via neural feature maps ​
Author: Anthony Stephenson
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2605.10285v2 Announce Type: replace-cross Abstract: We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels. We show that the learned feature map can be interpreted as an optimal low-rank approximation to a Gram matrix d...
219. Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform ​
Author: Oliver Hennh"ofer, Maximilian Kirsch, Christine Preisach
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.CO
arXiv:2605.13642v2 Announce Type: replace-cross Abstract: Most anomaly detection systems output scores rather than calibrated decisions, leaving practitioners to choose thresholds heuristically and without clear statistical interpretation. Conformal anomaly detection addresses this limitation by con...
220. Approximation and learning of anisotropic and mixed smooth functions by deep ReLU neural networks ​
Author: Yunfei Yang, Jun Fan
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA
arXiv:2605.31152v2 Announce Type: replace-cross Abstract: This paper studies how efficiently deep ReLU neural networks can approximate and learn smooth functions. When the error is measured in $L^p([0,1]^d)$ norm and the approximator is a network with width $W$ and depth $L$, recent works have prove...
221. Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017 ​
Author: Zach Moczkodan (Royal Military College of Canada, Kingston, Canada), Hany Ragab (Royal Military College of Canada, Kingston, Canada)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.LG
arXiv:2606.11098v2 Announce Type: replace-cross Abstract: Recent deep learning approaches for network intrusion detection increasingly incorporate temporal architectures such as recurrent networks and Transformers, often reporting near-perfect performance on CIC-IDS2017. However, many existing studi...
222. Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks ​
Author: Greg Ongie, Rahul Parhi
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.LG, math.OC, stat.ML
arXiv:2606.14954v4 Announce Type: replace-cross Abstract: We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its representation cost and native function space, prove existenc...
223. Local Multimodal Music Alignment from Global Supervision ​
Author: Irmak Bukey, Zachary Novack, Jongmin Jung, Dasaem Jeong, Chris Donahue
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.LG, cs.MM
arXiv:2607.10023v2 Announce Type: replace-cross Abstract: Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain-in pract...
224. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs ​
Author: Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2607.14186v5 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, eac...
225. Operator-Informed Gaussian Processes for Complex Helmholtz Wavefields: From Synthetic Benchmarks to In Vivo Brain Elastography ​
Author: Boyuan Deng, Kshitiz Upadhyay, Michael Shields
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA, physics.med-ph
arXiv:2607.14193v2 Announce Type: replace-cross Abstract: The Helmholtz equation governs time-harmonic wave propagation, and in dissipative media a complex modulus renders its squared wavenumber $\kappa^2$ complex. Inferring such fields from sparse, noisy data calls for solvers that also quantify th...
226. It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability ​
Author: Carson Rodrigues
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.16292v3 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a downstream ...
227. Entanglement geometry separates circuit cutting, classical hardness, and trainability ​
Author: Maria Gragera Garces, Sabina Dr\u{a}goi, Lirand"e Pira
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.DC, cs.ET, cs.LG
arXiv:2607.17872v2 Announce Type: replace-cross Abstract: Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability. We show that these properties are strongly constrained...
228. Multi-Mask Diffusion Language Models for Few-Step Generation ​
Author: Sijin Chen, Yinuo Ren, Heyang Zhao, Ziheng Cheng, Quanquan Gu, Lexing Ying
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.19686v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDMs, all forward trajectories collapse to a single fully masked state, leaving no terminal e...
229. DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making ​
Author: Raffi Khatchadourian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.20491v2 Announce Type: replace-cross Abstract: A financial AI agent can repeat a decision while changing the tools, order, or recorded arguments and results used to reach it. Outcome-only evaluation misses this variation, even when it matters for replay and change control. DFAH-Bench oper...
230. The Geometry of Personality: Activation Steering with Jungian Cognitive Functions ​
Author: Liu Zai (University of Glasgow), Yumeng Wang (Leiden University), Junchen Fu (University of Glasgow), Joemon M. Jose (University of Glasgow)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.20803v2 Announce Type: replace-cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled ...