Skip to content

arXiv cs.AI - 2026-07-14 ​

485 items collected.


1. From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation ​

Author: Anca Marginean, Adrian Groza
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.09664v1 Announce Type: new Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentation. This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing. Consider a cl...

📖 Read original article


2. Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking ​

Author: Deep Pankajbhai Mehta
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.09665v1 Announce Type: new Abstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FS...

📖 Read original article


3. Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent ​

Author: Zayx Shawn
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.09678v1 Announce Type: new Abstract: When LLM agents hand off information to one another, does the message format matter? Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction work finds that imposin...

📖 Read original article


4. Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes ​

Author: Yossi Eliaz
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, math.PR, math.ST, stat.TH

arXiv:2607.09689v1 Announce Type: new Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs--Boltzmann measure $\exp{-\beta E(\theta)}$ whose inverse temperature is the sample size, $\beta=n$. Three consequences...

📖 Read original article


5. Interpreting Latent CoT Reasoning as Dynamical Systems ​

Author: Sabari Iyyappan Duraipandian, Shreya Sanjay Boyane, Manju Nagesh, Jerome Francis, Archana Vaidheeswaran, Kevin Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.09698v1 Announce Type: new Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate traces in the hidden space at each step, unlike explicit- CoT, which follows a single transparent reas...

📖 Read original article


6. YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate ​

Author: Suyash Mishra
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG

arXiv:2607.09706v1 Announce Type: new Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once. For decisions that allocate real budget, effort, ...

📖 Read original article


7. GES-TSP: Graph Edge Sparsification for TSP ​

Author: Tianfeng Chen, Xianyue Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.CO

arXiv:2607.09708v1 Announce Type: new Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph sparsification methods to improve computational efficiency. Traditional sparsification methods typically rely on ...

📖 Read original article


8. The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation ​

Author: Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.09709v1 Announce Type: new Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministic, judge-free, ungameable filter -- whether a generated project launches...

📖 Read original article


9. Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction ​

Author: Yuchen Wang, Javal Vyas, Tong Liu, Mehmet Mercangoz
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.RO

arXiv:2607.09713v1 Announce Type: new Abstract: A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specifications, with minimal or no manual redesign. In this setting, policy generation by AI agents can be a c...

📖 Read original article


10. Feedback-Coupled Memory Systems in Continuous Time ​

Author: Stefano Grassi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.09714v1 Announce Type: new Abstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent update operator $f_i$ and the environmental update operator $\Psi$ - are left axiomatically undefined ...

📖 Read original article


11. AGM-like Paraconsistent Partial Meet Abductive Expansion Operation ​

Author: Ulisses Franceschi Eliano
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.09729v1 Announce Type: new Abstract: In his 1996 doctoral thesis, Maurice Pagnucco created the first AGM-like abductive expansion operation. Taking his operation as a basis, as well as a taxonomy -- inspired by Atocha Aliseda -- responsible for highlighting and formalizing the main compon...

📖 Read original article


12. Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks ​

Author: Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.09739v1 Announce Type: new Abstract: We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset select...

📖 Read original article


13. A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles ​

Author: Joshua Kofi Asamoah, Blessing Agyei Kyem, Eugene Denteh, Armstrong Aboah
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.09740v1 Announce Type: new Abstract: Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surrounding traffic scene is likely to evolve. However, many existing lane-change prediction methods remain centered on a s...

📖 Read original article


14. Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets ​

Author: Pratyush Singh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.GT

arXiv:2607.09743v1 Announce Type: new Abstract: We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture. Using Hotelling's linear city model as a diagnostic vehicle, we evaluat...

📖 Read original article


15. A Theory of Least Autonomy in AI ​

Author: Christophe Parisel
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.09744v1 Announce Type: new Abstract: Least privilege, the principle that an identity should hold only the permissions strictly required for its task, has been a foundational primitive of access control for decades. We argue that this principle is insufficient for agentic AI systems, which...

📖 Read original article


16. SupplyNetPy: An Open-Source Python Library for High-Fidelity Modeling and Simulation of Arbitrary Supply Chain and Inventory Networks ​

Author: Tushar Lone, Neha Karanjkar
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.09745v1 Announce Type: new Abstract: This paper introduces SupplyNetPy, an open-source, well-documented Python library for modeling and discrete-event simulation of supply chain networks with arbitrary multi-echelon structures. It supports multiple replenishment policies, perishable inven...

📖 Read original article


17. Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems ​

Author: Jun He, Deying Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LO, cs.MA

arXiv:2607.09748v1 Announce Type: new Abstract: In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where autonomous, stochas...

📖 Read original article


18. Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks ​

Author: Hamid Ebrahimy, Moritz Lucas, Martin Atzmueller
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.09751v1 Announce Type: new Abstract: Machine Learning (ML) algorithms have been widely used to estimate agricultural variables across diverse contexts. However, because the quantity and quality of training data strongly influence performance of ML algorithms, their use can be constrained ...

📖 Read original article


19. LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules ​

Author: Tanghui Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09755v1 Announce Type: new Abstract: Urban rail fare systems may be non-additive: the fare of a single paid journey from an origin to a destination can differ from the sum of fares over multiple legally separated journey legs. This paper presents LegalFarePlan, a fare-transparent route-pl...

📖 Read original article


20. BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking ​

Author: Tianwen Zhu, Hao Wang, Yonggang Wen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2607.09762v1 Announce Type: new Abstract: Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications. Current curation remains ...

📖 Read original article


21. How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm ​

Author: Igor Itkin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA, cs.SY, eess.SY, math.OC

arXiv:2607.09765v1 Announce Type: new Abstract: A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much one must spend, and where to place the oracles. We model the swarm as a consensus on a graph in which each oracle p...

📖 Read original article


22. Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems ​

Author: Yaowen Ye, Jacob Steinhardt
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2607.09766v1 Announce Type: new Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may ...

📖 Read original article


23. Verification of Adaptive Agentic Controllers through Finite Rule Revision ​

Author: Roberto Garrone
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SY, eess.SY

arXiv:2607.09770v1 Announce Type: new Abstract: Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production deployment. In particular, adaptive agents may generate plausible outputs while remaining difficult to verify under non-determinism, confidentiality co...

📖 Read original article


24. EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents ​

Author: Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover t...

📖 Read original article


25. From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction ​

Author: Shengyi Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.09781v1 Announce Type: new Abstract: We present a pipeline for constructing maze structures from input patterns such as text or shapes. The central path-synthesis problem is encoded in Satisfiability Modulo Theories as global constraints on adjacency, continuity, and pattern-constrained c...

📖 Read original article


26. Length Penalties Make Chain-of-Thought Less Monitorable ​

Author: Bryce Little
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.09786v1 Announce Type: new Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though...

📖 Read original article


27. PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language ​

Author: Xianglin Ji, Svetlana V. Boriskina
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci

arXiv:2607.09789v1 Announce Type: new Abstract: We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench comprises 282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit), synt...

📖 Read original article


28. Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems ​

Author: M. L. Kaluzhsky, V. A. Efirov
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2607.09790v1 Announce Type: new Abstract: The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid human-machine decision support systems (DSS) of a new generation. Based on a two-month continuous longitudinal experi...

📖 Read original article


29. Agentic Context Learning with Self-Discovered Specification ​

Author: Jike Zhong, Ming Li, Yuxiang Lai, Ziyan Yang, Jingyu Xie, Jihyung Kil, Zheda Mai, Shao-Yuan Lo, Ren Xiang, Konstantinos Psounis, Yuanyuan Lei
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.09794v1 Announce Type: new Abstract: Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduct a comprehen...

📖 Read original article


30. Exploring Agentic Workflows for Generating High Quality Math Visual Aids ​

Author: Rizwaan Malik, Ashna Khetan, Isabel Sieh, Samin Khan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.HC

arXiv:2607.09839v1 Announce Type: new Abstract: Mathematical diagrams play a crucial role in K 12 education, both as problem components and as scaffolding for student comprehension. However, current AI tools, including Large Language Models (LLMs), struggle to reliably generate accurate and pedagogi...

📖 Read original article


31. TopoExplore: Topological Discrimination for Archive-Based Exploration ​

Author: Jason Carlson
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09971v1 Announce Type: new Abstract: Archive-based exploration methods such as Go-Explore select which visited state to return to using visitation rarity, and frontier methods return to the boundary of the unknown; neither asks whether the unexplored region behind a boundary is enterable ...

📖 Read original article


32. Who&When Pro: Can LLMs Really Attribute Failures in AI Agents? ​

Author: Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, Huazheng Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.09996v1 Announce Type: new Abstract: Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmar...

📖 Read original article


33. A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution ​

Author: Jose Luis Lima de Jesus Silva
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, cs.LG, cs.NE

arXiv:2607.10021v1 Announce Type: new Abstract: Neural networks can learn algorithmic input-output mappings, but trusting a learned executor requires more than a correct final answer because the state transitions that produce it are usually hidden. To make those transitions visible, we introduce a t...

📖 Read original article


34. AgentAbstain: Do LLM Agents Know When Not to Act? ​

Author: Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, confl...

📖 Read original article


35. From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability ​

Author: Cosimo Spera, Ray Garcia
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10069v1 Announce Type: new Abstract: Semantic caching defines answer reuse on embedding similarity: two utterances share a stored answer when a similarity score clears a threshold, with no notion of authorization, versioning, or of what makes two demands the same. This note changes the ob...

📖 Read original article


36. MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation ​

Author: Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10079v2 Announce Type: new Abstract: Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a fe...

📖 Read original article


37. Looped State-Space Language Models with Adaptive Exit-State Selection ​

Author: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10110v1 Announce Type: new Abstract: Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leavi...

📖 Read original article


38. Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries ​

Author: Yubo Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10113v1 Announce Type: new Abstract: Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skills}: they may be code functions, natural-language instructions, SKILL.md packages, workflow graphs, or learned ad...

📖 Read original article


39. IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation ​

Author: Hengquan Guo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10144v2 Announce Type: new Abstract: Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Yet most existing resources capture isola...

📖 Read original article


40. UNIT: Unleash Large Language Models Potential for Graph Continual Learning ​

Author: Tairan Huang, Yili Wang, Beibei Hu, Yiting Shi, Qiutong Li, Changlong He, Jianliang Gao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10159v1 Announce Type: new Abstract: In real-world multimodal web scenarios, graph-structured data often arrives in a streaming manner, making graph continual learning a crucial paradigm for continuously modeling such evolving structures. However, existing graph continual learning methods...

📖 Read original article


41. GRATE: Temporal Extensions for Inductive KG Foundation Models via Gated Rotary Attention ​

Author: Jiaxin Pan, Osama Mohammed, Daniel Hern'andez, Steffen Staab
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10197v1 Announce Type: new Abstract: Knowledge graph foundation models such as Ultra and Trix achieve strong inductive transfer by learning relation-graph representations that generalise to unseen entities and relations. Extending this transferability to temporal knowledge graphs (TKGs) r...

📖 Read original article


42. KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text ​

Author: Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10212v1 Announce Type: new Abstract: Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-s...

📖 Read original article


43. When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control ​

Author: Daming Luo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.10226v1 Announce Type: new Abstract: We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenera...

📖 Read original article


44. Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models ​

Author: Xuankun Rong, Wenke Huang, Bo Du, Dacheng Tao, Mang Ye
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10251v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities. Human decision-making combines relatively persis...

📖 Read original article


45. Information-seeking failures of large language models in agentic clinical reasoning ​

Author: Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models mu...

📖 Read original article


46. Can Agentic Trading Systems Pay for Their Own Intelligence? ​

Author: Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but ...

📖 Read original article


47. SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models ​

Author: Dongxu Zhang, Yiding Sun, Zihao Guo, Xiangyang Yang, Kai Tang, Lin Chen, Cheng Tan, Jihua Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10296v1 Announce Type: new Abstract: Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failu...

📖 Read original article


48. Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems ​

Author: Rongping Zhou, Omid Tavallaie, Shuaijun Chen, Albert Y. Zomaya
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10309v1 Announce Type: new Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly and h...

📖 Read original article


49. Comparing Socio-technical Design Principles with Guidelines for Human-centered AI ​

Author: Thomas Herrmann
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2607.10331v1 Announce Type: new Abstract: Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines with princi-ples of socio-technical systems that emerged in the context of conventional in-formation technology. ...

📖 Read original article


50. ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ​

Author: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.10350v1 Announce Type: new Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-Age...

📖 Read original article


51. Co4ICF: Co-evolving Physics-Informed Surrogate and RL-based Pulse Optimizer for Inertial Confinement Fusion ​

Author: Jiatong Zhao, Tengyue Zhang, Yuhan Wang, Fuyuan Wu, Junchi Yan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10366v1 Announce Type: new Abstract: Offline-trained surrogates for Inertial Confinement Fusion (ICF) suffer a well-known failure mode that iterative optimizers drive inputs into out-of-distribution (OOD) regions where predictions become unreliable. Here we present Co4ICF, a co-evolving f...

📖 Read original article


52. ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm ​

Author: Kefan Song, Yanjun Qi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10455v1 Announce Type: new Abstract: Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, all with minimal human oversight. Does greater autonomy invite greater ri...

📖 Read original article


53. GRASP: GRanularity-Aware Search Policy for Agentic RAG ​

Author: Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2607.10463v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, w...

📖 Read original article


54. Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents ​

Author: Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10526v2 Announce Type: new Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting pref...

📖 Read original article


55. Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach ​

Author: Chengjun Zhang, Yang Gao, Jianna Hur, Jingjing Zhang, Sagar Samtani
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG

arXiv:2607.10534v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand...

📖 Read original article


56. AI YOU Town: Make Friends and Money with Your Digital Twin ​

Author: Yan Lin, Yuyang Dai, Jiahui Geng, Yuxia Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10539v1 Announce Type: new Abstract: Existing approaches to infer user traits and generate responses consistent with a persona rely on static prompting. They lack calibrated uncertainty, ignore sequential evidence, and drift during long interactions. We present \textbf{AI YOU}, a framewor...

📖 Read original article


57. Large language model agents accelerate inverse design of metal-organic frameworks for gas separation ​

Author: Zhaolin Hu, Hehe Fan, Wangyihan Guo, Meng Xu, Chenhao Rao, Qiwei Yang, Yi Yang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10559v1 Announce Type: new Abstract: Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneous constraints of chemical validity, separation performance, and structu...

📖 Read original article


58. CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps ​

Author: JungMin Yun, JuneHyoung Kwon, YoungBin Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10562v1 Announce Type: new Abstract: Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulner...

📖 Read original article


59. Laguerre Geometry for Interpreting Large Language Models ​

Author: Chunwei Ma, Russell Wolfinger
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10578v1 Announce Type: new Abstract: Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge. Here, we show that concept geometry can be precisely characterized via Laguerre G...

📖 Read original article


60. Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification ​

Author: Siyu Wang, Wei Tan, Lulu Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10588v1 Announce Type: new Abstract: Tasks such as customs tariff classification, export control categorization, and standards-based equipment coding require assigning an input instance to a fine-grained class under an explicit regulatory hierarchy. Unlike standard text classification, th...

📖 Read original article


61. MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis ​

Author: Haoran Ma, Yinfeng Yu, Liejun Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, eess.SP

arXiv:2607.10599v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion, background noise, motion blur, or imperfect transcripts, causing conventional fusion to over-trust unreliable m...

📖 Read original article


62. Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories ​

Author: Yixiong Chen, Alan Yuille
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10601v1 Announce Type: new Abstract: Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation. This recipe is simple and low-cost, but it only learns to imitate the...

📖 Read original article


63. The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory ​

Author: Yixiong Chen, Xinyi Bai, Alan Yuille
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what...

📖 Read original article


64. Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning ​

Author: Rongbo Qi, Yaqi Zhang, Shijun Yan, Xuemeng Liu, Xiangrui Cai, Chunyao Song
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10651v1 Announce Type: new Abstract: In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user preferences, and constraints such as opening hours. Effective solutions must dynamically accommodate diverse trav...

📖 Read original article


65. Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning ​

Author: Qing Lin, Mengmi Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10678v1 Announce Type: new Abstract: Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states. Despite recent advances in artificial intelligence (AI), current models remai...

📖 Read original article


66. WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs ​

Author: Mohannad Takrouri, Nicolas M. Cuadrado A., Martin Tak'a\v{c}
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2607.10720v1 Announce Type: new Abstract: The accelerating shift toward low-carbon power systems, together with the widespread adoption of behind-the-meter technologies such as rooftop solar and electric vehicles, is placing new operational and analytical demands on electricity grids. At the s...

📖 Read original article


67. Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF ​

Author: Chinmayi Dixit
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.10750v1 Announce Type: new Abstract: Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of training data for ...

📖 Read original article


68. Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems ​

Author: Yongchang Fu, Xinjie Huang, Chengjun Dai, Chengzhe Feng, Junshao Zhang, Hong Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10768v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct...

📖 Read original article


69. Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging ​

Author: Siyi Chen, Jiahe Ying, Yixuan Jia, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, Binhong Gao, Yubing Li, Tianhan Zhang, He Sun
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10789v1 Announce Type: new Abstract: Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious e...

📖 Read original article


70. STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA ​

Author: Xinkang Li, Rong Jiang, Xin Song, Ye Wang, Yue Han, Changjian Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10795v1 Announce Type: new Abstract: In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach to knowledge-intensive QA by combining retrieval with reasoning. Existing methods mainly improve open-domain multi-hop QA through reasoning paradigms, ...

📖 Read original article


71. Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning ​

Author: Sudipto Ghosh, Tanmoy Chakraborty
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.10836v1 Announce Type: new Abstract: Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We ...

📖 Read original article


72. Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health ​

Author: Asher Sprigler, Yang-Yang Feng, Iftach Amir, Jonathan E. Bogard, Todd S Braver, Yi Ding, David Kinney, Yixue Zhao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.10871v1 Announce Type: new Abstract: Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language ...

📖 Read original article


73. LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans ​

Author: Yuma Ichikawa, Yamato Arai, Kosaku Kimura, Akira Sakai, Hiromichi Kobashi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10878v1 Announce Type: new Abstract: AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior. The defining question for deployment is no longer merely what agents can d...

📖 Read original article


74. First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint) ​

Author: Christoph Benzm"uller, Daniel Kirchner
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, math.LO

arXiv:2607.10880v1 Announce Type: new Abstract: We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML) with constant-domain Kripke semantics. Three embeddings of FML into classical higher-order logic (HOL) are prov...

📖 Read original article


75. SETA: Scaling Environments for Terminal Agents ​

Author: Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru, Aznaur Aliev, Jay Rainton, Ahmed Awelkair, Zhichen Zeng, Jiajun Li, Shi Dong, Yueming Yuan, Boyuan Ma, Qizheng Zhang, Jiwei Fu, Yuzhen Mao, Wendong Fan, Ping Nie, Philip Torr, Bernard Ghanem, Changran Hu, Jonathan Lingjie Li, Urmish Thakker, Guohao Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10891v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, co...

📖 Read original article


76. Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures ​

Author: Giansalvo Cirrincione, Filippo Grassia
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, stat.ML

arXiv:2607.10896v1 Announce Type: new Abstract: Small-data inverse design is challenging in engineering informatics when observations are heterogeneous, mixed-type, and constrained by physical relations among design variables. This work proposes a topology-aware surrogate framework guided by an Incr...

📖 Read original article


77. Learning Linear Temporal Specifications from Demonstrations with Uncertainty ​

Author: Parastou Fahim, Constantino Lagoa, R^omulo Meira-G'oes
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2607.10918v1 Announce Type: new Abstract: Learning temporal logic specifications from system demonstrations is essential for tasks such as formal verification and controller synthesis, especially in safety-critical domains. Existing approaches typically assume demonstrations are correct or onl...

📖 Read original article


78. SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning ​

Author: Mingyuan Wu, Jingcheng Yang, Shengyi Qian, Xudong Wang, Jize Jiang, Qifan Wang, Aashu Singh, Khoi Pham, Fei Liu, Zhaolun Su, Zhuokai Zhao, Klara Nahrstedt, Jianyu Wang, Hanchao Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10966v1 Announce Type: new Abstract: We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and issues a binary self...

📖 Read original article


79. From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth ​

Author: Aleh Manchuliantsau
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic c...

📖 Read original article


80. QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics ​

Author: Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11019v2 Announce Type: new Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call ...

📖 Read original article


81. AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation ​

Author: Chenyang Li, Kaige Li, Zeyu Jiang, Changhao Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11063v1 Announce Type: new Abstract: Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed s...

📖 Read original article


82. Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists ​

Author: Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11079v1 Announce Type: new Abstract: Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference...

📖 Read original article


83. NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study ​

Author: Eddie Huang (NVIDIA AI Technology Center, NVIDIA Corporation), Ken Liao (NVIDIA AI Technology Center, NVIDIA Corporation), Iven Fu (NVIDIA AI Technology Center, NVIDIA Corporation), Yang-Hsien Lin (NVIDIA AI Technology Center, NVIDIA Corporation), Chao-Shun Zhan (NVIDIA AI Technology Center, NVIDIA Corporation), Andy Liao (NVIDIA AI Technology Center, NVIDIA Corporation), Virginia Chen (NVIDIA AI Technology Center, NVIDIA Corporation), Johnson Sun (NVIDIA AI Technology Center, NVIDIA Corporation), Pika Wang (NVIDIA AI Technology Center, NVIDIA Corporation), Richard Huang (NVIDIA AI Technology Center, NVIDIA Corporation), Jiun-Cheng Jiang (NVIDIA AI Technology Center, NVIDIA Corporation), Ting-Yuan Liu (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Master Program for Digital Health Innovation, China Medical University, Taichung, Taiwan), Hsing-Fang Lu (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Laboratory for Statistical and Translational Genetics, RIKEN Center for Integrative Medical Sciences, Yokohama, Japan), Ray Y. Lee (AI-Driven Genomic Medicine and Drug Discovery Lab, China Medical University Hospital, Taichung, Taiwan), Chi-Chou Liao (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan), Simon See (NVIDIA AI Technology Center, NVIDIA Corporation), Fuu-Jen Tsai (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Department of Medical Laboratory Science and Biotechnology, Asia University, Taichung, Taiwan)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11084v1 Announce Type: new Abstract: Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generation, or statistical analysis. However, deployment in institutional biomedical environments requires governed mech...

📖 Read original article


84. OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping ​

Author: Mohammed Ehab, Aymane El Gadarri, Vivek F. Farias, Adam Jozefiak, Ciamac C. Moallemi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11089v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase late...

📖 Read original article


85. A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery ​

Author: Prashant Devadiga, Abhishek, Adithya Mishra, Alok Singh, Amisha Sinha, Asit Desai, Gaurang Dahad, Harshit Bhushan, Mandati Pramod Reddy, Prakhar Gupta, Rupesh Patil, Siddhi Behere
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11138v1 Announce Type: new Abstract: The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options sim...

📖 Read original article


86. NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management ​

Author: Changlun Li, Peixian Ma, Qiqi Duan, Zhenyu Lin, Peineng Wu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11141v1 Announce Type: new Abstract: Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints. Current assessment practice, however, remains poorl...

📖 Read original article


87. The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation ​

Author: Chenglin Yu, Hongquan Gui, Ying Yu, Hongxia Yang, Ming Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11149v1 Announce Type: new Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework ben...

📖 Read original article


88. STAMP: Provenance-Guided Credit Assignment for Deep Search Agents ​

Author: Ke Xu, Han Xu, Xinran Chen, Yuqian Wang, Zhixuan Li, Xiaojian Liu, Changwo Wu, Jianqiang Xia, Yuchen Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11172v1 Announce Type: new Abstract: Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we c...

📖 Read original article


89. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy ​

Author: Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Cheewei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11175v1 Announce Type: new Abstract: The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, r...

📖 Read original article


90. SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL ​

Author: Bowen Lv, Xiao Liu, Yanyu Ren, Hanyu Lai, Bohao Jing, Hanchen Zhang, Yanxiao Zhao, Shuntian Yao, Jie Tang, Yuxiao Dong
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11185v1 Announce Type: new Abstract: Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scalin...

📖 Read original article


91. What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities ​

Author: Sukai Huang, Chenyuan Zhang, Fucai Ke, Zhixi Cai, Naim Rastgoo, Gholamreza Haffari, Hamid Rezatofighi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11197v1 Announce Type: new Abstract: When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variation may reflect distinct latent planning competencies rather than differen...

📖 Read original article


92. PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection ​

Author: Liming Liu, Chao Hu, Mingfei Lu, Yiwei Ge, Xingle Li, Heyuan Shi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11212v1 Announce Type: new Abstract: Relational fraud detection can exploit both label-free graph context and label-derived neighborhood evidence, but these two information sources obey different validity conditions. In particular, neighborhood risk becomes invalid when a queried node's o...

📖 Read original article


93. Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory ​

Author: Tengjiao Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11226v1 Announce Type: new Abstract: LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free reign with tools and multi-agent debate, and safety violations quickly follow. Rather than forcing...

📖 Read original article


94. Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans ​

Author: Jasmin Thelen, Oliver Wilhelm
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11263v2 Announce Type: new Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, ...

📖 Read original article


95. Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought ​

Author: Daeyeop Lee, Hwanjo Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11266v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs substantial computational costs due to over-reasoning: the generation of redundant, verbose, or irrelevant steps...

📖 Read original article


96. Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization ​

Author: Tian-Shuo Liu, Shiyuan Zhang, Zijie Geng, Haoyu Liu, Runjie Xu, Pengyuan Wang, Lei Yuan, Yang Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11307v1 Announce Type: new Abstract: Full-proof autoformalization bridges extensive mathematical proofs in natural language with formally validated reasoning, offering a pathway to elevate the ceiling of verifiable mathematical reasoning. Unlike statement-level formalization, proof autofo...

📖 Read original article


97. Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors ​

Author: El Hassane Ettifouri (Novelis Research, Paris, France), Ayoub Belfatmi (Novelis Research, Paris, France), Mahaman Sanoussi Yahaya Alassan (Novelis Research, Paris, France), Walid Dahhane (Novelis Research, Paris, France)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT

arXiv:2607.11317v1 Announce Type: new Abstract: Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered toke...

📖 Read original article


98. Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation ​

Author: Congren Dai, Danni Zhao, Enyang Liu, Michael Ching Yam, Zhancheng Guo, Siyi Gu, Wentao Yang, Bo Dai, Xiaobing Li, Maosong Sun
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11334v1 Announce Type: new Abstract: Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verificatio...

📖 Read original article


99. AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic ​

Author: Zhe Xiao, Longfei Li, Xu He, Haoying Wu, Zixing Zhang, Mingyu Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11338v1 Announce Type: new Abstract: Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and corre...

📖 Read original article


100. Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents ​

Author: Chenglin Yu, Li Yin, Ying Yu, Hongxia Yang, Ming Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.PL

arXiv:2607.11346v1 Announce Type: new Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the ...

📖 Read original article


101. From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory ​

Author: Manli Yan, Yuebin Lin, Yaowen Yu, Yong Zhao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11347v1 Announce Type: new Abstract: Neural networks increasingly guide decisions in high-stakes domains such as medical diagnosis, credit approval, and energy bidding. Audit in these settings requires case-level evidence: which training cases support an action and what outcomes they carr...

📖 Read original article


102. OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis ​

Author: Yongqian Sun, Rongchen Gao, Yu Luo, Wenwei Gu, Shenglin Zhang, Qingyi Guo, Qiuai Fu, Yaoliang Wu, Dan Pei
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.11357v1 Announce Type: new Abstract: Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience. Existing LLM-based methods improve diagnosis through agentic reasoning or knowledge augmentation, but they o...

📖 Read original article


103. StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure ​

Author: Wenyi Wu, Sibo Zhu, Kun Zhou, Aayush Salvi, Zixuan Song, Biwei Huang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2607.11388v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use. However, real-world tasks are often long-horizon and involve evolving contexts containing accumulated o...

📖 Read original article


104. Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA ​

Author: Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Weigao Sun, Junhan Shi, Lingrui Mei, Tianming Yang, Steven Hoi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11433v1 Announce Type: new Abstract: Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool tr...

📖 Read original article


105. The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning ​

Author: Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11436v1 Announce Type: new Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the inte...

📖 Read original article


106. Enhancing Query Efficiency for d-DNNF Representations Through Preprocessing ​

Author: Jean Marie Lagniez, Emmanuel Lonca
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11492v1 Announce Type: new Abstract: In this paper, we investigate preprocessing techniques aimed at improving the efficiency of accessing models of propositional formulas represented in conjunctive normal form (CNF). We focus on three fundamental tasks: uniform sampling, direct model acc...

📖 Read original article


107. Comparative Analysis of GAT and BERT for Human-Like Playtesting ​

Author: Kleio Fragkedaki, Theodoros Panagiotakopoulos, Matteo Biasielli, Hui Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11501v1 Announce Type: new Abstract: Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach involves collecting diverse user data to train predictive playtesting models that mimic player behavior. However,...

📖 Read original article


108. Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning ​

Author: Jiamian Li, Niall McShane, Attila Korik, Naomi du Bois, Karl McCreadie, Leen Jabban, Benjamin Metcalfe, "Ozg"ur \c{S}im\c{s}ek, Damien Coyle
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11530v1 Announce Type: new Abstract: Decoding continuous three-dimensional (3D) motor imagery (MI) using non-invasive electroencephalography (EEG)-based brain--computer interfaces (BCIs) remains challenging due to signal variability and residual decoding errors. Deep learning architecture...

📖 Read original article


109. HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference ​

Author: Yongqin Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11586v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces persistent expert hotness skew: a small set of hot experts continuously receives most tokens, while the remainin...

📖 Read original article


110. MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models ​

Author: Tsz Hei Fan, Choi Wing Fung, Yuxuan Wan, Shuqing Li, Michael R. Lyu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.GR

arXiv:2607.11594v1 Announce Type: new Abstract: Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both sides, each ...

📖 Read original article


111. Interaction Scaling: Grounding the Third Axis of Test-Time Compute ​

Author: Bojie Li, Noah Shi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal. Every extra token comes from the same frozen weights and the same prompt, so ...

📖 Read original article


112. Auditing the Risk Claims of Distributional Reinforcement Learning ​

Author: Hari Prasad
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML

arXiv:2607.11607v1 Announce Type: new Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measu...

📖 Read original article


113. Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns ​

Author: Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.11621v1 Announce Type: new Abstract: Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whethe...

📖 Read original article


114. Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling ​

Author: Jiangtao Han, Shoufeng Ma, Shuxian Xu, Geng Li, Shuai Ling, Ning Jia, Zhengbing He
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SI, physics.soc-ph

arXiv:2607.11632v1 Announce Type: new Abstract: Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such be...

📖 Read original article


115. Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction ​

Author: Huan Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11696v1 Announce Type: new Abstract: Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasonin...

📖 Read original article


116. Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement ​

Author: Ziv Ben-Zion, Teddy Lazebnik
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.11749v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes how recipients actually behave, and through what channel, remains unknown. Here, in a randomized crossover field experiment, 121 employe...

📖 Read original article


117. Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts ​

Author: Lea Roxanne Muth, Marian Margraf
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.08292v1 Announce Type: cross Abstract: The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance toward machine-readable compliance artifacts. The Open Security Controls Assessment Language (OSCAL) is a standa...

📖 Read original article


118. Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey ​

Author: Chengcheng Sun, Jiayun Tian, Cheng Zhai, Zhixiao Wang, Yajie Song, Xiaobin Rui, Jian Zhang, Philip S. Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SI

arXiv:2607.09666v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured data. However, there remains a lack of a systematic review about GNN-based methodologies across the en...

📖 Read original article


119. ECG-LDC: A Hardware-Efficient Low-Dimensional Computing Framework for ECG Arrhythmia Classification ​

Author: Anh Tran, Khanh Tran, Cuong Do
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2607.09680v1 Announce Type: cross Abstract: Continuous cardiac monitoring in wearable devices demands classifiers that are simultaneously accurate, energy-efficient, and deployable on resource-constrained hardware. While deep neural network approaches have demonstrated high classification accu...

📖 Read original article


120. Ablation, Statistical Inference, and Validation for KV-Cache Compression ​

Author: Paolo D'Alberto, Ashish Siarasao, Elliott Delaye, Rajeev Patwari
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT

arXiv:2607.09683v1 Announce Type: cross Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through a statistical validation methodology that separates systematic codec ...

📖 Read original article


121. SciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt ​

Author: Vrishank Sai Anand, Prathamesh Dinesh Joshi, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09684v1 Announce Type: cross Abstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Universal Differential Equations (UDEs) are most effective when structural priors reflect reliable gover...

📖 Read original article


122. Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems ​

Author: Roberto Garrone
Published: 7/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY, cs.LG

arXiv:2607.09685v1 Announce Type: cross Abstract: Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions. In adaptive socio-technical systems this assumption may fail: regulatory change can alter incentives, agents ...

📖 Read original article


123. What Context Does a Coding Agent Actually Need to Act? ​

Author: Brian Sam-Bodden
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09691v1 Announce Type: cross Abstract: A modern coding agent can hold an entire repository in its context window. Most of its reading is wasted -- and the interesting question is not how much context an agent can use, but what it actually \emph{needs}. We study that question at the moment...

📖 Read original article


124. Depth-Entropy Guided Sampling for Training-Free LLM Reasoning ​

Author: Zibin Meng, Peng Xie, Kani Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09693v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals. Recent work shows that sampling from sharpened base-...

📖 Read original article


125. Mitigating Early Training Collapse in CTR Models ​

Author: Ergun Bi\c{c}ici, Erkan \c{C}etinyama\c{c}
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09696v1 Announce Type: cross Abstract: Deep neural models for click-through rate prediction often exhibit a sharp decline in validation performance immediately after the first training epoch despite continued improvement in training loss. This instability restricts effective learning and ...

📖 Read original article


126. Model Collapse: On Recursion, Noise, and Uncharted Machine Visions ​

Author: Violaine Boutet de Monvel (LIRA, IRCAV)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.OH, cs.AI, cs.CY, cs.LG

arXiv:2607.09705v1 Announce Type: cross Abstract: Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance. Exemplifying a positive-feedback-driven failure, it produces effects such...

📖 Read original article


127. The Ramanujan Challenge For AI ​

Author: Michael Shalyt, Rotem Kalisch, Carsten Schneider, Hila Barkan, Elyasheev Leibtag, John Campbell, Shachar Weinbaum, Tali Monderer, Ashvni Narayanan, Ido Kaminer
Published: 7/14/2026, 4:00:00 AM
Categories: math.HO, cs.AI, math.NT

arXiv:2607.09721v1 Announce Type: cross Abstract: To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants. These problems are attractive for AI evaluation because they are concrete and can be checked numerically to arbitrary...

📖 Read original article


128. The Universal Language of CSI:Unifying Wireless Sensing Across Devices and Environments ​

Author: Jiayi Chen, Weiting Ou, Guangxu Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.09727v1 Announce Type: cross Abstract: WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in a Tower of Babel - fragmented into isolated silos where models are tailored to specific hardware dialects, fixe...

📖 Read original article


129. SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving ​

Author: Chengyue Wang, Bin Rao, Haicheng Liao, Bonan Wang, Chengzhong Xu, Zhenning Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA

arXiv:2607.09741v1 Announce Type: cross Abstract: Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents. However, most existing approaches are purely data-driven and lack structural priors, which limits their generaliz...

📖 Read original article


130. MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms ​

Author: Saiyang Feng, Yuanyun Zhang, Shi Li
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2607.09749v1 Announce Type: cross Abstract: Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives...

📖 Read original article


131. Data-Driven Forward and Inverse Modeling of V-Beam Thermal Sensors ​

Author: Tudor Bartha, Radu Chiorean, Adrian Groza
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.09752v1 Announce Type: cross Abstract: This paper presents a machine learning framework for data-driven inverse design of V-beam thermal sensors. The goal is to determine the optimal sensor geometry: beam inclination angle, beam length and beam width that achieves a target displacement un...

📖 Read original article


132. Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis ​

Author: Haksoo Lim, Myeongjin Lee, Wonjoon Chang, Jaesik Choi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.09753v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function. In this paper, we present a systematic, phase-aware analysis of diffusion comp...

📖 Read original article


133. Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization ​

Author: Mohammad Hosseini, Eray Erturk, Saba Hashemi, Maryam M. Shanechi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, q-bio.NC

arXiv:2607.09754v1 Announce Type: cross Abstract: Large-scale, multi-subject widefield calcium imaging provides unprecedented access to brain-wide cortical dynamics. However, the high dimensionality, complex spatiotemporal structure, and substantial task-irrelevant activity in widefield recordings h...

📖 Read original article


134. RSLoRA: Training-free Rank Allocation for LoRA via Representational Sensitivity Probing ​

Author: Jiaqi Liu, Haidong Kang, Qihui Zhao, Guo Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09757v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT); however, the conventional practice of uniform rank assignment ignores the functional heterogeneity of neural layers. Existing rank allocation methods typic...

📖 Read original article


135. ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams ​

Author: Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09759v2 Announce Type: cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted in...

📖 Read original article


136. Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting ​

Author: Fengchong Yao, Jianbing Li, Qing Liu, Qikun Liu, Kefeng Song, Haitao Li, Song Wang
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.09760v2 Announce Type: cross Abstract: Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for Internet of Things (IoT) devices, but deep models often degrade across acquisition environments. In multi-antenn...

📖 Read original article


137. Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design ​

Author: Wenhao Fan, Yuanwei Bin, Jianghan Gu, Wenfa Luo, Jiao Xiang, Yuntian Chen, Shiyi Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, physics.comp-ph

arXiv:2607.09763v1 Announce Type: cross Abstract: Engineering shape optimization faces challenges in both expert-dependent problem setup and surrogate-model reliability. In practical aerodynamic design, optimization settings such as editable regions, deformation ranges, and design-preservation const...

📖 Read original article


138. OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World ​

Author: Xiaoyun Dong, Qian Xu, Yang Lu, Yang Lou, Yung-Hui Li, Jianping Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09764v1 Announce Type: cross Abstract: The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust autonomous driving systems. A key aspect of this process involves editing agent states in both appearance and traj...

📖 Read original article


139. Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction ​

Author: Adrien Schneider (M-PSI), Kacper Zabkowski (M-PSI), Anderson Augusma (M-PSI), Fr'ed'erique Letu'e (SAM, SVH), Maria Camila Pinzon (M-PSI), Dominique Vaufreydaz (M-PSI)
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.09767v1 Announce Type: cross Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized...

📖 Read original article


140. Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-Cut Guided Pipeline ​

Author: Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Haoran Zhang, Fuxian Huang, Zhanhui Lin, Zijun Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09776v1 Announce Type: cross Abstract: When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post training are required because a single round of data cannot resolve all issues, making continuous iterations necessary to progressively address the weaknes...

📖 Read original article


141. Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models ​

Author: Sergi Masip, Alicja Dobrzeniecka, Jonathan Swinnen, Joachim Collin, Bart{\l}omiej Twardowski, Szymon {\L}ukasik, Tinne Tuytelaars
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09785v1 Announce Type: cross Abstract: Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams. This has led to the development of continual self-supe...

📖 Read original article


142. A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation ​

Author: Liuyi Wang, Kai Sheng, Zongtao He, Jinlong Li, Yongrui Qin, Haojie Dai, Xiangyi Wang, Jingwei Yang, Qingqing Yan, Chengju Liu, Qijun Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09792v1 Announce Type: cross Abstract: Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments. Vision-and-Language Nav...

📖 Read original article


143. Large Multimodal Model-Based Environment-Aware Mobility Management ​

Author: Seokhyun Jeong, Sangmok Shin, Seungnyun Kim, Jiao Wu, Byonghyo Shim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, eess.SP, math.IT

arXiv:2607.09795v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and autonomous vehicles, owing to their outstanding adaptability and reasoning abilities. Despite their huge potenti...

📖 Read original article


144. JEPA for AI-Native 6G: Predictive Representations and Open Challenges ​

Author: Sheikh Salman Hassan, Irshad A. Meer, Almoatssimbillah Saifaldawla, Yan Kyaw Tun, Mustafa Ozger, Madyan Alsenwi, Nguyen Van Huynh, Woong-Hee Lee, Cedomir Stefanovic, Mathini Sellathurai, Henk Wymeersch, Tharmalingam Ratnarajah
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NI

arXiv:2607.09798v1 Announce Type: cross Abstract: Sixth-generation (6G) networks are moving toward AI-native operation, where learning modules are embedded across the radio access network (RAN), edge, and core. This transition requires learning from limited labels, heterogeneous wireless and network...

📖 Read original article


145. Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B ​

Author: Avi-ad Avraam Buskila
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.09804v1 Announce Type: cross Abstract: Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recommending exact drug dosages, issuing definitive diagnoses, prescribing tr...

📖 Read original article


146. A Unified Model for Highly Accurate ECG-Free Dynamic Coronary Roadmapping Using Spatio-Temporal Transformers ​

Author: Saahil Islam, Sebastian Piat, Venkatesh N. Murthy, Serkan Cimen, Puneet Sharma, Andreas Maier, Florin C. Ghesu
Published: 7/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2607.09805v1 Announce Type: cross Abstract: Percutaneous Coronary Intervention (PCI) is a minimally invasive procedure used to restore coronary blood flow obstructed by atherosclerotic plaque. During PCI, repeated injections of iodine-based contrast agents are required to visualize the coronar...

📖 Read original article


147. An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery ​

Author: Dibakar Datta
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DL, cond-mat.mtrl-sci, cs.AI

arXiv:2607.09806v1 Announce Type: cross Abstract: Artificial intelligence (AI) is transforming scientific discovery, but its effectiveness is fundamentally limited by the availability of structured scientific knowledge. Although existing databases have accelerated data-driven materials research, muc...

📖 Read original article


148. TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging ​

Author: Shengzhuo Yang, Ronghao Yu, Chuanjie Lv, Linpeng Peng, Hang Yu, Jie Ren, Jiajun Lv, Yong Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.09818v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generation has driven the d...

📖 Read original article


149. Memory-Conditioned Tool Calling for Camera-First Visual Agents ​

Author: Xiaofan Wu (Chance AI), Xi Zeng (Chance AI), Miaoxia Chen (Chance AI), Peishan Chen (Chance AI), Shuyan Li (Chance AI), Jiyun Yao (Chance AI), Hanyong Zhong (Chance AI), Jiahao Zhu (Chance AI)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.HC, cs.IR

arXiv:2607.09822v1 Announce Type: cross Abstract: Recognition tells an agent what is in an image; personal memory affects what is worth looking up next. In a camera-first setting the user can send only an image, so the agent must form the lookups. We study whether personal visual memory improves age...

📖 Read original article


150. More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning ​

Author: Yi Li (TU Darmstadt), Alexandre Chapin (LIRIS), Liming Chen (LIRIS), Jan Peters (TU Darmstadt), Alap Kshirsagar (IIT Delhi, ADU)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09825v1 Announce Type: cross Abstract: Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features. Object-centric slot representations are a structured alternative: th...

📖 Read original article


151. Towards Objective Dysgraphia Detection: A Multi-Branch Deep Learning Approach for Online Handwriting Analysis ​

Author: Lydia Ouhib (LIASD), Yassine Ouzar (LIASD), Zo'e Pinseel (LIASD), St'ephane Bouilland (LIASD), Mehdi Ammi (LIASD)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.SP

arXiv:2607.09826v1 Announce Type: cross Abstract: Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluency, and legibility, often hindering academic achievement and early learning development. This motor coordination...

📖 Read original article


152. Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning ​

Author: Wenke Xia, Pei Ren, Wenbo Yu, Yizhuo Zhang, Jifan Li, Yixue Zhang, Yinuo Zhao, Qingyang Gao, Jianlong Fu, Jian Tang, Ji-Rong Wen, Zhengping Che, Di Hu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09866v1 Announce Type: cross Abstract: Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in prioritizing heterogeneous ...

📖 Read original article


153. Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data ​

Author: Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, q-bio.NC, q-bio.QM

arXiv:2607.09876v1 Announce Type: cross Abstract: Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potential to retrieve events of interest by describing them with simple text q...

📖 Read original article


154. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series ​

Author: Frank Nie, Ethan B. Liu, Yuan Zhu, Loe Yan, Wei Fan, Jindong Han
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09880v1 Announce Type: cross Abstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required f...

📖 Read original article


155. Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention ​

Author: Siddharth Pal, Viktoria Rojkova
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.NE

arXiv:2607.09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache ...

📖 Read original article


156. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection ​

Author: Aishwarya R. Fursule, Vamshi Nallaguntla, Shruti Kshirsagar, Anderson R. Avila
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.09891v1 Announce Type: cross Abstract: Audio deepfake detection models determine whether speech is genuine or artificially generated, but high overall accuracy can mask substantial performance disparities across demographic groups. In this work, we investigate gender bias in audio deepfak...

📖 Read original article


157. Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling ​

Author: Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao
Published: 7/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI

arXiv:2607.09892v1 Announce Type: cross Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid wit...

📖 Read original article


158. Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code ​

Author: Chunqiu Steven Xia, Courtney Miller
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.09902v1 Announce Type: cross Abstract: Agentic coding tools are increasingly used to make autonomous repository-level changes to real-world projects. Prior work has largely evaluated these contributions at the pre-merge stage, through outcomes such as pull request acceptance and review ef...

📖 Read original article


159. Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences ​

Author: Robert Williams
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09932v1 Announce Type: cross Abstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to hallucinate poses significant risks in this high-stakes context. This study introduces a benchmark e...

📖 Read original article


160. SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification ​

Author: Ivan Alejandro Montoya Sanchez, Anantaa Kotal, Aritran Piplai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2607.09936v1 Announce Type: cross Abstract: Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those threats first appear. Generalized zero-shot learning offers a natural solution by enabling recognition of unseen c...

📖 Read original article


161. A Knowledge-Based Multi-Agent Framework for Security Control Recommendation ​

Author: Carolina Fern'andez-Mart'inez, Shuaib Siddiqui, Vanesa Daza
Published: 7/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CR, cs.LG, cs.MA

arXiv:2607.09954v1 Announce Type: cross Abstract: Hardening IT on-premises environments can be a daunting task for teams without access to adequate cybersecurity expertise. In this regard, Decision Support Systems (DSS) with embedded expert knowledge can assist users by guiding them with security re...

📖 Read original article


162. A Foundation Model for Multimodal Event Sequences in Financial Applications ​

Author: Nikita Rusakov, Vladislav Meshkov, Konstantin Zorin, Gleb Zaripov, Alexander Uglov, Alexey Vasilev, Anton Klenitskiy
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09955v1 Announce Type: cross Abstract: Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separate models trained on manually engineered tabular features. This task-specific approach limits reuse and makes it...

📖 Read original article


163. Workload-Driven Optimization for On-Device Real-Time Subtitle Translation ​

Author: Tsz-To Wong
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09957v1 Announce Type: cross Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inference, low latency, and privacy constraints. These conditions limit the value of optimizations designed...

📖 Read original article


164. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization ​

Author: Ethan Smith
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09967v2 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps. Adaptive optimizers such as Adam normalize upd...

📖 Read original article


165. A Production-Oriented Framework for Evaluation of SFX Generation ​

Author: M'elodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.SY, eess.AS, eess.SP, eess.SY

arXiv:2607.09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluati...

📖 Read original article


166. An LLM-powered Agentic Recommendation System for Connected TV Content Discovery ​

Author: Lei Shi, Di Wang, Harry Tran, Helsing Xu, Yuchen Lu, Dhara Ghodasara, Wilson Chaney, Xueting Liao, Jerry Yu, Huayu Ding, Mingze Gao, Shike Mei, Shuo Tang, Zhe Zhang, Jianming He, Abhishek Kumar, Haotian Wu, Hamed Firooz, Li Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.09988v1 Announce Type: cross Abstract: Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse contextual signals, such as trending topics, breaking news, cultural events, and cross-surface user activities, ...

📖 Read original article


167. Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information ​

Author: Naman Aggarwal, Jonathan P. How
Published: 7/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.MA

arXiv:2607.09993v1 Announce Type: cross Abstract: Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs, require strategies that are robust to hidden opponent types, such as a hidden goal flag, and to deception. Und...

📖 Read original article


168. Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora ​

Author: Lars Henry Berge Olsen, Pierre Lison, Martin Jullum, Mark Anderson
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10020v1 Announce Type: cross Abstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel...

📖 Read original article


169. Geometric mean-based pairwise comparison method with the reference values -- statistical approach ​

Author: Konrad Ku{\l}akowski, Jacek Szybowski
Published: 7/14/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, math.ST, stat.TH

arXiv:2607.10038v1 Announce Type: cross Abstract: For many years, the pairwise comparison method has been widely used for decision-making involving experts. The best-known example of this method is the Analytic Hierarchy Process (AHP). In this now classic approach, the weights of alternatives are ca...

📖 Read original article


170. Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation ​

Author: Dongping Liu, Aoyu Zhang, Luyao Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision, a cost-aware evaluation framework for multimodal AI agents on quantum circuit visual understanding....

📖 Read original article


171. Efficiently Adapting Spoken Language Models for the Singaporean Context ​

Author: Ng Jia Sheng Jason
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10092v1 Announce Type: cross Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. ...

📖 Read original article


172. Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference ​

Author: Jiayin Hu, Kai Yuan, Vanessa Hu, Xuetao Yin, Jianhua Li, Sean Suchter
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.AR, cs.LG

arXiv:2607.10109v1 Announce Type: cross Abstract: Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhead inherent in static inference, which processes simple and complex tokens with uniform intensity. To address th...

📖 Read original article


173. Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety ​

Author: Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.10112v1 Announce Type: cross Abstract: Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmar...

📖 Read original article


174. Cost of Reasoning in non-English Languages: A Case Study on Japanese ​

Author: Yuu Jinnai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.10114v1 Announce Type: cross Abstract: Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and ...

📖 Read original article


175. When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation ​

Author: Cheng-Ting Chou, Duc Binh Hoang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10116v1 Announce Type: cross Abstract: We study robust generalization under spurious correlations: tasks where a shortcut feature is correlated with the true label in training but anti-correlated in an adversarial held-out split. Varying the spurious ratio $r$ (the fraction of training ex...

📖 Read original article


176. A Large-Scale Dataset of MCP Implementations on GitHub ​

Author: Benny Toeppe, Amine Barrak, Emna Ksontini
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.10123v1 Announce Type: cross Abstract: The rapid emergence of the Model Context Protocol (MCP) has introduced a new standard for connecting large language models to external tools and services. Despite its rapid adoption in open-source development, systematic understanding of how MCP is i...

📖 Read original article


177. ML in a Box: Analyzing Containerization Practices in Open Source ML Projects ​

Author: Faten Jebari, Emna Ksontini, Amine Barrak, Wael Kessentini
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.10126v1 Announce Type: cross Abstract: Containerization has become increasingly essential in the machine learning (ML) domain, providing reproducibility, portability, and environment consistency. While prior studies have analyzed Dockerfile structures and best practices, none have examine...

📖 Read original article


178. GAE: Graph-Augmented Evolution for Scientific Discovery via Reinforcement Optimization ​

Author: Xuanzhou Chen, Taoli Cheng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10127v1 Announce Type: cross Abstract: Evolutionary program search guided by Large Language Models (LLMs) has emerged as a powerful paradigm for automated scientific discovery. However, current approaches are fundamentally constrained by three bottlenecks: structurally blind parent select...

📖 Read original article


179. SALT-GNN: Handling Dense Neighborhoods in Anti-Money Laundering Graphs via Statistics-Aware Attention ​

Author: Lidia Losavio, Francesco Sovrano, Dario Fenoglio, Martin Gjoreski, Marc Langheinrich
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10131v1 Announce Type: cross Abstract: Money laundering threatens financial stability and exposes institutions to penalties, motivating automated detection. Because laundering schemes often emerge through relational patterns, graph neural networks (GNNs) are increasingly used for anti-mon...

📖 Read original article


180. LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning ​

Author: Ning Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models ...

📖 Read original article


181. FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion ​

Author: Yuang Meng, Chenyang Wu, Xianshun Liu, Chun-Le Guo, Zichen Liang, Lina Lei, Jie Liang, Hui Zeng, Chongyi Li, Lei Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation. Iterative methods, exemplified by RAFT, achieve high accuracy through recurrent refinement, but remain challenged by large displacement...

📖 Read original article


182. EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation ​

Author: Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10165v1 Announce Type: cross Abstract: Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in t...

📖 Read original article


183. Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers ​

Author: Rashin Gholijani Farahani, Azam Bastanfard
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.10168v1 Announce Type: cross Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend o...

📖 Read original article


184. Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization ​

Author: Zhicheng Cai, Xinyuan Guo, Hanlin Wu, Mingxuan Wang, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10169v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identi...

📖 Read original article


185. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception ​

Author: Weichen Zhang, Shiquan Yu, Yinan Zhu, Peizhi Tang, Shilong Ji, Zhiyuan Deng, Tianyi Lyu, Haoyang Wang, Xin Zeng, Chen Gao, Yong Li, Xinlei Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10180v1 Announce Type: cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question Answering ...

📖 Read original article


186. Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices ​

Author: Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG

arXiv:2607.10183v2 Announce Type: cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Existing o...

📖 Read original article


187. PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning ​

Author: Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.10190v1 Announce Type: cross Abstract: Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and...

📖 Read original article


188. Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization ​

Author: Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that thi...

📖 Read original article


189. Instruction Set and Language for Hypergraphs ​

Author: Mario Pascual-Gonzalez, Ezequiel Lopez-Rubio
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.PL

arXiv:2607.10194v1 Announce Type: cross Abstract: We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over a compact instruction alphabet $\Sigma_{\mathrm{HG}}$. The encoding is executed by a small virtual machine comp...

📖 Read original article


190. Comparing Socially-Equitable Renewable Energy Budget Allocation MDP Policies in Mature and Emerging Economies ​

Author: Riya Kinnarkar, Mansur M. Arief, Yan Pratama Akhra, Dino Arla
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY, math.OC

arXiv:2607.10201v1 Announce Type: cross Abstract: Equitable renewable-energy planning is a sequential decision problem, but the decision variables available to a public planner differ sharply between mature and emerging economies. In the former the government largely builds generation, while in the ...

📖 Read original article


191. Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter ​

Author: Achyuthan Sivasankar
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10203v2 Announce Type: cross Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively. In autoregressive rollouts, the first assumption requires depth's per...

📖 Read original article


192. Source-Lifted Flow Matching for Intervenable Multimodal Imitation ​

Author: He Zhang, Ying Sun, Pengteng Li, Ziyang Chen, Yiren Zhao, Ziyang Rao, Weiyu Guo, Yandong Guo, Hui Xiong
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10206v1 Announce Type: cross Abstract: Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sampling may yield diverse behaviors, but users cannot directly choose a...

📖 Read original article


193. Exploratory Analysis of Deep Learning Models for Forecasting Meteorological Parameters in the Agricultural Sector ​

Author: Piotr Sikora, Sotirios Kontogiannis
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10208v1 Announce Type: cross Abstract: Accurate meteorological forecasting is essential for agricultural planning, irrigation management, and environmental decision support. This study conducts a comparative evaluation of recurrent and hybrid deep learning architectures for multivariate f...

📖 Read original article


194. PhenoEmbed: Self-Supervised Multispectral UAV Time-Series Embeddings for Individual Tree Crown Phenology ​

Author: Taimur Khan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10231v1 Announce Type: cross Abstract: Tree crowns are a challenging target for resilient AI because they are not static objects: their spectral response, internal texture, translucency, and apparent boundaries change substantially across the growing season. We develop PhenoEmbed, a self-...

📖 Read original article


195. Partial Contracts Suffice: Sound, LLM-Inferred Regression Verification ​

Author: Yiannis Charalambous, Rafael Menezes, Youcheng Sun, Lucas C. Cordeiro
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.10291v1 Announce Type: cross Abstract: Software evolves continuously, yet ensuring that a patch preserves intended behavior without re-verifying an entire codebase remains difficult. Regression verification addresses this problem, but existing techniques require expensive whole-program re...

📖 Read original article


196. Program-Synthesis-Driven Autodesign of Universal Unitary Operators ​

Author: Yifei Zhang, Dong Chen, Fan Wang, Wenrui Zhang, Yan Chen, Dingding Han, Jianmin Yuan, Xiangjin Kong, Yu-Gang Ma
Published: 7/14/2026, 4:00:00 AM
Categories: physics.optics, cs.AI

arXiv:2607.10295v1 Announce Type: cross Abstract: We demonstrate that AI-driven program synthesis can autonomously discover fundamental strategies for decomposing unitary matrices in photonic networks. By extending DreamCoder to complex-valued linear algebra, the system generates decomposition progr...

📖 Read original article


197. Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings ​

Author: Muhammad Abdullahi Said
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10312v1 Announce Type: cross Abstract: The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate effectively across diverse linguistic contexts. This paper presents our system description for the POLAR Shared Ta...

📖 Read original article


198. Neutralizing Structural Inequality in the Nigerian FinTech Sector ​

Author: Muhammad Abdullahi Said
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC

arXiv:2607.10317v1 Announce Type: cross Abstract: Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper introduces a hierarchical human-AI triage model for Point of Sale fraud detection in the Nigerian FinTech sect...

📖 Read original article


199. From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement ​

Author: Ronald Sielinski
Published: 7/14/2026, 4:00:00 AM
Categories: stat.AP, cs.AI, cs.IR

arXiv:2607.10341v1 Announce Type: cross Abstract: AI visibility measurement is comparative: practitioners want to know which domains generative search engines cite most often and whether observed differences are large enough to support decisions. Yet the industry lacks a principled way to determine ...

📖 Read original article


200. GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification ​

Author: Yash Shah, Omar Todd, Philipp Seeb"ock, Georg Langs, Ben Glocker, Raghav Mehta
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10357v1 Announce Type: cross Abstract: The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in clinical practice. Recently, a hybrid pipeline (GRC-Net) for CVD classification was proposed, which leverages a...

📖 Read original article


201. Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift ​

Author: Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10358v1 Announce Type: cross Abstract: Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer stat...

📖 Read original article


202. A Hyperbolic Neural Closure for M1 Radiation Transfer ​

Author: Bongseok Kim, Jiahao Zhang, Johannes Krotz, Dinshaw Balsara, Ryan McClarren, Guang Lin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.NA, math.NA

arXiv:2607.10364v1 Announce Type: cross Abstract: In radiation transfer simulations, an M1 method achieves substantial computational savings by replacing the full angular transport equation with a low-order moment system. Because this reduced system is not closed, a closure model is required to repr...

📖 Read original article


203. VINE: Taming Generative Control Policies for Reinforcement Learning ​

Author: Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10369v1 Announce Type: cross Abstract: Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works obs...

📖 Read original article


204. ABot-N1: Toward a General Visual Language Navigation Foundation Model ​

Author: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Yang Cai, Jingjing Ma, Shihui Su, Zixiao Tang, Linbo Zheng, Zedong Chu, Xiaolong Wu, Ziqiao Li, Mu Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2607.10383v2 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observat...

📖 Read original article


205. Structured Thoughts For Improved Reasoning And Context Pruning ​

Author: Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu, Chia-Hsuan Lee, Anirban Das, Stephen Rawls, Kartik Balasubramaniam, Sambit Sahu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10386v1 Announce Type: cross Abstract: Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured Thoughts, a framework that organizes reasoning into alternating and blo...

📖 Read original article


206. The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy ​

Author: Evropi Toulkeridou, Jiafei Li, Leonardo Lari, Panagiotis Grammatikopoulos
Published: 7/14/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2607.10388v1 Announce Type: cross Abstract: Artificial intelligence (AI) is transforming electron microscopy by enabling quantitative analysis of increasingly large and complex datasets for nanoparticle characterization. Recent advances in machine learning (ML) and deep learning (DL) have expa...

📖 Read original article


207. A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization ​

Author: Lingyun Shen, Xuejia Guo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10390v1 Announce Type: cross Abstract: Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summarization remains challenging as their length often exceeds the input limits. As the agent investment, which prov...

📖 Read original article


208. SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding ​

Author: Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10400v1 Announce Type: cross Abstract: Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modalit...

📖 Read original article


209. Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability ​

Author: Lingwei Wei, Dou Hu, Wei Zhou, Songlin Hu, Philip S. Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SI

arXiv:2607.10402v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level security challenge. When misused, LLMs create risks beyond false content generation, enabling attacks on the social c...

📖 Read original article


210. Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control ​

Author: Boyu Li, Linjie Qiu, Lin-Ping Yuan, Duotun Wang, Yue Jiang, Zeyu Wang, Hongbo Fu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.GR

arXiv:2607.10405v1 Announce Type: cross Abstract: Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generative content. We propose Spatula, a proof-of-concept system that generat...

📖 Read original article


211. Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts ​

Author: Istiaq Ahmed Fahad, Kamruzzaman Asif, Md. Nurul Ahad Tawhid
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.10411v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics. However, their reliability in this context remains poorly explored, particularly under varying prompt conditions wh...

📖 Read original article


212. Learning the Brain's Dynamics as a Port-Hamiltonian System ​

Author: Dibakar Sigdel
Published: 7/14/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2607.10439v1 Announce Type: cross Abstract: We model human motor cortex during a wrist-extension BCI task as a port-Hamiltonian system (pHS): a conservative interconnection (gyroscopic coupling between neural phasors) plus a dissipative port (power-law energy decay driven by a GNN surrogate). ...

📖 Read original article


213. Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams ​

Author: Siddharth Pal, Viktoria Rojkova
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.IR

arXiv:2607.10441v1 Announce Type: cross Abstract: Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a bounded recurrent state, keeping a key-value entry for every token, or imposing a fixed budget through a window ...

📖 Read original article


214. Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer ​

Author: Benjamin Friedman
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10461v1 Announce Type: cross Abstract: Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose convention. We ask whether a single self-supervised token derived from object geometry can replace both, and study...

📖 Read original article


215. Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards ​

Author: Pengfei Cai, Utkarsh Utkarsh, Alan Edelman, Christopher Vincent Rackauckas, Rafael Gomez-Bombarelli
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE

arXiv:2607.10474v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability conditions, and bounda...

📖 Read original article


216. ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples ​

Author: Kexin Huang, Junkang Wu, Jinda Lu, Shuo Yang, Chiyu Ma, Jiancan Wu, Xiang Wang, Xiangnan He, Guoyin Wang, Jingren Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.10481v1 Announce Type: cross Abstract: Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this work, we investigate a critical source of this instability: over-optimiza...

📖 Read original article


217. Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents ​

Author: Igor Santos-Grueiro
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.10487v1 Announce Type: cross Abstract: LLM agents can commit durable effects from authority evidence that was valid earlier in execution: a DOM snapshot, approval epoch, version witness, branch token, or worker result. We study the commit boundary at which earlier authority evidence no lo...

📖 Read original article


218. Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision] ​

Author: Kyoungmin Kim, Anastasia Ailamaki
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2607.10508v1 Announce Type: cross Abstract: LLM agents that conduct research (proposing ideas, writing and running code, analyzing results) can already carry a study from research question to figures, yet cannot be fully trusted. The same question asked twice in a row returns different answers...

📖 Read original article


219. Conditional Optimal Bridge for Riemannian Activation Steering ​

Author: Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan, Mehrdad Mahdavi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10517v1 Announce Type: cross Abstract: Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-density-ratio objective between desired and undesired activation distri...

📖 Read original article


220. Towards Autonomous and Auditable Medical Imaging Model Development ​

Author: Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10522v1 Announce Type: cross Abstract: Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task...

📖 Read original article


221. Motif: Discovering and Automating Personal Web Workflows ​

Author: Shaokang Jiang, Daye Nam
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SE

arXiv:2607.10531v1 Announce Type: cross Abstract: Recent advances in LLMs and existing work on programming by demonstration have made it possible for end users to create automations by explicitly demonstrating their behavior to LLMs. However, these approaches rely on the assumption that users know w...

📖 Read original article


222. Tool-Adaptive LLM Reranker ​

Author: Zichuan Liu, Ruijin Hua
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.10555v1 Announce Type: cross Abstract: Generative Large Language Models (LLMs) have revolutionized information retrieval, yet their strictly parametric nature frequently leads to severe factual hallucinations when confronted with complex queries beyond their epistemic boundaries. While ex...

📖 Read original article


223. When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation ​

Author: Hong Yang, Qi Yu, Travis Desell
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.10569v1 Announce Type: cross Abstract: Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters. We run the missing crossed comparison: an int...

📖 Read original article


224. Learning from Local Walks on Dynamic Graphs with Bandit Feedback ​

Author: Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.10571v1 Announce Type: cross Abstract: We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In this setting, the learner is restricted to local movement, selecting only its current node or an immediate neigh...

📖 Read original article


225. DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders ​

Author: Syed Irfan Ali Meerza, Oktay Ozturk, Amir Sadovnik, Jian Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10580v1 Announce Type: cross Abstract: AI models are increasingly trained on personal images scraped from social media and public platforms, often without consent, leading to serious privacy violations, such as unauthorized facial recognition and targeted advertising. To counter this, res...

📖 Read original article


226. MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference ​

Author: Venkatesha Matam, Keon Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10582v1 Announce Type: cross Abstract: Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents, tool outputs, and intermediate reasoning, whose key-value (KV) cache can become a major memory bottleneck. Exis...

📖 Read original article


227. WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen ​

Author: Khush Kataruka, Harshit Maurya, Anuja Vats, Murari Mandal, Kiran Raja, Praveen Kumar Chandaliya
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10610v1 Announce Type: cross Abstract: Efficient waste segregation is critical for sustainable urban management and environmental governance. Existing automated systems are limited by single-modality visual processing, insufficient contextual understanding, and weak regulatory alignment. ...

📖 Read original article


228. Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation ​

Author: Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2607.10628v1 Announce Type: cross Abstract: We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source, and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey ...

📖 Read original article


229. World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning ​

Author: Tong Nie, Yuewen Mei, Junlin He, Yihong Tang, Jian Sun, Wei Ma
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10630v1 Announce Type: cross Abstract: Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods ...

📖 Read original article


230. Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts ​

Author: Alireza Dehghan, Negin Ashrafi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10633v1 Announce Type: cross Abstract: Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-stable risk hierarchies that are largely artefacts of how the outcome was constructed. We demonstrate this using ...

📖 Read original article


231. Coverage Path Planning: Classical Foundations, Recent Advances, and Future Directions ​

Author: Zongyuan Shen, Shalabh Gupta, Shancheng Zhao, Dehua Zhou, Gao Wang, Zhongqiang Ren, Yaming Ou, Yikui Zhai, C. L. Philip Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10649v1 Announce Type: cross Abstract: Coverage path planning (CPP) is a fundamental problem in robot motion planning, whose aim is to produce robot trajectories that provide complete coverage of target workspaces while minimizing task-specific objectives such as path length, overlap, num...

📖 Read original article


232. Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting ​

Author: Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10661v1 Announce Type: cross Abstract: Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, traditional speculative decoding typically relies on auxiliary draft modules, incurring significant training and...

📖 Read original article


233. Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs ​

Author: Shubham Rao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cond-mat.mtrl-sci, cs.AI, cs.LG

arXiv:2607.10666v1 Announce Type: cross Abstract: Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly a...

📖 Read original article


234. Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications ​

Author: Nasser Giacaman, Valerio Terragni, Paul Denny, Viraj Kumar
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY

arXiv:2607.10674v1 Announce Type: cross Abstract: As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and rely on these tools to generate code, shifting emphasis from code writing to specification. Yet little is known...

📖 Read original article


235. Learning to Fine-tune Foundation Models under Resource Limitations ​

Author: Thomas Tsouparopoulos, Iordanis Koutsopoulos
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10694v1 Announce Type: cross Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device. At each time slot, a new batch of training data arrives, and the controller is faced with two options: either use the data...

📖 Read original article


236. Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification ​

Author: Haojie Huang, Zhang Ye, Linfeng Zhao, Boce Hu, Mingxi Jia, Yu Qi, Ahmed Agha, Dian Wang, Robert Platt, Robin Walters
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG

arXiv:2607.10706v1 Announce Type: cross Abstract: The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A good choice of action representation and loss function can help to add...

📖 Read original article


237. MDQEC-QAS: Meta-Decoding for Quantum Error Correction with Hardware-Aware VQC Search and Confidence-Gated Recovery ​

Author: Prashant Kumar Choudhary, Nouhaila Innan, Muhammad Shafique, Rajeev Singh
Published: 7/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.AR, cs.LG

arXiv:2607.10707v1 Announce Type: cross Abstract: We propose a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabilizer codes and noise settings, without requiring separate decoders for each configuration. The benchmark include...

📖 Read original article


238. PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference ​

Author: Chen Gu, Hui Wan, Donghui Hu, Hui Wang, Zhuoer Gu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.10709v1 Announce Type: cross Abstract: Large Language Model (LLM) services introduce a fundamental privacy challenge. Sensitive information may be inferred not only from explicit identifiers, such as names or phone numbers, but also from contextual associations among otherwise innocuous s...

📖 Read original article


239. Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud ​

Author: B'alint Gyevn'ar, Atoosa Kasirzadeh, Nihar B. Shah
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DL

arXiv:2607.10712v1 Announce Type: cross Abstract: Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, Artificial Intel...

📖 Read original article


240. A Corpus of Persuasion Techniques in Slavic Languages ​

Author: Jakub Piskorski, Dimitar Iliyanov Dimitrov, Marina Ernst, Jacek Haneczok, Micha{\l} Marci'nczuk, Arkadiusz Modzelewski, Roman Yangarber
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10715v1 Announce Type: cross Abstract: Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of persuasion techniques, focusing on Slavic languages. The corpus contains documents in Bulgarian, Polish, and Russia...

📖 Read original article


241. To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning ​

Author: Fengji Zhang, Tianyu Fan, Yuxiang Zheng, Xinyao Niu, Chengen Huang, Jacky Keung, Bei Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.10738v1 Announce Type: cross Abstract: Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks. However, we argue that current training paradigms harbor a ...

📖 Read original article


242. Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series ​

Author: HaoChong Fu, Jian Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10740v1 Announce Type: cross Abstract: The analysis of Multivariate Time Series (MTS) plays an important role in a lot of real-world practical applications, but it still remains some challenging problem about capturing multi-granularity structural patterns and suppressing noise appropriat...

📖 Read original article


243. Lightning Fast Matching Dependency Discovery with Desbordante ​

Author: Alexey Shlyonskikh, Michael Sinelnikov, Daniil Nikolaev, Yurii Litvinov, George Chernishev
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG, cs.PF

arXiv:2607.10771v1 Announce Type: cross Abstract: Matching dependency is a generalization of the functional dependency concept, which allows users to apply custom similarity functions for matching individual attributes. Matching dependencies have a wide range of applications for solving various data...

📖 Read original article


244. LSTrans: Efficient Knowledge Transfer for Lightweight and Automated ECG Classification ​

Author: Yi Zhao, Jiajun Gao, Chenyang Xu, Yuxi Zhou, Hao Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10784v1 Announce Type: cross Abstract: Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging due to high computational costs. To address this, we propose LSTrans, a lightweight hybrid model designed for e...

📖 Read original article


245. Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs ​

Author: Shrestha Datta, Hongfu Liu, Anshuman Chhabra
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10803v1 Announce Type: cross Abstract: Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability. We introduce Weight-Adjusted Gradients (WAG), a simple yet effective approach for estimating p...

📖 Read original article


246. Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation ​

Author: Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.10806v1 Announce Type: cross Abstract: Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR) -- a set of pri...

📖 Read original article


247. Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models ​

Author: Shuning Zhao, Patrick Wong, Leran Zhang, Xiaolin Hu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-fin.RM

arXiv:2607.10810v1 Announce Type: cross Abstract: Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applications their usefulness depends on rare adverse scenarios rather than typical samples. Standard generative obje...

📖 Read original article


248. Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents ​

Author: Kai Yu, Lu Chen, Hanqi Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.10811v1 Announce Type: cross Abstract: AI engineering is shifting from passive text generation by large language models (LLMs) to agent-driven task execution, creating new reliability challenges for long-horizon tasks under resource constraints and environmental uncertainty. Conventional ...

📖 Read original article


249. Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games ​

Author: Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level in...

📖 Read original article


250. Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization ​

Author: Fabrizio Marozzo, Stefano Iannicelli
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.10825v1 Announce Type: cross Abstract: Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences, and concerns. However, the scale, redundancy, and imbalance of such corpora make it challenging to analyze opin...

📖 Read original article


251. 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects ​

Author: Zhenyu Zhao, Nanshan Jia, Jihyeon Je, Yifu Tang, Alvin Chan, Michael Spedden, Michael V. Palleschi, Sui Huang, Jingshen Wang, Zeyu Zheng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR

arXiv:2607.10826v1 Announce Type: cross Abstract: Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the entire evaluation pipeline, not only the underlying vision-language m...

📖 Read original article


252. How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study ​

Author: Yunbo Lyu, David Williams, Jieke Shi, Zhensu Sun, Chao Peng, Zhou Yang, Federica Sarro, David Lo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC

arXiv:2607.10856v1 Announce Type: cross Abstract: The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how de...

📖 Read original article


253. The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions ​

Author: Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan, Mingxue Wang, John D. Kelleher
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB

arXiv:2607.10911v1 Announce Type: cross Abstract: In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We explore interactions between several NL2SQL pipeline extensions to inspire development of more lightweight mod...

📖 Read original article


254. The Singularity Space: A Generative Diffusion Framework for Signal Representation ​

Author: Eli Bar-Yosef, Amir Averbuch, Eli Turkel
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, eess.SP, math.NA

arXiv:2607.10930v1 Announce Type: cross Abstract: Generative models often represent signals as dense grids of amplitudes, blurring sharp transients that are crucial for the correctness of physical signals. We introduce Singularity Space, a generative framework that represents signals through complex...

📖 Read original article


255. Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles ​

Author: Ashiyana Abdul Majeed, Mahmoud Meribout, Neethu Joseph, Abel Kidane Haile, Mohammad Abdullah Al Faruque
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.10942v1 Announce Type: cross Abstract: Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints. However, heterogeneous edge-GPU deployment remains limited by underutil...

📖 Read original article


256. Efficient Online Proportional Sampling with Applications to Smoothed Online Learning ​

Author: Amirmahdi Mirfakhar, Maria-Florina Balcan, Hedyeh Beyhaghi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CG, cs.GT, econ.TH

arXiv:2607.10963v1 Announce Type: cross Abstract: We study the problem of efficient online proportional sampling from a high-dimensional domain under a $\sigma$-smoothed adversary, where the sampling distribution is induced by a dynamically evolving weight function defined over a sequence of piecewi...

📖 Read original article


257. CGS: Configurable Graph Summarization with Bounded Neighborhood Loss and Query Support ​

Author: Shubhadip Mitra, Sona Elza Simon, C Oswald, Arnab Bhattacharya, Arindam Pal
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DS, cs.AI, cs.LG

arXiv:2607.10969v1 Announce Type: cross Abstract: Given a large graph, how to generate a compact summary graph that is configurable by the user and supports multiple graph queries with either no loss or with high accuracy? The ever growing size of graph datasets makes the above question on graph sum...

📖 Read original article


258. EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion ​

Author: Cecilia Curreli, Florian Hofherr, Dominik Muhle, Abhishek Saroha, Riccardo Marin, Daniel Cremers
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.HC, cs.LG

arXiv:2607.10984v1 Announce Type: cross Abstract: Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiting generalization, preventing cross-dataset training, and requiring complex data retargeting. We introduce Equi...

📖 Read original article


259. MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI ​

Author: Youngung Han, Induk Um, Kyeonghun Kim, Junga Kim, Hyunsu Go, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10988v1 Announce Type: cross Abstract: Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window M...

📖 Read original article


260. Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies ​

Author: Ali Ahmadi, Hamed Rahimi, Adrien Jacquet Cretides, Marie Samson, Mahdi Khoramshahi, Mohamed Chetouani
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.10991v1 Announce Type: cross Abstract: As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust. While reinforcement learning (RL) navigation policies provide the fast inference an...

📖 Read original article


261. LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI ​

Author: Youngung Han, Hyunsu Go, Kyeonghun Kim, Induk Um, Junga Kim, Jaewon Jung, Woo Kyoung Jeong, Won Jae Lee, Pa Hong, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10992v1 Announce Type: cross Abstract: Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby a...

📖 Read original article


262. Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion ​

Author: Solvi Arnold, Rin Karashima, Tadashi Adachi, Takafumi Mochizuki, Kimitoshi Yamazaki
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11004v1 Announce Type: cross Abstract: We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible futures in visual form, and evaluates candidate plans by agreement of predicted outcomes with text-based goals...

📖 Read original article


263. Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies ​

Author: Ziheng Cheng, Xin Guo, Huy^en Pham, Yufei Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG, stat.ML

arXiv:2607.11005v1 Announce Type: cross Abstract: This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedba...

📖 Read original article


264. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception ​

Author: Mingjie Xie, Guangjun He, Dongli Xu, Youtian Lin, Hongjue Li, Pengming Feng, Jian Guan, Yue Deng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11008v1 Announce Type: cross Abstract: Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inc...

📖 Read original article


265. Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration ​

Author: Pengping Tan, Baoquan Zhao, Zhenhui Peng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.11039v1 Announce Type: cross Abstract: Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities. However, passive feed browsing creates a paradox: the authentic peer content that provides emotional grounding also triggers po...

📖 Read original article


266. BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services ​

Author: Yuzhe Guo, Mengzhou Wu, Yuan Cao, Jialei Wei, Dezhi Ran, Wei Yang, Tao Xie
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11042v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, observe failures, and iteratively revise code. This shift raises a central evaluation question: can an agentic L...

📖 Read original article


267. Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry ​

Author: Hannah VanderHoeven, Abhijnan Nath, Nikhil Krishnaswamy
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11053v1 Announce Type: cross Abstract: Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use of LLMs in collaborative and agentic pipelines raises questions about the extent to which they exhibit these pr...

📖 Read original article


268. Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video ​

Author: Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11078v1 Announce Type: cross Abstract: Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, an...

📖 Read original article


269. Controlling Motion Transfer in Diffusion Transformers via Attention Heads ​

Author: Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11081v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challengin...

📖 Read original article


270. AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP ​

Author: Aritra Mazumder, Nusrat jahan Lia
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.11098v2 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, and confi...

📖 Read original article


271. The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning ​

Author: Joyjeet Singh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11116v1 Announce Type: cross Abstract: Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hami...

📖 Read original article


272. MusicMark: A Robust Generative Watermarking Framework for Music Generation ​

Author: Seohwan Yun, Jeeyoung Yun, Yongjin Kim, Juyeon Lee, Sungwoong Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR

arXiv:2607.11117v1 Announce Type: cross Abstract: AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution. However, existing audio watermarking research has largely focused on speech, and applying speech-orien...

📖 Read original article


273. VIA: Visual Interface Agent for Robot Control ​

Author: Hengyuan Hu, Priya Sundaresan, Jensen Gao, Dorsa Sadigh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11119v1 Announce Type: cross Abstract: Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control. General-purpose foundation models (FMs) have grown remarkably capable of some of these, especially vision and reasoning. T...

📖 Read original article


274. BeatEdit: Symbolic Music Generation as Explicit Editing ​

Author: Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.11124v1 Announce Type: cross Abstract: Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven effectiv...

📖 Read original article


275. AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation ​

Author: Yi Ting Shen, Kentaroh Toyoda, Alex Leung
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.11151v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially a...

📖 Read original article


276. Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation ​

Author: Haojie Huang, Linfeng Zhao, Haotian Liu, Zhang Ye, Si-Yuan Huang, Mingxi Jia, Boce Hu, Fangzhou Lin, Yu Qi, Dian Wang, Robin Walters, Robert Platt
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.11167v1 Announce Type: cross Abstract: Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precisio...

📖 Read original article


277. RepTran: Search-Based Repair of Transformer Models ​

Author: Yuta Ishimoto, Paolo Arcaini, Fuyuki Ishikawa, Masanari Kondo, Naoyasu Ubayashi, Yasutaka Kamei
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11193v2 Announce Type: cross Abstract: To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and repaired. Among AI components, Transformer models are increasingly integrated into software systems, which mak...

📖 Read original article


278. ProgramTab: Boosting Table Reasoning of LLMs via Programmatic Paradigm ​

Author: Pei Guo, Enjie Liu, Yunzhi Tan, Mochi Gao, Jianxin Zhang, Ruichao Zhong, Juntao Li, Bo Hu, Zang Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11207v1 Announce Type: cross Abstract: Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this ta...

📖 Read original article


279. HandFlow: Fully Generative 4D Hand Recovery with Flow Matching ​

Author: Mingxi Xu, Bowen Duan, Yi Gu, Zhengyang Shen, Renjing Xu, Yutao Yue
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are...

📖 Read original article


280. DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs ​

Author: Anqi Li, Jie Zhang, Zhongqi Wang, Songkai Xue, Jiahao Wang, Shiguang Shan, Xilin Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.11228v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial asses...

📖 Read original article


281. An Empirical Study for Android-to-OpenHarmony GUI Test Migration ​

Author: Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11245v2 Announce Type: cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing GUI test cases has become a critical problem. However, current research neither proposes solutions tailored f...

📖 Read original article


282. Multi-Agent LLMs Fail to Explore Each Other ​

Author: Hyeong Kyu Choi, Jiatong Li, Wendi Li, Xin Eric Wang, Sharon Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.11250v1 Announce Type: cross Abstract: Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do so, often exh...

📖 Read original article


283. Enhancing LLMs through human feedback: a journey towards self-improvement ​

Author: Tatiana Pelc, Gila Kamhi, Asaf Avrahamy, Adi Fledel-Alon
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.11267v1 Announce Type: cross Abstract: In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. This study introduces a novel methodology for refining the performance of a primary Retrieval Augmented Generatio...

📖 Read original article


284. Towards Predictive, Aligned, and Scalable Robot Learning ​

Author: Peijun Tang, Shangjin Xie, Baifu Huang, Binyan Sun, Haotian Yang, Kuncheng Luo, Weiqi Jin, Shilin Fang, Jianan Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11270v1 Announce Type: cross Abstract: Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in l...

📖 Read original article


285. Automated Textbook Auditing with Multi-Agent LLM Systems ​

Author: Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran, Gabriel Stefan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.MA

arXiv:2607.11276v1 Announce Type: cross Abstract: Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose gramma...

📖 Read original article


286. A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation ​

Author: Pooya Mohammadi Kazaj, Leo Fridolin Weber, Wen Xie, Seyed Amir Ahmad Safavi-Naini, Anselm Stark, Giovanni Baj, Ali Mokhtari, Toshiya Yoshida, Christoph Ryffel, Taishi Okuno, Yoshihiro Akashi, Ronny R. Buechel, Thomas Pilgrim, Waldo Valenzuela, George C. M. Siontis, Xiaowei Xu, Moritz Hundertmark, Stephan Windecker, Christoph Grani, Isaac Shiri
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11287v1 Announce Type: cross Abstract: Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalability of measurements, which makes routine use impractical. Here we present a unified framework for comprehensi...

📖 Read original article


287. Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation ​

Author: Praneeth Narisetty, Shiva Nagendra Babu Kore
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA

arXiv:2607.11288v1 Announce Type: cross Abstract: We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabilities, proving them aga...

📖 Read original article


288. The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students ​

Author: Alexis Popovici, Andrei Ionascu, Adrian-Marius Dumitran
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.11292v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regar...

📖 Read original article


289. Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis ​

Author: Adrian-Marius Dumitran, Iulia-Maria Popescu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.11314v1 Announce Type: cross Abstract: The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of unive...

📖 Read original article


290. PRISM Edit: One Vector for All Temporal Answers ​

Author: Chen Huang, Qi Zheng, Ruiqin Zheng, Long Zeng, Yuantong Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11327v2 Announce Type: cross Abstract: Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should becom...

📖 Read original article


291. Fail-Aware and Explainable Test Oracle Prediction ​

Author: Yue Zhao, Binish Tanveer, Jelena Zdravkovic
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11342v1 Announce Type: cross Abstract: Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learning based methods address this challenge by automatically generating test assertions, yet even if syntactically correct, they are oft...

📖 Read original article


292. Longitudinal Multi-View Breast Cancer Risk Prediction ​

Author: Solveig Thrun, Zijun Sun, Suaiba A. Salahuddin, Kristoffer Wickstr{\o}m, Elisabeth Wetzer, Stine Hansen, Robert Jenssen, Michael Kampffmeyer
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11343v1 Announce Type: cross Abstract: Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recent deep learning methods have shown the value of longitudinal data and explicit temporal alignment. H...

📖 Read original article


293. Understanding the Impact of AI Code Assistants on Security API Usage: An Empirical Study ​

Author: Zahra Mousavi, Chadni Islam, M. Ali Babar, Alsharif Abuadbba, Kristen Moore
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11348v1 Announce Type: cross Abstract: AI code assistants are transforming software development, but their implications for software security remain a major concern, particularly in the context of security APIs. These APIs are critical for safeguarding software systems, yet their complexi...

📖 Read original article


294. Characterising AI Models for Cataloguing ​

Author: Miguel Arana-Catania, Neil Jefferies
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL, cs.IR, cs.LG

arXiv:2607.11353v1 Announce Type: cross Abstract: The creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it. This often-overlooked task requires slow and costly expert manual work. In this project, we have evaluated the appli...

📖 Read original article


295. Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points ​

Author: Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' f...

📖 Read original article


296. BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation ​

Author: Ajitesh Jamulkar, Aritra Hazra
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MM

arXiv:2607.11364v1 Announce Type: cross Abstract: Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with...

📖 Read original article


297. A Glimpse into Long-term Physical Coexistence with Intelligent Robots ​

Author: Weiqi Jin, Peijun Tang, Kuncheng Luo, Baifu Huang, Binyan Sun, Haotian Yang, Shangjin Xie, Jianan Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11377v1 Announce Type: cross Abstract: Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across ro...

📖 Read original article


298. Agentic Routing: The Harness-Native Data Flywheel ​

Author: Xinchen Liu, Hang Zhou, Yingjie Zong, Yuchuan Tian, Liuyang Song, Shuo Zhang, Yulong Li, Wei He, Mengyu Zheng, Runke Liu, Siyang Cheng, Xiang Kuang, Hailin Hu, Kai Han, Yunhe Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11399v1 Announce Type: cross Abstract: Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becoming structu...

📖 Read original article


299. Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation ​

Author: Ritu Yadav, Andrea Nascetti, Yifang Ban
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11412v1 Announce Type: cross Abstract: Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applications in urban planning, forest monitoring, and climate policy, where both accuracy and reliability are critical. ...

📖 Read original article


300. A Multimodal Dataset for Large Language Model Applications in the Energy Domain ​

Author: Costas Mylonas, Magda Foti
Published: 7/14/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2607.11459v1 Announce Type: cross Abstract: This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million n...

📖 Read original article


301. LightMem-Ego: Your AI Memory for Everyday Life ​

Author: Yijun Chen, Boyi Xiao, Yixian Zhao, Haoting Xia, Buqiang Xu, Jizhan Fang, Yanya Li, Yaqi Zheng, Xuehai Wang, Zirui Xue, Liuxin Zhang, Hui Li, Ningyu Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.HC, cs.MM

arXiv:2607.11487v1 Announce Type: cross Abstract: Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously accumulate...

📖 Read original article


302. Agentic Skill Optimization over Lie Algebroids ​

Author: Sridhar Mahadevan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.CT

arXiv:2607.11493v1 Announce Type: cross Abstract: Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structured artifact...

📖 Read original article


303. IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry ​

Author: Ying Yan, Liwei Hu, Xiaoming Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11497v1 Announce Type: cross Abstract: Existing generative models learn data distributions in flat Euclidean space. However, most data in our real world are manifolds embedded in high dimensional Euclidean space. Therefore, we propose an intrinsic-geometry-based generative adversarial net...

📖 Read original article


304. See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models ​

Author: Byungkun Lee, Dongyoon Hwang, Dongjin Kim, Hojoon Lee, Minho Park, Jaegul Choo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11498v1 Announce Type: cross Abstract: Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mism...

📖 Read original article


305. Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals ​

Author: Daocheng Fu, Rong Wu, Yu Yang, Xuemeng Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Yong Liu, Botian Shi, Yu Qiao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11505v1 Announce Type: cross Abstract: Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling ...

📖 Read original article


306. CDFM: Towards a General-Purpose Causal Discovery Foundation Model ​

Author: Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines. Over the past decades, numerous algorithms have been developed to tackle this challenge through ...

📖 Read original article


307. Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence ​

Author: Fernanda Miyuki Yamada, Jo~ao Paulo Gois, Hiroki Takahashi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.GR

arXiv:2607.11512v1 Announce Type: cross Abstract: As extended reality becomes more popular for social interaction and entertainment, 3D avatars must represent the full diversity of body types. Most 3D avatar systems only support normative bodies and do not accurately depict people with limb differen...

📖 Read original article


308. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos ​

Author: Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11523v1 Announce Type: cross Abstract: When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait...

📖 Read original article


309. AutoMatBench: An Automatic Optimization Toolkit for the Acceleration of Material Properties Prediction Benchmarking ​

Author: Hongxiao Li, Wanling Gao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11526v1 Announce Type: cross Abstract: Material property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimization of novel materials. In the realm of MPP, MatBench is a widely accepted benchmarking tool that defines over t...

📖 Read original article


310. Technical Report on the CVPR 2026@AdvML Workshop Challenge ​

Author: Tianyuan Zhang, Zonglei Jing, Jiangfan Liu, Ligong Zhang, Ke Ma, Chengzhi Sun, Xiaohai Xu, Zhirui Zhang, Qianqian Xu, Qingming Huang, Hanyu Fang, Junhua Liu, Zheng Wang, Xiaoliang Liu, Yuanbo Li, Shuai Gui, Bin Wang, Menghe Zheng, Jing Nie, Hanyang Meng, Zeyang Zhang, Xiang Zhang, Yongxuan Zhu, Rui Ding, Hainan Li, Yongkang Zhang, Zhilei Zhu, Xianglong Kong, Jin Hu, Zonghao Ying, Yisong Xiao, Lei Chen, Haotong Qin, Jiakai Wang, Aishan Liu, Ruikai Li, Julia Karbing, Yinpeng Dong, Zhenfei Yin, Shao Jing, Xia Hu, Jingyi Xu, Juntao Dai, Xinyun Chen, Vishal M. Patel, Xianglong Liu, Dawn Song, Alan Yuille, Philip H. S. Torr, Dacheng Tao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11560v1 Announce Type: cross Abstract: Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VL...

📖 Read original article


311. Heuristic Learning for Active Flow Control Using Coding Agents ​

Author: Paul Garnier, Jonathan Viquerat, Elie Hachem
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.flu-dyn

arXiv:2607.11565v1 Announce Type: cross Abstract: Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design particularly challenging. Deep reinforcement learning (DRL) has emerged as a powerful framework for such proble...

📖 Read original article


312. Structure-Feature Aligned Graph Learning via Alternating Constrained Optimization ​

Author: Chengcheng Yan, Qingsong Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11577v1 Announce Type: cross Abstract: We introduce a constrained two-view framework for node prediction that aligns structure-conditioned GNN embeddings with a structure-free feature prior learned by an anchor model. Conventional Graph Neural Networks (GNNs) couple feature transformation...

📖 Read original article


313. DiffEEG: A Self-Supervised Denoising Diffusion Model for Learning EEG Generic Representations ​

Author: Abdulkader Helwan, Lina Abou-Abbas, Hussein El Amouri, Belkacem Chikhaoui, Khadidja Henni
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, eess.SP

arXiv:2607.11578v1 Announce Type: cross Abstract: Deep learning for EEG-based seizure detection faces critical challenges: severe annotation scarcity and extreme class imbalance, where ictal events comprise less than 10% of clinical recordings. We present DiffEEG, a 9.6M-parameter self-supervised f...

📖 Read original article


314. Extending LLM Context via Associative Recurrent Memory ​

Author: Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recur...

📖 Read original article


315. Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model ​

Author: Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment c...

📖 Read original article


316. Closing the Loop: An Access-Control Architecture for Automated, Anomaly-Driven Network Revocation in IoT Deployments ​

Author: Muhammet Emir Korkmaz, Kemal Bicakci, Yusuf Uzunay
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.11649v1 Announce Type: cross Abstract: Network-based anomaly detection for IoT devices has matured to the point of reporting strong detection accuracy, yet most published systems stop at raising an alert and leave the question of automated enforcement to future work or to a programmable d...

📖 Read original article


317. RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM ​

Author: Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka, Oleg Sedukhin, Roman Shuvalov, Yana Dementyeva, Matvey Solovyov, Nikolay O. Nikitin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11683v1 Announce Type: cross Abstract: Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-sourc...

📖 Read original article


318. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence ​

Author: Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, Xuelong Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.11689v1 Announce Type: cross Abstract: Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly prom...

📖 Read original article


319. Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming ​

Author: Xutao Mao, Xiang Zheng, Cong Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.11698v1 Announce Type: cross Abstract: Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approa...

📖 Read original article


320. VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion ​

Author: Aastha Sharma, Guangjing Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.11706v1 Announce Type: cross Abstract: Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestima...

📖 Read original article


321. An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory ​

Author: Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare
Published: 7/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CR, cs.HC

arXiv:2607.11707v1 Announce Type: cross Abstract: Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. E...

📖 Read original article


322. Active Offline-to-Online Reinforcement Learning ​

Author: Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11720v1 Announce Type: cross Abstract: Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particular...

📖 Read original article


323. Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories ​

Author: Ziheng Zhang, Wei Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC

arXiv:2607.11725v1 Announce Type: cross Abstract: Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop. A major complication is decisive: long post-operation time-lags caused by concrete curing, water...

📖 Read original article


324. Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal ​

Author: Umm-e- Habiba, Lucas Mauser, Jonas Fritzsch, Justus Bogner, Stefan Wagner
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11771v1 Announce Type: cross Abstract: Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although prior research has proposed frameworks, patterns, and user-centered approaches to support explainability, there...

📖 Read original article


325. StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description ​

Author: Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11798v1 Announce Type: cross Abstract: Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language ...

📖 Read original article


326. Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models ​

Author: Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu, An-Yu Cheng, Hung-yi Lee
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.11801v1 Announce Type: cross Abstract: Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective...

📖 Read original article


327. Introducing Human-Centeredness in AI-Assisted Lexicography ​

Author: Antonio San Martin, Catherine Trekker
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11808v1 Announce Type: cross Abstract: This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role of lexic...

📖 Read original article


328. MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents ​

Author: Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11818v1 Announce Type: cross Abstract: We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn...

📖 Read original article


Author: Romain Amigon
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2607.11826v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive computational resources, often measured in thousands of GPU-days. In this paper, we propose a frugal and memetic NAS framework design...

📖 Read original article


330. LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments ​

Author: Divya Mereddy, Jeevan Beedareddy
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11839v1 Announce Type: cross Abstract: This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific ada...

📖 Read original article


331. Evidence-Backed Video Question Answering ​

Author: Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.11862v1 Announce Type: cross Abstract: Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse ...

📖 Read original article


332. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias ​

Author: Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.11871v1 Announce Type: cross Abstract: Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the jud...

📖 Read original article


333. A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation ​

Author: Yunhai Feng, Natalie Leung, Jiaxuan Wang, Lujie Yang, Haozhi Qi, Preston Culbertson
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.11874v1 Announce Type: cross Abstract: Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous ...

📖 Read original article


334. Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks ​

Author: Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet, Thomas Hofmann
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11875v2 Announce Type: cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized ...

📖 Read original article


335. Metacognition in LLMs: Foundations, Progress, and Opportunities ​

Author: Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11881v1 Announce Type: cross Abstract: Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI sy...

📖 Read original article


336. Measuring AI Ability to Complete Long Software Tasks ​

Author: Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Chris Painter, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2503.14499v4 Announce Type: replace Abstract: Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is ...

📖 Read original article


337. Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message ​

Author: Wei Duan, Li Qian
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2507.04673v2 Announce Type: replace Abstract: The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. However, this reliance introduces an unexplored attack surface. This paper introduces Trojan Horse Prompting, a nov...

📖 Read original article


338. LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning ​

Author: Bo Hou, Xin Tan, Kai Zheng, Fang Liu, Yinghao Zhu, Li Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2507.16395v3 Announce Type: replace Abstract: Atomic commits, which address a single development concern, are a best practice in software development. In practice, however, developers often produce tangled commits that mix unrelated changes, complicating code review and maintenance. Prior unta...

📖 Read original article


339. InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation ​

Author: Wen-Xi Yang, Tian-Fang Zhao, Guan Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2508.03174v4 Announce Type: replace Abstract: Collaborative partnerships play a crucial role in inquiry-oriented education. However, most learning partners are currently assigned through experience-driven heuristics or rule-based machine assistants, which often result in limited knowledge expa...

📖 Read original article


340. Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions ​

Author: St'ephane Aroca-Ouellette, Ian Berlot-Attwell, Panagiotis Lymperopoulos, Abhiramon Rajasekharan, Tongqi Zhu, Herin Kang, Kaheer Suleman, Sam Pasupalak
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.15830v2 Announce Type: replace Abstract: Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decision making. Practical domains, such as business management, require optimizing an open-ended and multi-facete...

📖 Read original article


341. Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs ​

Author: Yutong Song, Jiang Wu, Shaofan Yuan, Chengze Shen, Jian Wang, Yu Wang, Nikil Dutt, Amir M. Rahmani
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.06362v2 Announce Type: replace Abstract: Personalized large language models are often expected to follow explicit style instructions, yet we find that such instructions can undermine the user-specific characteristics that personalization methods aim to preserve. We call this failure mode ...

📖 Read original article


342. BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation ​

Author: Xin Guo, Rongjunchen Zhang, Guilong Lu, Xuntao Guo, Shuai Jia, Zhi Yang, Liwen Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2601.06401v2 Announce Type: replace Abstract: Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual ...

📖 Read original article


343. FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights ​

Author: Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge. Existing benchmarks face a trade-off: th...

📖 Read original article


344. JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks ​

Author: Lanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai, Yuanwu Xu, Lei Wei, Sicong Xie, Guannan Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.06486v4 Announce Type: replace Abstract: Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response strategies, while LLM-as-a-jud...

📖 Read original article


345. A Model-Free Universal AI ​

Author: Yegon Kim, Juho Lee
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.23242v4 Announce Type: replace Abstract: In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the first model-free agent proven t...

📖 Read original article


Author: Avrile Floro (UPHF), Tamara Dhorasoo (UPHF), Soline Pellez (UPHF), Nils Holzenberger
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.22973v2 Announce Type: replace Abstract: Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a concrete task: detecting implicit citations of the French Civil Code, where a court applies a statutory rule...

📖 Read original article


347. When Sensing Varies with Contexts: Context Probing for Tactile Few-Shot Class-Incremental Learning ​

Author: Yifeng Lin, Aiping Huang, Wenxi Liu, Si Wu, Tiesong Zhao, Zechao Li, Zheng-Jun Zha
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.25115v2 Announce Type: replace Abstract: Few-shot class-incremental learning (FSCIL) aims to recognize novel classes from only a few labeled samples while retaining previously learned knowledge. Although recent FSCIL methods have achieved substantial progress on visual benchmarks, they re...

📖 Read original article


348. Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents ​

Author: Hy Dang, Quang Dao, Meng Jiang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2604.00137v2 Announce Type: replace Abstract: Tool-integrated LLMs retrieve information, perform computations, and take real-world actions, but their reliability depends on both tool-use accuracy and intrinsic tool accuracy, including tool correctness, stability, and safety. While prior work p...

📖 Read original article


349. GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning ​

Author: DeepReinforce Team, Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 Deep Think, attained ...

📖 Read original article


350. Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs ​

Author: Kevin Murphy
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.18576v4 Announce Type: replace Abstract: We present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on the ForecastBench benchmark. The system is built on three ideas. (1) Linguistic belief state: a semi-structu...

📖 Read original article


351. Algorithm Selection with Zero Domain Knowledge via Text Embeddings ​

Author: Stefan Szeider
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2604.19753v2 Announce Type: replace Abstract: We propose a feature-free approach to algorithm selection: instead of hand-crafted instance features, we use pretrained text embeddings. Our method, ZeroFolio, proceeds in three steps. First, it reads the raw instance file as plain text. Second, it...

📖 Read original article


352. Ideological Bias in LLMs' Economic Causal Reasoning ​

Author: Donggyu Lee, Hyeok Yun, Jungwon Kim, Junsik Min, Sungwon Park, Sangyoon Park, Jihee Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.CL, cs.LG, econ.GN, q-fin.EC

arXiv:2604.21334v2 Announce Type: replace Abstract: Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directionally correct causal judgments are essential, t...

📖 Read original article


353. Recursive Multi-Agent Systems ​

Author: Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus J. Buehler, Jingrui He, James Zou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2604.25917v2 Announce Type: replace Abstract: Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states to deepen reasoning. We extend such scaling principle from a single model to multi-agent systems, ...

📖 Read original article


354. A Low-Latency Fraud Detection Layer for Detecting Adversarial Interaction Patterns in LLM-Powered Agents ​

Author: Sheldon Yu, Yingcheng Sun, Hanqing Guo, Qianqian Tong
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.01143v2 Announce Type: replace Abstract: Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning. However, their increasing autonomy also introduces a new attack surface: adversarial interactions can manipu...

📖 Read original article


355. 2.5-D Decomposition for LLM-Based Spatial Construction ​

Author: Paul Whitten, Li-Jen Chen, Sharath Baddam
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.07066v4 Announce Type: replace Abstract: Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLMs) make systematic coordinate errors when generating three-dimensional block placements. We present a neuro-s...

📖 Read original article


356. EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents ​

Author: Ruofei Ju, Xinrui Wang, Xin Ding, Yifan Yang, Hao Wu, Shiqi Jiang, Qianxi Zhang, Hao Wen, Xiangyu Li, Weijun Wang, Kun Li, Yunxin Liu, Haipeng Dai, Wei Wang, Ting Cao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.10332v2 Announce Type: replace Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied environments vary across layouts, object states, and other execution factors, these skills must self-e...

📖 Read original article


357. Learning Developmental Scaffoldings to Guide Self-Organisation ​

Author: Milton L. Montero, Elias Najarro, Jakob Schauser, Sebastian Risi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY, q-bio.QM

arXiv:2605.14998v4 Announce Type: replace Abstract: From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local interactions that collectively give rise to global structure without any blueprint of the outcome. Yet a significan...

📖 Read original article


358. MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention ​

Author: Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Jianlong Fu, Wen-Huang Cheng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.01063v2 Announce Type: replace Abstract: Theory of Mind (ToM) enables an agent to reason about another actor's beliefs, goals, and intentions, which is essential for human-centered embodied assistance. Existing ToM benchmarks have advanced text and multimodal mental-state recognition, but...

📖 Read original article


359. Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation ​

Author: Ahmed Alansary, Molham Mohamed, Ali Hamdi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.05510v3 Announce Type: replace Abstract: Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language models often struggle to provide consistent and contextually appropriate medical responses across varying levels...

📖 Read original article


360. When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference ​

Author: Yasushi Sakai, Allen Song, Kent Larson
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.08098v3 Announce Type: replace Abstract: Majority voting is the default unsupervised aggregator for multi-sample LLM inference, but it discards two signals: within-group answer entropy and between-group reasoning geometry. We aggregate by delegation instead (Propagational Proxy Voting, PP...

📖 Read original article


361. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation ​

Author: Kailin Lyu, Di Wu, Pengwei Zhang, Yuhang Zheng, Yingxin Lai, Long Xiao, Kangyi Wu, Pengna Li, Chen Gao, Lianyu Hu, Xiaobin Hu, Jie Hao, Ce Hao, Weihao Yuan, Shuicheng Yan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.11637v3 Announce Type: replace Abstract: Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remai...

📖 Read original article


362. NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning ​

Author: Shiyun Zhao, Xinwei Song, Tianyu Guo, Xiaomeng Gao, Mingyuan Liu, Xu Han, Yuanyuan Zhang, Zhenliang Zhang, Xue Feng, Bo Dai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.27826v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requires not only achieving instructed goals but also acting in socially appropriate ways. While explicit goals m...

📖 Read original article


363. Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem ​

Author: Yutian Tang, Yuming Zhou, Huaming Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.29116v2 Announce Type: replace Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs. LLM agents are LLM systems tha...

📖 Read original article


364. Safety from Honesty in a Disinterested AI Predictor ​

Author: Yoshua Bengio, Oliver Richardson, Tom'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.29657v2 Announce Type: replace Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Pre...

📖 Read original article


365. FARS: A Fully Automated Research System Deployed at Scale ​

Author: Qiong Tang, Tianxiang Sun, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao, Bobo Li, Changze Lv, Cheng Xu, Chengsong Huang, Chunyang Li, Dizhan Xue, Hao Bai, Haodong Duan, Hengquan Guo, Hongyang He, Hongyi Chen, Hui Shen, Jiahao Yuan, Jiankai Sun, Jikang Cheng, Jinfeng Xu, Jingqi Tong, Jingye Chen, Jinxiu Liu, Jixuan Leng, Junchi Yu, Kaixun Jiang, Kun Xiang, Kunpeng Yao, Lang Feng, Liangqi Yuan, Longsen Gao, Meng Li, Qi Jia, Qiushi Sun, Shengyuan Ding, Shizhan Gong, Siru Zhong, Terry Jingchen Zhang, Tianle Gu, Tianyi Liang, Weijie Liu, Weikai Yang, Weizhi Fei, Xin Wang, Xinpeng Liu, Xuanwen Ding, Yihong Tang, Yuanli Wang, Yukun Jiang, Yuming Yang, Zhengbao He, Zhikai Chen, Zhikun Xu, Zhuang Li, Zihao Huang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.31651v2 Announce Type: replace Abstract: Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tas...

📖 Read original article


366. Agentic generation of verifiable rules for deterministic, self-expanding reaction classification ​

Author: Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas, Helena Avila, Octavian Susanu, J'er^ome Waser, Philippe Schwaller
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.01061v3 Announce Type: replace Abstract: Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encodi...

📖 Read original article


367. Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment ​

Author: Yufan Lu, Xinhui Liu, Chenyang Xu, Yuxi Zhou, Hao Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.01674v3 Announce Type: replace Abstract: In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freezing a pretrained backbone and assigning each source an isolated classifier prevents parameter interference, b...

📖 Read original article


368. SUNTA: Hierarchical Video Prediction with Surprise-based Chunking ​

Author: Tomoshi Iiyama, Masahiro Suzuki, Yutaka Matsuo
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.02087v2 Announce Type: replace Abstract: Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how chunk boundaries are determined. While prior HSSMs typically rely o...

📖 Read original article


369. Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces ​

Author: Andrew Zhang, Chengzhan Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.04419v3 Announce Type: replace Abstract: When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable. Yet the same frozen transition can change sign with the scoring channel. Process rewards vary agent states, while evaluator ...

📖 Read original article


370. MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models ​

Author: Zhi Song, Ximing Xing, Zhenchao Tang, hanbo Huang, Weilong Yan, Tianxu Lv, Minghao Yang, Zhongzheng Niu, Bing He, Lusheng Wang, Jianhua Yao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.05238v2 Announce Type: replace Abstract: JEPA world models commonly predict the next latent state with one regressor. Under stochastic transitions, squared and cosine regression return the conditional mean and its normalized direction, respectively: a single compromise that may match no v...

📖 Read original article


371. When do prophets profit in prediction markets? ​

Author: Anri Gu, Nicole Kagan, Alec Sun, Jibang Wu, Haifeng Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.GT

arXiv:2607.06166v2 Announce Type: replace Abstract: Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automate...

📖 Read original article


372. Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents ​

Author: Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.07405v2 Announce Type: replace Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transit...

📖 Read original article


373. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading ​

Author: Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, Leowei Liang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.08964v2 Announce Type: replace Abstract: AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlo...

📖 Read original article


374. LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making ​

Author: Zihan Xu, Yanzhen Chen, Xiaocheng Zhang, Zhiting Fan, Weiqi Zhai, Hongxia Xu, Zuozhu Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09322v2 Announce Type: replace Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world ...

📖 Read original article


375. Research on Cross-media Science and Technology Information Data Retrieval ​

Author: Yang Jiang, Zhe Xue, Ang Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2204.04887v3 Announce Type: replace-cross Abstract: Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has become an integral part of people's daily life. Unlike news data and social data on the Internet, cross-medi...

📖 Read original article


376. Research on Intellectual Property Resource Profile and Evolution Law ​

Author: Yuhui Wang, Yingxia Shao, Ang Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.LG

arXiv:2204.06221v2 Announce Type: replace-cross Abstract: In the era of big data, intellectual property-oriented scientific and technological resources show the trend of large data scale, high information density, and low value density, which brings severe challenges to the effective use of intellec...

📖 Read original article


377. Profiling and Evolution of Intellectual Property ​

Author: Bowen Yu, Yingxia Shao, Ang Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2204.09333v3 Announce Type: replace-cross Abstract: In recent years, with the rapid growth of Internet data, the number and types of scientific and technological resources are also rapidly expanding. However, the increase in the number and category of information data will also increase the co...

📖 Read original article


378. HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA ​

Author: Xinyue Chen, Pengyu Gao, Jiangjiang Song, Xiaoyang Tan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2402.01767v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) has rapidly advanced the language model field, particularly in question-answering (QA) systems. By integrating external documents during the response generation phase, RAG significantly enhances the accura...

📖 Read original article


379. Constrained Reinforcement Learning for Safe Heat Pump Control ​

Author: Baohe Zhang, Lilli Frison, Thomas Brox, Joschka B"odecker
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY

arXiv:2409.19716v2 Announce Type: replace-cross Abstract: Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is crucial for enhancing safety and performance across diverse control tasks. In the context of heating s...

📖 Read original article


380. Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs ​

Author: Max Weltevrede, Caroline Horsch, Matthijs T. J. Spaan, Wendelin B"ohmer
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.03565v4 Announce Type: replace-cross Abstract: In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of contexts and must generalize to new ones. Recent work has demonstrated that training on additional states,...

📖 Read original article


381. On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes ​

Author: Rajat Modi, Vibhav Vineet, Yogesh Singh Rawat
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CY

arXiv:2410.19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic occlusions, OVIS-UCF an...

📖 Read original article


382. Asynchronous Perception Machine For Efficient Test-Time-Training ​

Author: Rajat Modi, Yogesh Singh Rawat
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2410.20535v5 Announce Type: replace-cross Abstract: In this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM can process patches of an image one at a time in any order asymmetrically and still encode semantic-awa...

📖 Read original article


383. Training-Free, Identity-Preserving Image Editing for Fashion Pose Alignment and Normalization ​

Author: Potito Aghilar, Vito Walter Anelli, Michelantonio Trizio, Eugenio Di Sciascio, Tommaso Di Noia
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.SE

arXiv:2501.13692v2 Announce Type: replace-cross Abstract: Diffusion models have recently unlocked new possibilities in editing images of real-world objects. Yet, transforming objects in non-rigid ways, such as modifying poses or applying image-based conditioning, continues to present significant cha...

📖 Read original article


384. Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers ​

Author: Zixuan Gong, Shijia Li, Yong Liu, Jiaye Teng
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2502.20681v3 Announce Type: replace-cross Abstract: Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantica...

📖 Read original article


385. Hyperflux: Pruning Reveals Importance ​

Author: Eugen Barbulescu, Antonio Alexoaie, Lucian Busoniu
Published: 7/14/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2504.05349v5 Announce Type: replace-cross Abstract: Network pruning is used to reduce inference latency and power consumption in large neural networks. However, most methods focus on empirical results at the expense of understanding the pruning process. We introduce Hyperflux, a novel $L_0$ me...

📖 Read original article


386. Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning ​

Author: Yang Xu, Swetha Ganesh, Vaneet Aggarwal
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2506.07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learning and actor-critic algorithms for robust average-reward MDPs ...

📖 Read original article


387. Adaptive Reinforcement Learning for Unobservable Random Delays ​

Author: John Wikman, Alexandre Proutiere, David Broman
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2506.14411v2 Announce Type: replace-cross Abstract: In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects a...

📖 Read original article


388. On the Necessity of Output Distribution Reweighting for Effective Class Unlearning ​

Author: Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.20893v5 Announce Type: replace-cross Abstract: In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class. We further propose a simple unlearning strategy to mitigate...

📖 Read original article


389. Can Argus Judge Them All? Comparing VLMs Across Domains ​

Author: Harsh Joshi, Gautam Siddharth Kashyap, Rafiq Ali, Ebad Shabbir, Niharika Jain, Sarthak Jain, Jiechao Gao, Usman Naseem
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2507.01042v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly used in industry VLM applications such as retrieval systems, content generation platforms, and decision-support workflows, where model selection is commonly guided by benchmark rankings. These ra...

📖 Read original article


390. Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI ​

Author: Nikhita Joshi, Daniel Vogel
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2507.03670v3 Announce Type: replace-cross Abstract: Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs to them. To encourage users to write longer prompts, we evaluated two interaction techniques that modi...

📖 Read original article


391. SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks ​

Author: Pavel Adamenko, Mikhail Ivanov, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitry Babaev, Alena Fenogenova, Valentin Malykh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2507.11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset. Recent studies have uncovered severe data contamination is...

📖 Read original article


Author: Xiaoya Li, Albert Wang, Guoyin Wang, Chris Shum, Jiwei Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.DB

arXiv:2508.02091v3 Announce Type: replace-cross Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-augmented generation (RAG) and agent-based LLM applications. In this paper, we present CRINN, a new ...

📖 Read original article


393. Beyond Na\"ive Prompting: Strategies for Improved Context-aided Forecasting with LLMs ​

Author: Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng, Irina Rish, Nicolas Chapados, 'Etienne Marcotte, Valentina Zantedeschi, Alexandre Drouin
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2508.09904v3 Announce Type: replace-cross Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting, critical challenges...

📖 Read original article


394. Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts ​

Author: Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2508.10123v3 Announce Type: replace-cross Abstract: Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT frameworks, a behavior model generates multiple completions with answe...

📖 Read original article


395. PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains ​

Author: Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema, Sohee Yang, Wai-Chung Kwan, Xuanli He, Wenda Li, Pasquale Minervini, Eleonora Giunchiglia, Shay B. Cohen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.21787v3 Announce Type: replace-cross Abstract: Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate solutions and selecting the one with the highest reward. The key challenge for reasoning tasks is desi...

📖 Read original article


396. TENET: One Step Toward Test-Driven Development for Repository-Level Code Generation ​

Author: Yiran Hu, Nan Jiang, Shanchao Liang, Yi Wu, Lin Tan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2509.24148v3 Announce Type: replace-cross Abstract: Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation. With recent advances in Large Language Models (LLMs), developers can shift from manually writing the cod...

📖 Read original article


397. Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm ​

Author: Yunhao Liang, Pujun Zhang, Yuan Qu, Jingyuan Yang, Shaochong Lin, Zuo-jun Max Shen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.24256v2 Announce Type: replace-cross Abstract: The pretrain-transfer paradigm, which underpins the success of large language models (LLMs), has demonstrated the immense power of creating foundation models that learn generalizable representations from vast datasets. However, extending this...

📖 Read original article


398. Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge ​

Author: Pengxiao Lin, Zheng-An Chen, Zhi-Qin John Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.24653v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse of two-hop reasoning. In this work, we argue that this phenomenon can be attributed to a missing super...

📖 Read original article


399. Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning ​

Author: Noah Barnes, Ji Woong Kim, Lingyun Di, Hannah Qu, Anuruddha Bhattacharjee, Miroslaw Janowski, Dheeraj Gandhi, Bailey Felix, Shaopeng Jiang, Olivia Young, Mark Fuge, Ryan D. Sochol, Jeremy D. Brown, Axel Krieger
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2510.09497v2 Announce Type: replace-cross Abstract: In endovascular surgery, endovascular interventionists push a thin tube called a catheter, guided by a thin wire to a treatment site inside the patient's blood vessels to treat various conditions such as blood clots, aneurysms, and malformati...

📖 Read original article


400. People use fast and flat simulation to reason about new games ​

Author: Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum
Published: 7/14/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.GT

arXiv:2510.11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play. But real life also pushes human intelligence along a different fro...

📖 Read original article


Author: Wangjiaxuan Xin, Shuhua Yin, Shi Chen, Yaorong Ge
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.18908v2 Announce Type: replace-cross Abstract: Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic. However, the brevity, informality, and noise of social media short texts often hinder the...

📖 Read original article


402. Enabling Agents to Communicate Entirely in Latent Space ​

Author: Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2511.09149v5 Announce Type: replace-cross Abstract: While natural language is the de facto communication medium for LLM-based agents, it presents a fundamental constraint. The process of downsampling rich, internal latent states into discrete tokens inherently limits the depth and nuance of in...

📖 Read original article


403. Enhancing Adversarial Transferability through Block Stretch and Shrink ​

Author: Quan Liu, Feng Ye, Chenhao Lu, Shuming Zhen, Guanliang Huang, Lunzhe Chen, Xudong Ke
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.17688v2 Announce Type: replace-cross Abstract: Input transformation-based attacks improve adversarial transferability by aggregating gradients over transformed inputs. Existing analyses mainly explain their efficacy from image diversity, semantic preservation, attention variance or hypoth...

📖 Read original article


404. SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation ​

Author: Mohammed Talha Alam, Nada Saadi, Fahad Shamshad, Nils Lukas, Karthik Nandakumar, Fahkri Karray, Samuele Poppi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV, cs.LG

arXiv:2511.19558v2 Announce Type: replace-cross Abstract: Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied af...

📖 Read original article


405. On the Condition Number Dependency in Bilevel Optimization ​

Author: Lesi Chen, Jingzhao Zhang
Published: 7/14/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG

arXiv:2511.22331v3 Announce Type: replace-cross Abstract: Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-level problem. We study the oracle complexity of finding an $\epsilon$-stationary point with first-order ...

📖 Read original article


406. CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning ​

Author: Songqiao Su, Xiaoya Li, Albert Wang, Guoyin Wang, Jiwei Li, Chris Shum
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2512.02551v3 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels. Using CUDA execution speed as the RL...

📖 Read original article


407. The Theory of Strategic Evolution: Games with Endogenous Players and Strategic Replicators ​

Author: Kevin Vallier
Published: 7/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, econ.TH

arXiv:2512.07901v3 Announce Type: replace-cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides the synthesis. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize u...

📖 Read original article


408. Graph-Based Bayesian Optimization for Quantum Circuit Architecture Search with Uncertainty Calibrated Surrogates ​

Author: Prashant Kumar Choudhary, Nouhaila Innan, Muhammad Shafique, Rajeev Singh
Published: 7/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG, cs.NE, cs.NI

arXiv:2512.09586v2 Announce Type: replace-cross Abstract: Quantum circuit design is a key bottleneck for practical quantum machine learning on complex, real-world data. We present an automated framework that discovers and refines variational quantum circuits (VQCs) using graph-based Bayesian optimiz...

📖 Read original article


409. Emotion Recognition in Signers ​

Author: Kotaro Funakoshi, Yaoxiong Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2512.15376v2 Announce Type: replace-cross Abstract: Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatical and affective facial expressions and the scarcity of data for model training. This paper addresses th...

📖 Read original article


410. MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture ​

Author: Hui Li, Fu-Yun Wang, Haoyuan Xia, Jiayue Lyu, Kaihui Cheng, Siyu Zhu, Jingdong Wang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2512.19311v2 Announce Type: replace-cross Abstract: This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data th...

📖 Read original article


411. SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution ​

Author: Habiba Kausar, Saeed Anwar, Omar Jamal Hammad, Ibrahim Radwan, Abdul Bais
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.01406v3 Announce Type: replace-cross Abstract: Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due to the loss of fine structural details and identity-specific features. This work introduces SwinIFS, a ...

📖 Read original article


412. BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts ​

Author: William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault, Vitor D. de Moura, Bertan Ucar, Jose O. Gomes
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.06861v2 Announce Type: replace-cross Abstract: Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions. Existing bias-eva...

📖 Read original article


413. Stable On-Policy Distillation through Adaptive Target Reformulation ​

Author: Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggyu Lim, Taesup Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.07155v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and inf...

📖 Read original article


414. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models ​

Author: Xin Cheng, Rui Tian, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Chengqi Deng, Shangyan Zhou, Chenggang Zhao, Zhewen Hao, Yukun Li, Han Zhang, Zhengyan Zhang, Yixu Wei, M. Y Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.07372v2 Announce Type: replace-cross Abstract: While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forcing them to inefficiently simulate retrieval through computation. To address this, we introduce conditi...

📖 Read original article


415. PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour ​

Author: Liang Wang, Kanzhong Yao, Yang Liu, Weikai Qin, Jun Wu, Zhe Sun, Qiuguo Zhu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2601.15995v2 Announce Type: replace-cross Abstract: Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion. While human athletes can effectively perceive environmental characteristics to select appropriate footholds for obstacle traversal, endowing legged robot...

📖 Read original article


416. Referential Regimes: Transformation-Invariant Identity for Neutral Substrates ​

Author: Denise M. Case
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2601.16152v2 Announce Type: replace-cross Abstract: Data systems increasingly operate under persistent legal, political, and analytic disagreement, where no single interpretive authority can be assumed. A neutral substrate provides stable shared reference without requiring agreement about caus...

📖 Read original article


417. SFO: Learning PDE Operators via Spectral Filtering ​

Author: Noam Koren, Rafael Moschopoulos, Kira Radinsky, Elad Hazan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.17090v2 Announce Type: replace-cross Abstract: Partial differential equations (PDEs) govern complex systems, yet neural operators often struggle to efficiently capture the long-range, nonlocal interactions inherent in their solution maps. We introduce Spectral Filtering Operator (SFO), a ...

📖 Read original article


418. Rethinking Zero-Shot Time Series Classification: From Task-specific Classifiers to In-Context Inference ​

Author: Juntao Fang, Shifeng Xie, Shengbin Nie, Yuhui Ling, Yuming Liu, Zijian Li, Keli Zhang, Lujia Pan, Themis Palpanas, Ruichu Cai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.00620v2 Announce Type: replace-cross Abstract: The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-specific classifier. However, this practice violates the training-free premise of zero-shot deployment and...

📖 Read original article


419. Disentangling Intrinsic Importance from Emergent Structure in Multi-Expert Orchestration ​

Author: Sudipto Ghosh, Sujoy Nath, Sunny Manchanda, Tanmoy Chakraborty
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2602.04291v3 Announce Type: replace-cross Abstract: Multi-expert systems, where multiple Large Language Models (LLMs) collaborate to solve complex tasks, are increasingly adopted for high-performance reasoning and generation. However, the orchestration policies governing expert interaction and...

📖 Read original article


420. Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM) ​

Author: Stephan Vonschallen, Friederike Eyssel, Theresa Schmiedel
Published: 7/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2602.11483v2 Announce Type: replace-cross Abstract: Generative social agents (GSAs) use artificial intelligence to autonomously communicate with human users in a natural and adaptive manner. Currently, there is a lack of theorizing regarding interactions with GSAs, and likewise, few guidelines...

📖 Read original article


421. SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data ​

Author: David Chanin, Adri`a Garriga-Alonso
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.14687v2 Announce Type: replace-cross Abstract: Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. Current LLM-based SAE benchmarks are too noisy to differentiate architectural improvements, while commonly used synthetic-data exp...

📖 Read original article


422. Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision ​

Author: Pengcheng Pan, Yonekura Shogo, Yasuo Kuniyosh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context. Task-driven hard-attention models for vision are often evaluated by how well their scanpaths match human gaze. However, common scanpat...

📖 Read original article


423. BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking ​

Author: Bowen Yu, Sheng Zhang, Binhao Wang, Yi Wen, Jingtong Gao, Bowen Liu, Zimo Zhao, Shanshan Ye, Wanyu Wang, Maolin Wang, Xiangyu Zhao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17686v5 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabilities into smaller models remains challenging due to the capacity mismatch between verbose teachers and co...

📖 Read original article


424. Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers ​

Author: Mohan Tang, Sidi Lu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17993v2 Announce Type: replace-cross Abstract: Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next. In this work, we adopt the perspective that the reasoning power of Transformers is fundamen...

📖 Read original article


425. A General Equilibrium Theory of Orchestrated AI Agent Systems ​

Author: Jean-Philippe Garnier (Br.AI.K)
Published: 7/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, math.OC

arXiv:2602.21255v2 Announce Type: replace-cross Abstract: We establish a general equilibrium theory for systems of large language model (LLM) agents operating under centralized orchestration. The framework is a production economy in the sense of Arrow-Debreu (1954), extended to infinite-dimensional ...

📖 Read original article


426. MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models ​

Author: Kejing Xia, Mingzhe Li, Lixuan Wei, Zhenbang Du, Xiangchi Yuan, Dachuan Shi, Qirui Jin, Wenke Lee
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2603.01331v3 Announce Type: replace-cross Abstract: Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate continuous representatio...

📖 Read original article


427. RVN-Bench: A Benchmark for Reactive Visual Navigation ​

Author: Jaewon Lee, Jaeseok Heo, Gunmin Lee, Howoong Jun, Jeongwoo Oh, Songhwai Oh
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2603.03953v2 Announce Type: replace-cross Abstract: Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglect collisions or are designed for outdoor scenarios, making them unsuitable for indoor visual navigation...

📖 Read original article


428. VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery ​

Author: Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2603.04277v2 Announce Type: replace-cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene. As L...

📖 Read original article


429. Context-Dependent Affordance Computation in Vision-Language Models ​

Author: Murad Farzulla
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2603.04419v2 Announce Type: replace-cross Abstract: We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-VL-30B-A3B ($n = 3{,}213$ scene-context pairs from COCO-2017: 479 images under 7 agentic personas), with...

📖 Read original article


430. Local Message-Passing for Discrete Graph Generation ​

Author: Jay Revolinsky, Harry Shomer, Jiliang Tang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.08825v2 Announce Type: replace-cross Abstract: Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on Graph Transformers or higher order architectures. We revisit this design assumption by introducing GenG...

📖 Read original article


431. MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models ​

Author: Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, eess.AS

arXiv:2603.09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments r...

📖 Read original article


432. Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning ​

Author: Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov, Artem Iudin, Oleg Kiriukhin, Mikhail Pautov, Dmitrii Korzh, Oleg Y. Rogov
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2603.10725v3 Announce Type: replace-cross Abstract: The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to ...

📖 Read original article


433. ECoLAD: Selecting Anomaly Detectors for Automotive Deployment via Compute-Reduction Evaluation ​

Author: Kadir-Kaan "Ozer, Ren'e Ebeling, Markus Enzweiler
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.10926v2 Announce Type: replace-cross Abstract: Automotive anomaly detectors are often selected from accuracy only benchmarks on workstation class hardware, whereas in-vehicle monitoring requires predictable scoring latency under limited CPU parallelism. This mismatch can make methods that...

📖 Read original article


434. Evolutionarily Stable Stackelberg Equilibrium ​

Author: Sam Ganzfried
Published: 7/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA, econ.TH, q-bio.PE

arXiv:2603.18385v3 Announce Type: replace-cross Abstract: We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game setting in which there is a single leading player and a symmetric population of followers. The leader se...

📖 Read original article


435. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models ​

Author: Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-o...

📖 Read original article


436. Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection ​

Author: Ivan Pasichnyk
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.28921v3 Announce Type: replace-cross Abstract: The critical damping condition of the damped harmonic oscillator model of SGD with momentum (Qian, 1999) yields a momentum schedule with no tuned hyperparameters: mu(t) = 1 - 2*sqrt(alpha(t)). Across five seeds on ResNet-18/CIFAR-10 (200-epoc...

📖 Read original article


437. StanceMoE: Mixture-of-Experts Architecture for Stance Detection ​

Author: Abdullah Al Shafi, Md. Milon Islam, Sk. Imran Hossain, K. M. Azharul Hasan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.00878v2 Announce Type: replace-cross Abstract: Actor-level stance detection aims to determine an author expressed position toward specific geopolitical actors mentioned or implicated in a text. Although transformer-based models have achieved relatively good performance in stance classific...

📖 Read original article


438. How Annotation Trains Annotators: Competence Development in Social Influence Recognition ​

Author: Maciej Markiewicz, Beata Bajcar, Wiktoria Mieleszczenko-Kowszewicz, Aleksander Szcz\k{e}sny, Tomasz Adamczyk, Grzegorz Chodak, Karolina Ostrowska, Aleksandra Sawczuk, Jolanta Babiak, Jagoda Szklarczyk, Przemys{\l}aw Kazienko
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.02951v2 Announce Type: replace-cross Abstract: Human data annotation, especially when involving experts, is often treated as an objective reference. However, many annotation tasks are inherently subjective, and annotators' judgments may evolve over time. This study investigates changes in...

📖 Read original article


439. From Paper to Program: Knowledge Externalization and Bottleneck Diagnosis in AI-Assisted Quantum Many-Body Programming ​

Author: Yi Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: physics.comp-ph, cond-mat.str-el, cs.AI, cs.HC

arXiv:2604.04089v5 Announce Type: replace-cross Abstract: Large language models can write scientific code, but direct paper-to-program translation remains fragile when correctness depends on tacit conventions rather than explicit equations. We frame this as a knowledge-externalization problem: index...

📖 Read original article


440. Pickalo: Leveraging 6D Pose Estimation for Low-Cost Industrial Bin Picking ​

Author: Alessandro Tarsi, Matteo Mastrogiuseppe, Saverio Taliani, Simone Cortinovis, Ugo Pattacini
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2604.04690v2 Announce Type: replace-cross Abstract: Bin picking in real industrial environments remains challenging due to severe clutter, occlusions, and the high cost of traditional 3D sensing setups. We present Pickalo, a modular 6D pose-based bin-picking pipeline built entirely on low-cost...

📖 Read original article


441. MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation ​

Author: Sijun Dai, Qiang Huang, Xiaoxing You, Jun Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2604.04969v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle with complex cross-modal reasoning. Flat vector retrieval often ignores structural dependencies, while cu...

📖 Read original article


442. Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation ​

Author: Shutong Zhang, Dylan Zhou, Yinxiao Liu, Yang Yang, Huiwen Luo, Wenfei Zou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.06205v2 Announce Type: replace-cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. While large language models (LLMs) are effective, their high computational cost and latency pr...

📖 Read original article


443. Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation ​

Author: Qian Ma, Sarah Rajtmajer
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2604.07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose ...

📖 Read original article


444. Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces ​

Author: Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.11996v3 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a fundamental limitati...

📖 Read original article


445. SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation ​

Author: Tianhao Fu, Austin Wang, Charles Chen, Roby Aldave-Garza, Yucheng Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2604.15271v3 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decision support. Many strong uncertainty methods require repeated inference, while efficient sin...

📖 Read original article


446. Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning ​

Author: Nicy Scaria, Silvester John Joseph Kennedy, Deepak Subramani
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.HC

arXiv:2604.22770v2 Announce Type: replace-cross Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using gram...

📖 Read original article


447. PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging ​

Author: Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities. Model merging provides a cost-effective mechanism for ...

📖 Read original article


448. Graph Construction and Matching for Imperative Programs using Neural and Structural Methods ​

Author: Arshad Beg, Diarmuid O'Donoghue, Rosemary Monahan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2604.26578v3 Announce Type: replace-cross Abstract: Reusing verification artefacts requires identifying structural and semantic similarities across programs and their specifications. In this paper, we focus on graph construction as a foundational step toward this goal. We present a pipeline th...

📖 Read original article


449. Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration ​

Author: Nihanth W. Cherukuru, Matt Rehme, Kirsten J. Mayer, David John Gagne, John Schreck, John Clyne, Charlie Becker
Published: 7/14/2026, 4:00:00 AM
Categories: physics.data-an, cs.AI, cs.CV, cs.IR

arXiv:2605.00972v2 Announce Type: replace-cross Abstract: Earth system science is producing increasingly large, high-dimensional datasets from both physics-based and AI-driven models. While embedding-based representations make these data searchable and serve as foundational building blocks for AI-dr...

📖 Read original article


450. IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation ​

Author: Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen, Laurence Tianruo Yang, Yurun Jin, Haishan Liu, Changti Wu, Hang Yuan, Cong Huang, Kai Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.CV

arXiv:2605.14712v2 Announce Type: replace-cross Abstract: Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-co...

📖 Read original article


451. Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA ​

Author: Sterling Huang, Abigayle Brown, Jiyoo Noh, Jiakang Xu, Wantong Huo, Kaung Myat Kyaw, Jonathan Chan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.17932v2 Announce Type: replace-cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive architectures. This study examines whether LLMLingua-2 transfers effectively to diffusion large langua...

📖 Read original article


452. Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build ​

Author: Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn, Eyad Kurd-Misto
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2605.21629v3 Announce Type: replace-cross Abstract: How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning outcomes? Self-report surveys show little change, while small-scale behavioral studies report widespread...

📖 Read original article


453. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization ​

Author: Ahmed Alansary, Ali Hamdi
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.05494v3 Announce Type: replace-cross Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model Summarization Framework designed to improve the robustness and quality of abstractive tex...

📖 Read original article


454. Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models ​

Author: Yitong Chen, Shiduo Zhang, Jingjing Gong, Xipeng Qiu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO

arXiv:2606.05737v2 Announce Type: replace-cross Abstract: Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus aligns with image-to-text, not text-to-image. We formalize this...

📖 Read original article


455. The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders ​

Author: Yousef Radwan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.07882v2 Announce Type: replace-cross Abstract: Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different internal representations. We report that they do not. After training, the top sixteen principal d...

📖 Read original article


456. Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models ​

Author: Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2606.11324v2 Announce Type: replace-cross Abstract: We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward gen...

📖 Read original article


457. ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories ​

Author: Siyuan Luo, Nairong Zheng, Lin Zhou, Tiankuo Yao, Shengyou Yuan, Haojia Yu, Cong Pang, Jiapeng Luo, Lewei Lu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.11520v4 Announce Type: replace-cross Abstract: Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution--properties absent from existing datasets. We propose ISE (Intent -> Simulate -> Execute), ...

📖 Read original article


458. Pipette: An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics ​

Author: Zhe Liu, Huanbo Jin, Zhaohui Du, Zhe Wang, Dongzhan Zhou, Minting Pan, He Xu, Peijia Li, Jiaming Gu, Quan Lu, Qi Wang, Bin Ji, Ting Xiao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2606.12936v2 Announce Type: replace-cross Abstract: Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires customizable simulators for safe and reproducible task generation, open editable laboratory assets, and effi...

📖 Read original article


459. Gefen: Optimized Stochastic Optimizer ​

Author: Nadav Benedek, Tomer Koren, Ohad Fried
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV

arXiv:2606.13894v2 Announce Type: replace-cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining. We propose Gefen, a ...

📖 Read original article


460. RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos ​

Author: Zhenhua Wu, Yun Pang, Mingkun Chang, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.16278v2 Announce Type: replace-cross Abstract: Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a promising alternative by reconstructing re...

📖 Read original article


461. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation ​

Author: Renzhi Wu, Zikun Cui, Junjie Yang, Tai Guo, Hong Li, Xian Chen, Li Yu, Ke Pan, Sri Reddy, Mahesh Srinivasan, Nipun Mathur, Haomin Yu, Hong Yan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2606.18379v3 Announce Type: replace-cross Abstract: Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation learning, and real-time serving -- yet existing work addresses each in isolation. We present RankGraph-2...

📖 Read original article


462. FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving ​

Author: Bonan Wang, Letian Tao, Bin Shuai, Jiaxin Gao, Wenxin Zhao, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.21587v2 Announce Type: replace-cross Abstract: Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the prematu...

📖 Read original article


463. Small edits, large models: How Wikipedia advocacy shapes LLM values ​

Author: Jasmine Brazilek, Maria Navas, Alexa Gnauck
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2606.24890v3 Announce Type: replace-cross Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawl...

📖 Read original article


Author: Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Min Zhang, Shaoping Ma, Qingyao Ai
Published: 7/14/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2606.24894v3 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit summarization-oriented metrics, using lexical or semantic sim...

📖 Read original article


465. What Does It Mean to Break a Distillation Defense? ​

Author: Lena Libon, Pura Peetathawatchai, Michael Aerni, Daniel Paleka, Florian Tram`er
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2606.25059v2 Announce Type: replace-cross Abstract: Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work proposes output perturbation defenses that modify the teacher's...

📖 Read original article


466. Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning ​

Author: Yinuo Wang, Gavin Tao, Yuze Liu, John V. Ringwood
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.ET, cs.SY, eess.SP, eess.SY

arXiv:2606.25680v2 Announce Type: replace-cross Abstract: Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task while drawing less thruster power directly extends mission range and endurance. Reinforcement learning yi...

📖 Read original article


467. Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training ​

Author: Jasmine Brazilek, Juliana Seawell
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2606.26102v3 Announce Type: replace-cross Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether th...

📖 Read original article


468. Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare ​

Author: Jasmine Brazilek, Harper Dunn
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.26104v3 Announce Type: replace-cross Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare be...

📖 Read original article


469. NaviCache: Test-Time Self-Calibration Caching for Video Generation ​

Author: Zheqi Lv, Zhibo Zhu, Jinke Wang, Qi Tian, Shengyu Zhang, Zhengyu Chen, Chengxi Zang, Zhou Zhao, Fei Wu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2606.26795v2 Announce Type: replace-cross Abstract: Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive calibration duration, and susceptibility to distribution shifts, o...

📖 Read original article


470. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction ​

Author: Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG

arXiv:2606.28186v2 Announce Type: replace-cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations,...

📖 Read original article


471. CMSL: Constructive Multi-Sequence Learning for Recommendation Systems ​

Author: Zikun Cui, Renzhi Wu, Junjie Yang, Li Sheng, Jijie Wei, Linfeng Liu, Tai Guo, Tao Jia, Xiaodong Wang, Hong Li, Li Yu, Sri Reddy, Hong Yan
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2606.28533v2 Announce Type: replace-cross Abstract: Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Models (DLRM) by capturing the temporal nuances of user behavior. However, current state-of-the-art archite...

📖 Read original article


472. Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation ​

Author: Ananto Nayan Bala, Faisal Muhammad Shah
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR, cs.MA

arXiv:2606.28925v2 Announce Type: replace-cross Abstract: Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agents, while over-selection increases execution cost. The benchmark introduced here is derived from WildCh...

📖 Read original article


473. Freeform Preference Learning for Robotic Manipulation ​

Author: Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2606.32027v2 Announce Type: replace-cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of ...

📖 Read original article


474. FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology ​

Author: Benquan Wang, Jingdao Chen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.00025v2 Announce Type: replace-cross Abstract: While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or sensory deprivation. In contrast, biological systems exhibit remarkable tolerance to these challenges. W...

📖 Read original article


475. Diffusion-GR2: Diffusion Generative Reasoning Re-ranker ​

Author: Zhuoxuan Zhang (Yang), Kangqi Ni (Yang), Yuhang Chen (Yang), Mingfu Liang (Yang), Xiaohan Wei (Yang), Yunchen Pu (Yang), Fei Tian (Yang), Chonglin Sun (Yang), Frank Shyu (Yang), Adam (Yang), Song, Sandeep Pandey, Luke Simon, Tianlong Chen, Xi Liu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.01170v4 Announce Type: replace-cross Abstract: Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per r...

📖 Read original article


476. DemoPSD: Disagreement-Modulated Policy Self-Distillation ​

Author: Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.02502v3 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access. However, rece...

📖 Read original article


477. Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval ​

Author: Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang
Published: 7/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.CV

arXiv:2607.04605v2 Announce Type: replace-cross Abstract: Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, ye...

📖 Read original article


478. x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability ​

Author: Xin Peng, Ang Gao
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.06114v2 Announce Type: replace-cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs). This remains a practical challenge for released checkpoints, since many accelerators r...

📖 Read original article


479. AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning ​

Author: Kyuan Oh, Bumsoo Kim
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.07033v2 Announce Type: replace-cross Abstract: Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and toke...

📖 Read original article


480. LieBN: Batch Normalization over Lie Groups ​

Author: Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, Nicu Sebe
Published: 7/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08783v2 Announce Type: replace-cross Abstract: Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, collect...

📖 Read original article


481. EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins ​

Author: Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
Published: 7/14/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, cs.SY, eess.SY, math.OC

arXiv:2607.08793v2 Announce Type: replace-cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during...

📖 Read original article


482. TACTIC: Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation ​

Author: Rishabh Madan, Angchen Xie, Samantha Saak, Andres Blanco, Dohyeok Lee, Sarah Grace Brown, Yunting Yan, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee
Published: 7/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09218v2 Announce Type: replace-cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break. This setting breaks common implicit assumptions in many lea...

📖 Read original article


483. A Sovereign, Open-Source Foundation Model for German and English ​

Author: The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben H"arle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom R"ohr, Sebastian von Rohrscheidt, J"org Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim K"ohler, Alexander L"oser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max L"ubbering
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.09424v2 Announce Type: replace-cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near...

📖 Read original article


484. Conceptual Networks for Cross-Linguistic Idiomatic Expressions: A Feature-Based Graph Approach ​

Author: Kiran Pala, Punam Silu, Luxin Yu
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.ET

arXiv:2607.09576v2 Announce Type: replace-cross Abstract: We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic. Each express...

📖 Read original article


485. 4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception ​

Author: Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
Published: 7/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09629v2 Announce Type: replace-cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera ...

📖 Read original article