Skip to content

arXiv cs.AI - 2026-07-15 ​

256 items collected.


1. Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets ​

Author: Minmin Zeng, Yi Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11888v1 Announce Type: new Abstract: We develop a rigorous theoretical framework for optimal market making in perpetual futures markets with zero maker fees. We model the market maker's problem as a stochastic optimal control problem on a filtered probability space, where the controls are...

📖 Read original article


2. In-Context Reinforcement Learning under Non-Stationarity: A Survey ​

Author: A Run, Ziluo Ding
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-tuned decision model...

📖 Read original article


3. Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study ​

Author: Thanh Luong Tuan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2607.11948v1 Announce Type: new Abstract: Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter. This paper combines two related FAOS studies into one mechanism-and-control article. First, it repo...

📖 Read original article


4. GRID: Grammar-Railed Decoding for Enterprise SQL Generation ​

Author: Mohsen Arjmandi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.FL, cs.PL

arXiv:2607.11951v1 Announce Type: new Abstract: Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as g...

📖 Read original article


5. Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses ​

Author: Yuhui Bie, Guowei Xu, Yaojun Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11959v1 Announce Type: new Abstract: Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smart-greenhouse control, however, a single simulator return is not enough: a grower or control engineer...

📖 Read original article


6. Optimization Is Not All You Need ​

Author: Minh Hua, Rita Raley
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.11977v2 Announce Type: new Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead...

📖 Read original article


7. LP Mining with LP2Graph: A Use Case for Railway Rescheduling ​

Author: J"orn Maurischat, Nikola Be\v{s}inovi'c, Michael F"arber
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11980v1 Announce Type: new Abstract: Like many optimization-driven domains, railway rescheduling relies on Mixed-Integer Linear Programming (MILP), yet the field's modeling knowledge is scattered across hundreds of papers in incompatible notations, and narrative surveys organize it subjec...

📖 Read original article


8. Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability ​

Author: Said Elnaffar, Farzad Rashidi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.12056v1 Announce Type: new Abstract: Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and...

📖 Read original article


9. Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations ​

Author: Samer Saab Jr, Chaouki Abdallah
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.12077v1 Announce Type: new Abstract: Multi-agent language-model systems increasingly route local interactions, yet the runtime interaction graph is often treated as an implementation detail. We study convention formation in open-weight LM populations spanning 1.1B-32B parameters with a na...

📖 Read original article


10. Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking ​

Author: Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to h...

📖 Read original article


11. Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning ​

Author: Emily Halina, Matthew Guzdial
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12097v1 Announce Type: new Abstract: Video games are a dynamic medium experienced over time. While there are many Procedural Content Generation (PCG) approaches for generating video game levels, they often use representations that abstract away this dynamic nature. In this paper, we intro...

📖 Read original article


12. Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems ​

Author: Ke Sun, Xinyuan Zhang, Xinwu Qian
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12127v1 Announce Type: new Abstract: Learning-based methods for the traveling salesman problem (TSP) are often evaluated through the tours produced after decoding or search, but the learned object itself frequently lives in a surrogate space such as heatmaps, assignments, construction pol...

📖 Read original article


13. The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning ​

Author: Shelley Cazares
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.12177v1 Announce Type: new Abstract: The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML) models pre-tr...

📖 Read original article


14. Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems ​

Author: Navnit Shukla
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.IR

arXiv:2607.12188v1 Announce Type: new Abstract: Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed share...

📖 Read original article


15. A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models ​

Author: Rahul Gupta, Abhinav Mohanty, Payal Motwani, Venkatesh Saligrama, Satyapriya Krishna, Connor Harris, Gary Anthony Ackerman, Brandon Behlendorf, Tom Hobson, Theodore Wilson, Spyros Matsoukas
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY

arXiv:2607.12200v1 Announce Type: new Abstract: As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misu...

📖 Read original article


16. Good Benchmarks ​

Author: Ivan Bercovich
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12217v1 Announce Type: new Abstract: Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons. The best tasks describe a real problem an experienced practitioner would recognize, in language a practitioner would use, with tests that verify the outcome...

📖 Read original article


17. Rethinking the Evaluation of Harness Evolution for Agents ​

Author: Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12227v1 Announce Type: new Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises tw...

📖 Read original article


18. On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage ​

Author: Vinay Kumar Chaganti
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG

arXiv:2607.12257v1 Announce Type: new Abstract: On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful, and at what cost, is unmeasured for a deployable small model. This study fixes one 4B generator on a 24 GB lapt...

📖 Read original article


19. How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks ​

Author: Wei-Jung Huang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study ...

📖 Read original article


20. PM-Bench: Evaluating Prospective Memory in LLM Agents ​

Author: Genglin Liu, Saadia Gabriel
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12385v1 Announce Type: new Abstract: A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabi...

📖 Read original article


21. Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents ​

Author: Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12397v1 Announce Type: new Abstract: LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Re...

📖 Read original article


22. Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions ​

Author: Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, Hongyu Luo, Wun Yu Chan, Wei Fan, Haoran Li, Yangqiu Song
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior ...

📖 Read original article


23. Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting ​

Author: Abdurrahman Javat, Allan Kazakov
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12422v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them. A common design goal is therefore to improve draft quality while reducing auxiliary p...

📖 Read original article


24. EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading ​

Author: Jie Mao, Changlun Li, Xiang Li, Qiqi Duan, Jinhui Yuan, Xiang Liu, Yuyu Luo, Jing Tang, Xiaowen Chu, Nan Tang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CE

arXiv:2607.12455v1 Announce Type: new Abstract: Quantitative strategy optimization remains largely manual, requiring domain experts to identify weak signals, tune risk-control rules, and repeatedly validate iterative revisions. Large language models can accelerate this process, but directly relying ...

📖 Read original article


25. Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting? ​

Author: Qihang Zhang, Siyao Zhang, Letao Kang, Wenzhe Liang, Miao Zhang, Zhao Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12462v1 Announce Type: new Abstract: Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each individual node and all nodes across the ...

📖 Read original article


26. Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models ​

Author: Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.12463v1 Announce Type: new Abstract: Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent i...

📖 Read original article


27. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery ​

Author: Ingmar Posner, Anson Lei, Bernhard Sch"olkopf
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12474v2 Announce Type: new Abstract: Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from protein folding to weather forecasting. Yet prediction alone does not constitute scientific discovery....

📖 Read original article


28. TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments ​

Author: Edward Y. Chang, Emily J. Chang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for writing records against it, and one operating discipline, no durable state change without a record. ...

📖 Read original article


29. The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank ​

Author: Bojie Li, Noah Shi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12520v1 Announce Type: new Abstract: What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first description a human sees, making that parametric corpus presence a measurement problem. Citations explain about a third of ...

📖 Read original article


30. Evidence-Grounded AI for Musculoskeletal Care ​

Author: Wenjie Li, Yujie Zhang, Fanrui Zhang, Haoran Sun, Renhao Yang, Junjun He, Weiran Huang, Yuanfeng Ji, Chenrun Wang, Kailing Wang, Hongcheng Gao, Kaipeng Zhang, Hanyu Wang, Angela Lin Wang, Xingqi He, Yilin Huang, Shiyi Yao, Lilong Wang, Yankai Jiang, Yirong Chen, Chenglong Ma, Jiyao Liu, Ming Hu, Gen Li, Yidong Xu, Chengyu Zhuang, Jiawei Liu, Yin Zhang, Lequan Yu, Lu Chen, Yinpeng Dong, Lei Liu, Carlos Gutierrez Sanroman, Yu Qiao, Weijie Ma, Xiaosong Wang, Lei Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12527v2 Announce Type: new Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery, remodelling and degeneration often unfold over months to years, musculoskeletal care requires longit...

📖 Read original article


31. Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring ​

Author: Anna Gatzioura, Vrettos Moulos, Nina Baranowska
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12588v1 Announce Type: new Abstract: According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to specific areas, like risk management, data quality and governance, logging and traceability, technical documentation, ...

📖 Read original article


32. Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing ​

Author: Amin Beheshti, Rong N. Chang, Boualem Benatallah, Fabio Casati, Schahram Dustdar, Geoffrey Fox, Quan Z. Sheng, Yan Wang, Jian Yang, Albert Zomaya
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.ET

arXiv:2607.12619v1 Announce Type: new Abstract: The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move from isolated cognit...

📖 Read original article


33. Atomic Units of X: The Compression Layer of Intelligence ​

Author: Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12634v1 Announce Type: new Abstract: This paper proposes a theoretical framework for understanding intelligence as a process of atomic compression and compositional reuse. We argue that cognitive, biological, computational, and organizational systems achieve scalable intelligence by decom...

📖 Read original article


34. A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism ​

Author: Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and...

📖 Read original article


35. Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration ​

Author: Quanyan Zhu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SY, eess.SY

arXiv:2607.12662v1 Announce Type: new Abstract: The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical AI, edge computing, and digital twins into a unified closed-loop orchestration framework. The propo...

📖 Read original article


36. MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku ​

Author: Pedro Orvalho, Guillem Aleny`a, Felip Many`a
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.12711v1 Announce Type: new Abstract: Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles. However, despite strong perceptual capabilities, these models lack explicit mechanisms for enforcing log...

📖 Read original article


37. LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos ​

Author: Julius Steiglechner, Lucas Mahler, Gabriele Lohmann
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.12733v1 Announce Type: new Abstract: Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after...

📖 Read original article


38. Tracing Agentic Failure from the Flow of Success ​

Author: Samuel Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.12747v1 Announce Type: new Abstract: Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which...

📖 Read original article


39. Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative ​

Author: Koen Oostermeijer
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12767v1 Announce Type: new Abstract: Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilities sum over tokens, longer answers tend to be penalized relative to shorter ones in practice. A common mitigatio...

📖 Read original article


40. Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? ​

Author: Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.MM

arXiv:2607.12787v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. Howeve...

📖 Read original article


41. Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents ​

Author: Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We make three claims. First,...

📖 Read original article


42. Visual Access Boundaries in Vision-Language Model Reasoning ​

Author: Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image...

📖 Read original article


43. Human-AI Agent Interaction as a Neuroplastic Training Environment ​

Author: Eranga Bandara, Ross Gore, Asanga Gunaratna, Ravi Mukkamala, Nihal Siriwardanagea, Gihan Siriwardanagea, Sachini Rajapakse, Isurunima Kularathna, Pramoda Karunarathna, Chalani Rajapakse, Sachin Shetty, Christopher K. Rhea, Ng Wee Keong, Kasun De Zoysa, Amin Hass, Shaifali Kaushik, Wathsala Herath, Preston Samuel, Anita H. Clayton, Atmaram Yarlagadd
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12823v1 Announce Type: new Abstract: Interaction with AI agents has become one of the most frequent activities of everyday digital life. Whether conversing with an assistant, working with a coding copilot, or generating images, the interaction follows a common iterative loop: a request is...

📖 Read original article


44. Solution of the Hempel's statistical ambiguity problem and Causal AI ​

Author: Evgenii Vityaev
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12826v1 Announce Type: new Abstract: This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement o...

📖 Read original article


45. A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study ​

Author: Cameron Cagan, Pedram Fard, Jiazi Tian, Jingya Cheng, Shawn N. Murphy, Hossein Estiri
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12886v1 Announce Type: new Abstract: Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on superv...

📖 Read original article


46. MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations ​

Author: Xixuan Hao, Zeyu Zhang, Zehao Lin, Yihang Sun, Ziliang Guo, Xichong Zhang, Yuxuan Liang, Feiyu Xiong, Zhiyu Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.12893v1 Announce Type: new Abstract: Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, sc...

📖 Read original article


47. Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes ​

Author: Jonas Ehrhardt, Ren'e Heesch, Oliver Niggemann
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12924v2 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine ...

📖 Read original article


48. FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation ​

Author: Ruoran Xu, Wending Gao, Qiufeng Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SC

arXiv:2607.12982v1 Announce Type: new Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram gener...

📖 Read original article


49. Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs ​

Author: Sen Yang, Yuen-Hei Yeung
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12985v1 Announce Type: new Abstract: Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC...

📖 Read original article


50. Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation ​

Author: Aleh Manchuliantsau
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.12986v1 Announce Type: new Abstract: Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while ...

📖 Read original article


51. Dynamic Resource Allocation for Ensemble Determinization MCTS ​

Author: Jakub Kowalski, Adam Ci\k{e}.zkowski, Artur Krzy.zy'nski, Mark H. M. Winands
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13007v1 Announce Type: new Abstract: Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomness and hidden information. In particular, several Monte Carlo Tree Search (MCTS) variants are commo...

📖 Read original article


52. Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model ​

Author: Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.SD

arXiv:2607.13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoisi...

📖 Read original article


53. Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution ​

Author: Junjie Yin, Xinyu Feng
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE, cs.SY, eess.SY

arXiv:2607.13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencie...

📖 Read original article


54. Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain ​

Author: Qiaoni Shi, Kai Zhu, Kai Gu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, econ.GN, q-fin.EC

arXiv:2607.07652v1 Announce Type: cross Abstract: Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream...

📖 Read original article


55. FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis ​

Author: Marlena Fl"uh, Soo-Yon Kim, Carolin Victoria Schneider, Sandra Geisler
Published: 7/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.DB

arXiv:2607.11464v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions. Graph-based RAG approaches, such as GraphRAG, enhance retrieval by capturing semantic relationships ...

📖 Read original article


56. Scaling Point-in-Time Language Models ​

Author: Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-in-time ...

📖 Read original article


57. So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis ​

Author: Abdullah Akinde, Mariam Akinde, Rasheedat Emiola, Ahmed Akinsola
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach...

📖 Read original article


58. CANDI: Contextual Alignment for Niche Domains Question Answering ​

Author: Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often fail to capture the n...

📖 Read original article


59. G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis ​

Author: Xingyu Xiao, Mao Du, Jiejuan Tong, Jingang Liang, Haitao Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11892v1 Announce Type: cross Abstract: Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert interpretation of narrative reports and guideline-based reasoning.Existing data-driven or one-shot larg...

📖 Read original article


60. I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs ​

Author: Abdullah Abdullah
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11893v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-critical modalities such as Braille remains unclear. We evaluate state-of-the-art LLMs on bidirectional Korean-Brail...

📖 Read original article


61. Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels ​

Author: Yuliia Vistak, Viktoriia Makovska, Vera Schmitt, Veronika Solopova
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11894v1 Announce Type: cross Abstract: Detecting disinformation narratives on social media is challenging due to the scale of amplification, rapid evolution, and linguistic variability of online content. We propose a graph-based framework for identifying and analyzing disinformation narra...

📖 Read original article


62. OmniPMNet: Bridging discrete and gridded PM10 forecasts via omni-query neural processes ​

Author: Shuangshuang He, Shuo Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.ao-ph

arXiv:2607.11896v1 Announce Type: cross Abstract: Forecasting particulate matter (PM10) requires both station-scale accuracy and continuous spatial fields, especially during severe dust storms. Chemical transport models (CTMs) provide gridded forecasts but retain local biases, whereas graph neural n...

📖 Read original article


63. SeqGPT: A Constrained Transformer Agent for the Inverse Designof Multi-Panel Composite Structures ​

Author: Driss Chraibi (Toulouse INP), Alejandro Garc'ia Pis (IRIT), St'ephane Grihon (IRIT), Sixin Zhang (IRIT)
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.11910v1 Announce Type: cross Abstract: Optimizing composite stacking sequences to match continuous targets (e.g., Lamination or Buckling Parameters) with discrete manufacturing constraints represents a challenging combinatorial inverse problem that regularly occurs in composite design esp...

📖 Read original article


64. Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming ​

Author: Ali Kohan, Mohamad Roshanzamir, Roohallah Alizadehsani, Seyedali Mirjalili
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.11913v1 Announce Type: cross Abstract: Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-linear reasoning workflows. A prominent example is Genetic Network Programming (GNP), a self-evolving ...

📖 Read original article


65. Burst Spiking Neural Networks ​

Author: Jiahong Zhang, Sijun Shen, Man Yao, Han Xu, Mingqiang Huang, Yonghong Tian, Bo Xu, Guoqi Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.11914v1 Announce Type: cross Abstract: A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Artificial Neural Networks (ANNs). This work further argues that realizing this ambition requires improving not onl...

📖 Read original article


66. QDEvo: A Multi-Objective Quality-Diversity Framework for Automated Heuristic Design ​

Author: Nam Do Khanh, Nhat Nguyen Tran Minh, Dat Pham Vu Tuan, Long Doan, Binh Huynh Thi Thanh
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CL

arXiv:2607.11916v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) with evolutionary computation has emerged as a powerful paradigm for automated heuristic design in combinatorial optimization. However, existing approaches suffer from mode collapse, converging to homog...

📖 Read original article


67. AAAI-26 Dual Submissions: Novel Challenges ​

Author: Kiri L. Wagstaff, Joydeep Biswas, Erich Merrill III, Bo An, Ida Camacho, David J. Crandall, Matthew E. Taylor
Published: 7/15/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.CY, cs.LG

arXiv:2607.11918v1 Announce Type: cross Abstract: Dual submissions, in which identical or substantially similar papers are simultaneously submitted to one or more archival venues, without cross-citation or disclosure, are a growing problem for the AAAI Conference and other scientific publication ven...

📖 Read original article


68. Do You Remember? Toward Memory-Centric Multimodal AI ​

Author: Xuguang Yu, Weigang Zheng, Minyue Yu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV, cs.LG

arXiv:2607.11919v1 Announce Type: cross Abstract: Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-shot text output, and discard internal representations. We present DoYo...

📖 Read original article


69. Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making ​

Author: Jeff Helzner
Published: 7/15/2026, 4:00:00 AM
Categories: econ.EM, cs.AI, stat.ME

arXiv:2607.11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of a...

📖 Read original article


70. Mathematics of Data Science ​

Author: Afonso S. Bandeira, Amit Singer, Thomas Strohmer
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT, math.PR

arXiv:2607.11938v1 Announce Type: cross Abstract: This book is about the mathematical foundations of data science. 1. Introduction 2. Curses, Blessings, and Surprises in High Dimensions 3. Singular Value Decomposition and Principal Component Analysis 4. Linear Regression and Regularization 5. Graphs...

📖 Read original article


71. CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA ​

Author: Gengyu Zhang, Haiyin Ran, Zhengbao He, Yuhang Liu, Hanling Tian, Zhehao Huang, Xiaolin Huang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11940v1 Announce Type: cross Abstract: As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging. Low-Rank Adaptation (LoRA), currently one of the most widely adopted parameter-efficient fine-tuning (PEFT)...

📖 Read original article


72. How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit ​

Author: Daming Luo, Christy Liang, Junyu Xuan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11942v1 Announce Type: cross Abstract: KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol. Yet the economic case for a compressed KV cache is reuse: compress a document once, answer many future quest...

📖 Read original article


73. BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification ​

Author: Raghvender Raghvender, Mahdi Abid, Ferran Brosa Planella, Charles Delacourt, Arnaud Demorti`ere
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11943v1 Announce Type: cross Abstract: Long-horizon physics-based simulations of battery degradation provide mechanistic insight but remain computationally expensive, limiting their use for dense exploration of operating conditions over extended cycle life. Here, we propose a hybrid physi...

📖 Read original article


74. Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite ​

Author: Yushi Hirose, Hiroo Irobe, Takafumi Kanamori
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11947v2 Announce Type: cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted to binary...

📖 Read original article


75. BAT-RM: A Boundary-Aware Transformer with Region-Aware Multi-Directional Mamba for Clinically Deployed Cervical Cancer Radiotherapy Auto-Contouring ​

Author: Istiak Ahmed, Kazi Shahriar Sanjid, Galib Ahmed, Md. Tanzim Hossain, Md. Anwarul Islam, Shahrukh Khan, Md. Ashrif Rahman Arian, Md. Nishan Khan, Md. Misbah Khan, S M Hasibul Hoque, Rahnuma Shahrin Rista, Md. Jobairul Islam, Sheikh Anisul Haque, Md Arifur Rahman, Syed Md. Akram Hussain, Syeda Nashra, Sayeed Shafayet Chowdhury, Md. Mostafa Kamal Sarker, M. Monir Uddin
Published: 7/15/2026, 4:00:00 AM
Categories: eess.IV, cs.AI

arXiv:2607.11949v1 Announce Type: cross Abstract: We present a clinically deployed end-to-end auto-contouring system for cervical cancer radiotherapy planning, anchored by the Boundary-Aware Transformer with Region-Aware Mamba (BAT-RM), a hybrid architecture that integrates Sobel-gated boundary atte...

📖 Read original article


76. Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG ​

Author: Dibakar Sigdel
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11950v1 Announce Type: cross Abstract: Brain field potentials are scale-free: their power spectra follow a $1/f^{\beta}$ law whose aperiodic exponent $\beta$ tracks cortical state, and sleep depth in particular is a shift in $\beta$. We ask whether a transformer endowed with an explicit r...

📖 Read original article


77. Graph-Constrained Policy Learning for Extreme Clinical Code Prediction ​

Author: Amritpal Singh, Sebastian Torres, Khawar Shakeel, Syed Ahmad Chan Bukhari
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2607.11954v1 Announce Type: cross Abstract: Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space. Most systems treat the task as flat multi-label classification, scoring codes independently and providing ...

📖 Read original article


78. Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction ​

Author: Zongye Lyu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS

arXiv:2607.11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal. Exact ...

📖 Read original article


79. Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability ​

Author: Mashrul Hossain, Nafesa Kibria, Fahim Shahriar
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11963v1 Announce Type: cross Abstract: The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially mislea...

📖 Read original article


80. Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse ​

Author: Truong Xuan Khanh, Phan Thanh Duc
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11967v1 Announce Type: cross Abstract: Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate freedom. Neural Collapse sharpens this problem: networks converge toward a shared, low-dimensional geometry, raisi...

📖 Read original article


81. Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems ​

Author: Yubo Zhang, Xiaodong Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT

arXiv:2607.11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed scheme integrates the ICL-Transformer backbone with the pilot ...

📖 Read original article


82. Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance ​

Author: Zixuan Shen (Central South University), Bingchuan Wang (Central South University), Zhi Wang (Nanjing University), Yong Wang (Central South University)
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11974v1 Announce Type: cross Abstract: Most neural partial differential equation (PDE) surrogates learn how fields evolve after a grid has already been chosen. However, before any operator is applied, the grid has already determined how modeling capacity is allocated across space, resolut...

📖 Read original article


83. Signal-Guided Optimization for Machine Unlearning ​

Author: Xujia Li, Dan Li, Jian Lou, Wenjie Feng
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11975v1 Announce Type: cross Abstract: Current machine unlearning methods predominantly rely on global, coarse-grained intervention strategies. They lack precise pilot signals to guide the unlearning process and fail to provide differentiable guidance across different unlearning tasks. Du...

📖 Read original article


84. Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design ​

Author: Hang Yuan, Chen Li, Wenjun Ma, Tadahiko Murata, Yuncheng Jiang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11978v1 Announce Type: cross Abstract: Precision molecular design aims to discover personalized drug candidates through joint control of multiple conditions, such as biological relevance and molecular design strategies. Biological relevance reflects cellular functional states under diseas...

📖 Read original article


85. Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring ​

Author: Yi Gui
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11981v1 Announce Type: cross Abstract: Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generaliza...

📖 Read original article


86. Removable Defects: The Economics and Limits of Deliberate Deficiency ​

Author: Cheng Qian
Published: 7/15/2026, 4:00:00 AM
Categories: econ.EM, cs.AI, cs.LG, stat.ML

arXiv:2607.11983v2 Announce Type: cross Abstract: A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a deficiency can be kept because it pays and removed on demand in the rare situation where it would be...

📖 Read original article


87. Sparse Inter-Layer Dependencies of Transformer FFN Neurons ​

Author: Johannes Knittel, Hanspeter Pfister
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.11990v1 Announce Type: cross Abstract: Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream...

📖 Read original article


88. Mitigating The Effect of Class Imbalance in Data with Hierarchical and Dependable Structure ​

Author: Bipin Chhetri, Deepika Giri, Avishek Kadel, Rabin Kumar Karki, Akbar Siami Namin
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11994v1 Announce Type: cross Abstract: Classifying cybersecurity vulnerabilities using the Common Weakness Enumeration (CWE) taxonomy is challenging due to extreme class imbalance and strong hierarchical dependencies among weakness categories. Although oversampling techniques such as Synt...

📖 Read original article


89. Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs ​

Author: Nikita Kozodoi, Zainab Afolabi, Jack Butler
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.11997v2 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically study...

📖 Read original article


90. HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning ​

Author: Abubakar Hamisu Kamagata, Dharm Singh Jat, Attlee Munyaradzi Gamundani, Saravanakumar Paramasivam, Babangida Sani, Aliyu Zakariyya
Published: 7/15/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.LG

arXiv:2607.11998v1 Announce Type: cross Abstract: High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper presents a video-based and high performance computing (HPC) enabled deep learning framework for joi...

📖 Read original article


91. An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering ​

Author: Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed, Dilnaz Utemissova, Ufaq Khan, Mohammad Yaqub
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2607.12048v1 Announce Type: cross Abstract: Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without forgetting previously acquired knowledge. Continual learning (CL) provides a practical framework for...

📖 Read original article


92. Representation and Reference Selection in Training-Free Synthetic Image Attribution ​

Author: Meiling Li, Pietro Bongini, Benedetta Tondi, Mauro Barni
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR

arXiv:2607.12052v1 Announce Type: cross Abstract: Synthetic image attribution aims at identifying the generator responsible for a given AI-generated image. Training-free reference-based attribution methods are easily scalable, since newly emerging generators can be incorporated by adding source-spec...

📖 Read original article


93. AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration ​

Author: Arastoo Zibaeirad, Marco Vieira, Thomas Zimmermann
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR

arXiv:2607.12058v1 Announce Type: cross Abstract: Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer demands interp...

📖 Read original article


94. Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation ​

Author: Robel Mamo, Rajitha de Silva, Grzegorz Cielniak, Taeyeong Choi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.12065v1 Announce Type: cross Abstract: While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying autonomous robots at night offers significant advantages, including 24-hour crop and soil monitoring, f...

📖 Read original article


95. Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation ​

Author: Md. Sadibul Hasan Sadib, Md. Mohayminul Mukit, Rahmatul Kabir Rasel Sarker, Tahmid Alam Tamim, Md. Monir Hossain Shimul
Published: 7/15/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2607.12075v1 Announce Type: cross Abstract: Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift. Methods: We ...

📖 Read original article


96. Sparse Autoencoders for Interpretable Out-of-Distribution Detection ​

Author: Ayush Karmacharya (Purdue University), Luke Luschwitz (Purdue University), Lucia Romero (Purdue University), Yanan Niu (EPFL), Joseph Campbell (Purdue University)
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks often produce overconfident predictions for inputs that deviate from their training data, leading to significant de...

📖 Read original article


97. PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs ​

Author: Jing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, Yang Liu, Wei Zhou
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.MA, cs.NI

arXiv:2607.12111v1 Announce Type: cross Abstract: Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Lan...

📖 Read original article


98. Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning ​

Author: Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC

arXiv:2607.12112v1 Announce Type: cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic f...

📖 Read original article


99. Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap ​

Author: Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, Fr'ed'eric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie
Published: 7/15/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2607.12113v1 Announce Type: cross Abstract: One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have ...

📖 Read original article


100. GaitSpan: Growing Humanoid Locomotion from Walking to Running ​

Author: Kwan-Yee Lin, Zilin Wang, Janelle J. Liu, Stella X. Yu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.12114v1 Announce Type: cross Abstract: A humanoid that can walk should not relearn locomotion from scratch to jog or run. Yet current approaches often obtain gait diversity by prescribing gait schedules, imitating motion clips, training experts to switch between or distilling skills into ...

📖 Read original article


101. Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models ​

Author: Xu Han, Jiajing Hu, Li-Ping Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.12171v1 Announce Type: cross Abstract: In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous velocity or the data endpoint, to perform denoising. Although prior work shows that these parameterizations lead ...

📖 Read original article


102. From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data ​

Author: Pradyumna Elavarthi, Arun J. Bhattacharjee, Harrison Lisabeth, Anca Ralescu, Petrus H. Zwart, Dilworth Parkinson, Elizabeth G. Clark
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12175v1 Announce Type: cross Abstract: X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated volumetric data acquisition and reconstruction. However, rapid interpretation remains limited by image segmentat...

📖 Read original article


103. TRAIL: A Platform for Configurable Human--AI Teaming Experiments ​

Author: Mohammad Amin Samadi, Pedro Martins De Bastos, Jaeyoon Choi, Spencer JaQuay, Seehee Park, Nia Nixon
Published: 7/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.12180v1 Announce Type: cross Abstract: An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions. Studying this rigorously demands infrastructure no existing tool provides: reproducible configuration of an A...

📖 Read original article


104. Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing ​

Author: Gabriel Paris-Colombo, Rodrigo M. Cabral-Carvalho, Felipe D. Toro-Hern'andez
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.12195v1 Announce Type: cross Abstract: Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By appl...

📖 Read original article


105. The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests ​

Author: Edgar Dobriban
Published: 7/15/2026, 4:00:00 AM
Categories: math.ST, cs.AI, stat.ME, stat.TH

arXiv:2607.12208v1 Announce Type: cross Abstract: We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-sided Gaussian $p$-values. We construct a factor model for which, at level $\alpha=0.01$, a rigorous interval-ar...

📖 Read original article


106. RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls ​

Author: Brenda Lelis, Rodrigo Cabral-Carvalho
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA

arXiv:2607.12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task. This creates a practical allocatio...

📖 Read original article


107. Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor ​

Author: Burhan Ogut, Michelle Yin
Published: 7/15/2026, 4:00:00 AM
Categories: econ.EM, cs.AI

arXiv:2607.12219v1 Announce Type: cross Abstract: We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly nonlinear function of the latent variable. The problem is acute in the measurement of occupational exposure to a...

📖 Read original article


108. Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals ​

Author: Mohotarema Rashid, Lingzi Hong, Junhua Ding, K. S. M. Tozammel Hossain
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.12233v1 Announce Type: cross Abstract: Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with little evidence from live deployment. We present Fin-Analyst, a hybrid agent for FinMMEval 2026 Task 3: an eight-s...

📖 Read original article


109. Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents ​

Author: Ning Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed, what it suspects, a...

📖 Read original article


110. Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs ​

Author: Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable aut...

📖 Read original article


111. The Sound of Absence: Audio-Language Embedding Models Struggle with Negation ​

Author: Chun-Yi Kuan, Hung-yi Lee
Published: 7/15/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD

arXiv:2607.12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping a...

📖 Read original article


112. A Longitudinal Analysis of Public Discourse on AI Ethics in Education Using Twitter Data ​

Author: Akriti Bagale, Nafisa Mehjabin, Ali "Unl"u, Aditya Johri
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.12295v1 Announce Type: cross Abstract: The rapid integration of artificial intelligence (AI) and generative AI (GenAI) into education presents significant opportunities to enhance teaching and learning, while raising ethical concerns about the responsible use of these technologies in educ...

📖 Read original article


113. A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education ​

Author: Amrita Ganguly, Aditya Johri, Nora McDonald, Areej Ali, Umama Dewan, Aayushi Hingle Collier
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.12296v1 Announce Type: cross Abstract: With the increased use of generative AI (GenAI) applications such as ChatGPT, higher education institutions (HEIs) have released a range of guidelines and policies to direct adoption within their institutions. In computer science (CS) courses GenAI a...

📖 Read original article


114. LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes ​

Author: Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti, Jesse C. Cresswell, Victor Zhong
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.12310v1 Announce Type: cross Abstract: While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, weakly s...

📖 Read original article


115. Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study) ​

Author: Farnaz Farid, Raihan Alam, Al Al-Areqi, Farhad Ahamed, Muhammad Hassan Khan, Sadia Hossain, Irena Veljanova, Anika Tabassum Binte Hossain
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.ET, cs.HC

arXiv:2607.12336v1 Announce Type: cross Abstract: Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivalent effect by simultaneously acting as a combatant against and a spread vector for misinformation. A ...

📖 Read original article


116. Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction ​

Author: Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12364v1 Announce Type: cross Abstract: EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted, and low-detail, causing SSIM, LPIPS, and CLIP to penalize semantically recoverable outputs or reward plausible...

📖 Read original article


117. IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment ​

Author: Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV

arXiv:2607.12375v1 Announce Type: cross Abstract: Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, ...

📖 Read original article


118. Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification ​

Author: Zhirui Zhang, Tianhang Nan, Yong Ding, Zhuolun Song, Dayu Hu, Xiaoyu Cui
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12376v1 Announce Type: cross Abstract: Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in digital pathology. These methods adjust feature distributions of subtle evidence sub-images to correctly associa...

📖 Read original article


119. ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning ​

Author: Zijie Wang, Wei Zhang, Weiming Zhang, Xiao Tan, Weikai Chen, Xiaoxu Li, Guanbin Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12433v1 Announce Type: cross Abstract: Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect th...

📖 Read original article


120. The Computational Basis of Confidence in Large Language Models ​

Author: Dharshan Kumaran, Viorica Patraucean, Maks Ovsanikov, Petar Veli\v{c}kovi'c, Nathaniel Daw
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12447v1 Announce Type: cross Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibr...

📖 Read original article


121. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge ​

Author: Shuming Fang, Shuifei Zeng
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM...

📖 Read original article


122. Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric ​

Author: Oleg Solozobov
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores recons...

📖 Read original article


123. OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning ​

Author: Emil Mittag, Richard Dazeley, Peter Vamplew
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12523v1 Announce Type: cross Abstract: Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow environmental shifts. The detection of out-of-distribution conditions is pivotal to determining when an agent's...

📖 Read original article


124. Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel ​

Author: Niccol`o Caselli, Francesco Massafra, Samuele Punzo, Salvatore Lo Sardo, Ippokratis Pantelidis, Sathya Kamesh Bhethanabhotla
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.12547v2 Announce Type: cross Abstract: We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an extension that freezes the pretrained low-level LeWM and adds high-level planning over latent subgoals. We evaluate ...

📖 Read original article


125. Traceback Translators Against Forgetting in Continual Fake Speech Detection ​

Author: Enrico Gottardis, Mattia Tamiazzo, Simone Milani
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM, cs.SD

arXiv:2607.12569v1 Announce Type: cross Abstract: Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new da...

📖 Read original article


126. Deep Learning-based Surrogate Modelling of the LOD Method for Multiscale Problems ​

Author: Marc Haltmayer, Jaemin Seo, Yuseung Lee, Sungyeop Lee, Jaehoon Jeong, Jae Yong Lee
Published: 7/15/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.LG, cs.NA

arXiv:2607.12570v1 Announce Type: cross Abstract: Multiscale problems are notoriously difficult to tackle using traditional numerical methods, as accurately resolving fine-scale features often requires prohibitively fine discretizations. This challenge is particularly pronounced in applications such...

📖 Read original article


127. Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction ​

Author: Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR, cs.MM

arXiv:2607.12584v1 Announce Type: cross Abstract: The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that ...

📖 Read original article


128. Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs ​

Author: Zhili Huang, Ling Xu, Hongyu Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.12605v1 Announce Type: cross Abstract: Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce dif...

📖 Read original article


129. Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making? ​

Author: Minh Khoi Ho, Zihao Zhu, Runchuan Zhu, Levina Li, Zhiwen Fan, Zhangyang Wang, Junyuan Hong
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.12631v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors that may modulate their decision-making becomes critical. While LLMs are trained to perceive and resonate with use...

📖 Read original article


130. Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs ​

Author: Junyu Ren
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY, cs.SE

arXiv:2607.12650v1 Announce Type: cross Abstract: Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Agentic Reason...

📖 Read original article


131. Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference ​

Author: Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, Meng Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.12659v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational complexity, w...

📖 Read original article


132. Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings ​

Author: Yan Gong, Bohao Li, Bowen Du, Junchen Ye
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12678v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding. Among CAD analysis tasks, panoptic symbol spottin...

📖 Read original article


133. From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation ​

Author: Mehak Dhaliwal, Rasta Tadayon, Andong Hua, Haewon Jeong, Yao Qin
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.12687v1 Announce Type: cross Abstract: LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, making it critical to know not only what a model predicts, but when its predictions can be trusted. W...

📖 Read original article


134. Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts ​

Author: Jincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DC

arXiv:2607.12696v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activation patterns. Speculative decoding (SD) accelerates autoregressive gen...

📖 Read original article


135. Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing ​

Author: William Franz Lamberti
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.12713v1 Announce Type: cross Abstract: Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback is a lever on all three. We compare two deliveries of the same requested changes: a holistic prompt (control) ve...

📖 Read original article


136. Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities ​

Author: Qiyuan Fan, Zhi Li, Junjie Li, XiaoFeng Wang, Bin Yuan, Deqing Zou
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE

arXiv:2607.12723v1 Announce Type: cross Abstract: Filesystem isolation in container ecosystems is often weakened by cross-boundary path misresolution, causing path traversal (PaTra) vulnerabilities. These vulnerabilities stem from insecure host-container interactions and have become increasingly per...

📖 Read original article


137. Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty ​

Author: Sarah Al-Shareeda, Gulcihan Ozdemir, Heung Seok Jeon
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12730v1 Announce Type: cross Abstract: Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs. Missing features must then be reconstructed, and their errors can propagate through...

📖 Read original article


138. Weakly Supervised Spatio-Temporal Candidate Discovery of Dairy Farm Sites from Seasonal Satellite Imagery ​

Author: Usman Haider, Fatima Khalid, Karl Mason
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12748v1 Announce Type: cross Abstract: Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pasture, field boundaries, roads, buildings, and seasonal vegetation patterns. Direct farm labels are often incomplet...

📖 Read original article


139. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation ​

Author: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12752v2 Announce Type: cross Abstract: While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated struc...

📖 Read original article


140. Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems ​

Author: Nathan G. Wood, Andrew P. Rebera
Published: 7/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY, cs.ET, cs.RO

arXiv:2607.12755v1 Announce Type: cross Abstract: AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions and capabilities. I.e., many systems take inputs and give outputs, but without users having any ability to see h...

📖 Read original article


141. Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination ​

Author: Usman Haider, Karl Mason
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.12763v1 Announce Type: cross Abstract: Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation methods such as FedAvg do not account for system-level constraints, often leading to unsafe global ...

📖 Read original article


142. HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition ​

Author: Aleksei Bakin, Andrey V. Savchenko
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12774v1 Announce Type: cross Abstract: This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with simultaneous prediction of valence, arousal, facial expressions, and action units on s-Aff-Wild2 dataset, we use f...

📖 Read original article


143. When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis ​

Author: Mehdi Saeedi, Eddie Richter, Paul Hartke
Published: 7/15/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET

arXiv:2607.12780v1 Announce Type: cross Abstract: Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford resources such as T gates. We study this problem using a compact 44.8M-parameter encoder-decoder transformer with ...

📖 Read original article


144. Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels ​

Author: Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to the benchmark's grading procedure and capture only observed behavior on a given set of at...

📖 Read original article


145. The One-Word Census: Answer-Choice Conformity Across 44 Language Models ​

Author: Tapan Parikh
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2607.12796v1 Announce Type: cross Abstract: When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the ti...

📖 Read original article


146. Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVs ​

Author: Wei-Hao Liou, Teng-Hu Cheng
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2607.12801v1 Announce Type: cross Abstract: This study introduces a unified control framework for fixed-wing unmanned aerial vehicles (UAVs) fitted with a pan-tilt (PT) camera, intended to perform an end-to-end mission spanning from initial target detection to accurate terminal engagement. The...

📖 Read original article


147. PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops ​

Author: Sarthak Chittawar, Vansh Garg, Aditya Vadali, Krish Pandya, Rohit Jayanti, Sourav Garg, Madhava Krishna
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.12811v1 Announce Type: cross Abstract: Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in purely topological representations has received relatively little attention. Importantly, loop closure over topolog...

📖 Read original article


148. Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques ​

Author: Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.12829v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized...

📖 Read original article


149. Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity ​

Author: Zhengfei Chen, Alex Welbourne, Matthew O. A. Ellis, Dan A. Allwood, Eleni Vasilaki, Thomas J. Hayward
Published: 7/15/2026, 4:00:00 AM
Categories: cs.ET, cond-mat.mes-hall, cs.AI, cs.LG

arXiv:2607.12840v1 Announce Type: cross Abstract: Unconventional computing systems must demonstrate robust performance under real-world environmental conditions to enable practical deployments. We have recently proposed superparamagnetic nanodot ensembles driven by strain-induced magnetoelectric cou...

📖 Read original article


150. ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation ​

Author: Jhen-Ke Lin
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.12857v1 Announce Type: cross Abstract: A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a...

📖 Read original article


151. Unveiling Complex Collective Behaviors from Simple Rewards ​

Author: Yize Mi, Jianan Li, Liang Li, Shiyu Zhao
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2607.12861v1 Announce Type: cross Abstract: Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic analysis, limiting multi-robot applications. Furthermore, complex swarm behaviors can surprisingly eme...

📖 Read original article


152. UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies ​

Author: Lirui Zhao, Modi Shi, Li Chen, Qi Liu, Ping Luo, Hongyang Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.12892v1 Announce Type: cross Abstract: Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect task completion, making the quality of these signals critical. Since such dense labels are rarely av...

📖 Read original article


153. Real-time fall detection based on vision for low-power edge platforms ​

Author: Wenjun Xia, Zhicheng Peng, Haopeng Li, Zhengdi Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CV

arXiv:2607.12909v1 Announce Type: cross Abstract: Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instabilit...

📖 Read original article


154. ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark ​

Author: Minh Hoang Nguyen
Published: 7/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.12946v1 Announce Type: cross Abstract: Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a resource is challenging for three reasons: cross-platform hotel names must be reconciled before interac...

📖 Read original article


155. Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models ​

Author: Mehmet Iscan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.12962v1 Announce Type: cross Abstract: Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexampl...

📖 Read original article


156. PalmClaw: A Native On-Device Agent Framework for Mobile Phones ​

Author: Hongru Cai, Yongqi Li, Ran Wei, Wenjie Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13027v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support too...

📖 Read original article


157. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale ​

Author: Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Tim Wang, Wei Zhan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2607.13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that lo...

📖 Read original article


158. DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents ​

Author: Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience. Despite its benefits, existing studies rely on hand-craft...

📖 Read original article


159. Rethinking Reward Models for Multi-Domain Test-Time Scaling ​

Author: Dong Bok Lee, Seanie Lee, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, Lei Song
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic. Prior work has studied both outcome reward model...

📖 Read original article


160. CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain? ​

Author: Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.09483v3 Announce Type: replace Abstract: While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored. CrochetBench presented in this paper evaluates this shift from describing to doing through fine-grained proce...

📖 Read original article


161. RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories ​

Author: Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, Rohit Gandikota
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.04144v3 Announce Type: replace Abstract: Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies...

📖 Read original article


162. JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks ​

Author: Lanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai, Yuanwu Xu, Lei Wei, Sicong Xie, Guannan Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.06486v4 Announce Type: replace Abstract: Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response strategies, while LLM-as-a-jud...

📖 Read original article


163. Calculating Mutual Information between a Reward Maximizer and its Environment ​

Author: Alfred Harwood, Jose Faustino, Alex Altair
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.12963v2 Announce Type: replace Abstract: An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of information an optimal policy provides about the underlying environment. ...

📖 Read original article


164. Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation ​

Author: Hua Yan, Heng Tan, Yingxue Zhang, Yu Yang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2602.16727v2 Announce Type: replace Abstract: Simulating large-scale human mobility is fundamental to understanding population movement patterns and supporting real-world geospatial applications such as urban planning, epidemic response, and transportation analysis. Recent works treat large la...

📖 Read original article


165. Learning When to Trust in Contextual Social Bandits ​

Author: Majid Ghasemi, Mark Crowley
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2603.13356v2 Announce Type: replace Abstract: Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget. We identify a more subtle failure mode that escapes this dichotomy, which we call \emph{Contextual Syco...

📖 Read original article


166. NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing ​

Author: Ming Yang, Zhi Zhou, Shi-Yu Tian, Kun-Yang Yu, Lan-Zhe Guo, Yu-Feng Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.16307v3 Announce Type: replace Abstract: Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex scenes and constraints and make reliable decisions. Current remote-sensing benchmarks mainly focus on evaluati...

📖 Read original article


167. ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing ​

Author: Yaopei Zeng, Congchao Wang, Blake JianHang Chen, Lu Lin
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.24787v2 Announce Type: replace Abstract: Routing has emerged as a promising strategy for balancing performance and cost in large language model (LLM) systems that combine lightweight models with powerful but expensive large models. Recent studies show that \emph{probe routing}, which pred...

📖 Read original article


168. Quantification of Credal Uncertainty: A Distance-Based Approach ​

Author: Xabier Gonzalez-Garcia, Siu Lun Chau, Julian Rodemann, Michele Caprio, Krikamol Muandet, Humberto Bustince, S'ebastien Destercke, Eyke H"ullermeier, Yusuf Sale
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, stat.ML

arXiv:2603.27270v2 Announce Type: replace Abstract: Credal sets, i.e., closed convex sets of probability measures, provide a natural framework to represent aleatoric and epistemic uncertainty in machine learning. Yet how to quantify these two types of uncertainty for a given credal set, particularly...

📖 Read original article


169. Mistake gating leads to energy and memory efficient continual learning ​

Author: Aaron Pache, Mark CW van Rossum
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.14336v2 Announce Type: replace Abstract: Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. However, when artificial neural networks are trained, the network parameters are typically updated on every sa...

📖 Read original article


170. Action-Aware Generative Sequence Modeling for Short Video Recommendation ​

Author: Wenhao Li, Zihan Lin, Zhengxiao Guo, Jie Zhou, Shukai Liu, Yongqi Liu, Chuan Luo, Chaoyi Ma, Ruiming Tang, Han Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2604.25834v2 Announce Type: replace Abstract: With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content consumption platforms. However, short videos often contain diverse segments, and users may not hold the same a...

📖 Read original article


171. Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR ​

Author: Sidi Chang, Peiying Zhu, Yuxiao Chen, Rongdong Chai
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.27374v2 Announce Type: replace Abstract: As LLMs become credible readers of earnings calls, investor-relations Q&A, guidance, and disclosure language, supervised financial NLP benchmarks increasingly function as decision evidence for model selection and deployment. A hidden assumption is...

📖 Read original article


172. Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination ​

Author: Qiyao Liang, Risto Miikkulainen, Ila Fiete
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.05686v3 Announce Type: replace Abstract: Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory, WM). We study two mechanistically distinct failure modes--conflict, when PM and WM disagree and interfere; a...

📖 Read original article


173. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World ​

Author: Pedro Conde, Henrique Branquinho, Valerio Mazzone, Bruno Mendes, Andr'e Baptista, Nuno Moniz
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2605.10834v2 Announce Type: replace Abstract: AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined g...

📖 Read original article


174. Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms ​

Author: Anni Zhao, Alex Bateh, Tyler Baldridge, Sandra Billinger, Xiao Hu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.16969v2 Announce Type: replace Abstract: Defining vascular age in terms of physiological function has become one focal point of the extensive studies to categorize and track chronological age. Transcranial Doppler (TCD) is a method by which cerebral blood flow velocity is measured along t...

📖 Read original article


175. How Inference Compute Shapes Frontier LLM Evaluation ​

Author: Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test tim...

📖 Read original article


176. Matilda: Engine-Agnostic Search with Human Policy Guidance ​

Author: Jason Carlson
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.25176v2 Announce Type: replace Abstract: Chess engines have evolved from search-based systems optimized solely for strength to neural policies capable of modeling human decisions across much of the rating spectrum. Maia-3, the strongest human-like move policy for chess, models the typical...

📖 Read original article


177. When Does Personality Composition Matter for Multi-Agent LLM Teams? ​

Author: Aryan Keluskar, Amrita Bhattacharjee, Huan Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.27443v2 Announce Type: replace Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored. Prior work shows that agents prompted with low agreeableness produce adversarial language...

📖 Read original article


178. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks ​

Author: Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Saaket Agashe, Xing Han Lu, Manpreet Kaur, Zhengyang Qi, Vincent Sunn Chen, Frederic Sala, Dayiheng Liu, Junyang Lin, Zhou Yu, Yu Su, Siva Reddy, Xin Eric Wang, Peng Qi, Tianbao Xie, Tao Yu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.29537v2 Announce Type: replace Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-ho...

📖 Read original article


179. DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models ​

Author: Xi Fang, Weijie Xu, Yingqiang Ge, Yuhui Xu, Stephanie Eckman, Chandan K. Reddy
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.02374v2 Announce Type: replace Abstract: Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then inject...

📖 Read original article


180. Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale ​

Author: Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li, Jiale Bai, Gao Cong
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.06233v2 Announce Type: replace Abstract: LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise set...

📖 Read original article


181. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation ​

Author: Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, Sergey Nikolenko
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE

arXiv:2607.06624v2 Announce Type: replace Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how t...

📖 Read original article


182. Infinity-Parser2 Technical Report ​

Author: Zuming Huang, Jun Huang, Kexuan Ren, Baode Wang, Weizhen Li, Jianming Feng, Yu Wang, Yichen Yao, Shijun Lin, Yige Tang, Cheng Peng, Weidi Xu, Wei Chu, Yinghui Xu, Yuan Qi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.07836v3 Announce Type: replace Abstract: We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing cor...

📖 Read original article


183. MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation ​

Author: Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.10079v2 Announce Type: replace Abstract: Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking ...

📖 Read original article


184. IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation ​

Author: Hengquan Guo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10144v2 Announce Type: replace Abstract: Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Yet most existing resources capture i...

📖 Read original article


185. Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents ​

Author: Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10526v2 Announce Type: replace Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting ...

📖 Read original article


186. QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics ​

Author: Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11019v2 Announce Type: replace Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics c...

📖 Read original article


187. Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans ​

Author: Jasmin Thelen, Oliver Wilhelm
Published: 7/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11263v3 Announce Type: replace Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measureme...

📖 Read original article


188. Propheticus: Machine Learning Framework for the Development of Predictive Models for Reliable and Secure Software ​

Author: Jo~ao R. Campos, Marco Vieira, Ernesto Costa
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:1809.01898v2 Announce Type: replace-cross Abstract: The growing complexity of software calls for innovative solutions that support the deployment of reliable and secure software. Machine Learning (ML) has shown its applicability to various complex problems and is frequently used in the dependa...

📖 Read original article


189. Diversity-Enriched Option-Critic ​

Author: Anand Kamat, Doina Precup
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2011.02565v2 Announce Type: replace-cross Abstract: Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The option-critic framework has been demonstrated to learn temporally extended actions, represented as opt...

📖 Read original article


190. Seeing Through Uncertainty: Free-Energy-Inspired Real-Time Adaptation for Robust Visual Navigation ​

Author: Maytus Piriyajitakonkij, Rishabh Dev Yadav, Mingfei Sun, Mengmi Zhang, Wei Pan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2403.01977v4 Announce Type: replace-cross Abstract: Navigation in the natural world is a feat of adaptive inference, where biological organisms maintain goal-directed behaviour despite noisy and incomplete sensory streams. Central to this ability is the Free Energy Principle (FEP), which posit...

📖 Read original article


191. Enabling Energy-Efficient Simultaneous Multi-Task Reinforcement Learning through Spiking Neural Networks with Active Dendrites for Bio-inspired Generalist Agents ​

Author: Rachmad Vidya Wicaksana Putra, Avaneesh Devkota, Muhammad Shafique
Published: 7/15/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG

arXiv:2412.04847v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has demonstrated remarkable capabilities in training agents to solve complex tasks autonomously, such as mobile robots, UAVs/UGVs, and game-playing agents). However, scaling RL to master multiple tasks simultaneous...

📖 Read original article


192. Modeling Story Expectations: A Generative Framework using LLMs ​

Author: Hortense Fong, George Gui, Bo Yang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, econ.GN, q-fin.EC, stat.ME

arXiv:2412.15239v4 Announce Type: replace-cross Abstract: Consumers' engagement with stories is shaped by their expectations about what will happen next, yet modeling these forward-looking beliefs over unstructured narrative content has remained challenging. We develop a framework that uses large la...

📖 Read original article


193. Toward Metaphor-Fluid Conversation Design for Voice User Interfaces ​

Author: Smit Desai, Jessie Chin, Dakuo Wang, Benjamin Cowan, Michael Twidale
Published: 7/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY, cs.ET

arXiv:2502.11554v3 Announce Type: replace-cross Abstract: Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, human-centric metaphors that fail to adapt to diverse contexts and user needs. This paper introduces Meta...

📖 Read original article


194. Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI ​

Author: Santhosh Parampottupadam, Melih Co\c{s}\u{g}un, Sarthak Pati, Maximilian Zenk, Saikat Roy, Dimitrios Bounias, Benjamin Hamm, Sinem Sav, Ralf Floca, Klaus Maier-Hein
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.DC

arXiv:2505.22108v4 Announce Type: replace-cross Abstract: Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is limited by privacy concerns, heterogeneous institutional compliance, and resource disparities; standar...

📖 Read original article


195. SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation ​

Author: Xi Cheng, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Lyuhao Chen, Brian Zhu, Daniel Jin, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao
Published: 7/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2506.12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions. In this paper, we introduce a hierarchical agentic system consisting of three speciali...

📖 Read original article


196. Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers ​

Author: Ian Chuang, Jinyu Zou, Andrew Lee, Dechen Gao, Iman Soltani
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2507.15833v3 Announce Type: replace-cross Abstract: Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform proces...

📖 Read original article


197. Real-Time Model Checking for Closed-Loop Robot Reactive Planning ​

Author: Christopher Chandler, Bernd Porr, Giulia Lafratta, Alice Miller
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.FL

arXiv:2508.19186v2 Announce Type: replace-cross Abstract: Reactive obstacle avoidance methods often cause agents to become trapped in local minima, because they can often only reason one step ahead (i.e., the next action based on the current state). In this paper, we use model checking to achieve re...

📖 Read original article


198. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models ​

Author: Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.22415v4 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token hidden state into t...

📖 Read original article


199. Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning ​

Author: Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, Risto Miikkulainen
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2509.24372v3 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art LLMs. In contrast, e...

📖 Read original article


200. Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task ​

Author: Brady Bhalla, Honglu Fan, Nancy Chen, Tony Yue YU
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.18315v2 Announce Type: replace-cross Abstract: We investigate how embedding dimension affects the emergence of an internal "world model" in a transformer trained with reinforcement learning to perform bubble-sort-style adjacent swaps. Models achieve high accuracy even with very small embe...

📖 Read original article


201. Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks ​

Author: Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, Nitesh V. Chawla
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over fixed item sets, t...

📖 Read original article


202. A Neurosymbolic Approach to Natural Language Formalization and Verification ​

Author: Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, R'emi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovi'c, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joseph Lilien, Divya Raghunathan, Zvonimir Rakamari'c, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.LO

arXiv:2511.09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and health-care that operate under strict policies...

📖 Read original article


203. Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning ​

Author: Dongyue Li, Zhenshuo Zhang, Minxuan Duan, Edgar Dobriban, Hongyang R. Zhang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS

arXiv:2512.01113v2 Announce Type: replace-cross Abstract: Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reasoning abilities, designed for graph neural networks and also for transformer models. Prior work has evalu...

📖 Read original article


204. First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations ​

Author: David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk i...

📖 Read original article


205. Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning ​

Author: Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.19624v3 Announce Type: replace-cross Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unans...

📖 Read original article


206. PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models ​

Author: Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi, Johannes Hoffart, Carlos Guestrin, Jure Leskovec
Published: 7/15/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG

arXiv:2602.04029v2 Announce Type: replace-cross Abstract: Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the diverse relational databases needed to train such models are rarely public due to privacy constraints. Whi...

📖 Read original article


207. DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter ​

Author: Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng, Yuan Meng, Haojun Jiang, Shaoxuan Xie, Guocai Yao, Alois Knoll, Zhenshan Bing, Xinlong Wang, Zhenguo Sun
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2602.05513v3 Announce Type: replace-cross Abstract: Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of effectively combining these modalities, we propose DECO, a decoupled multimodal diffusion transformer th...

📖 Read original article


208. Self-Regulated Reading with AI Support: An Eight-Week Study with Students ​

Author: Yue Fu, Joel Wester, Niels Van Berkel, Alexis Hiniker
Published: 7/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2602.09907v2 Announce Type: replace-cross Abstract: College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement. We conducted an eight-week longitudinal study wit...

📖 Read original article


209. Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability ​

Author: Shafiuddin Rehan Ahmed, Sourabh Deshpande
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2602.11198v2 Announce Type: replace-cross Abstract: Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding assistants can generate correct, framework-specific code. We introduce \textit{AI-assistability} ($\mathca...

📖 Read original article


210. SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise ​

Author: Yuejie Li, Ke Yang, Yueying Hua, Berlin Chen, Jianhao Nie, Yueping He, Caixin Kang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2602.12783v3 Announce Type: replace-cross Abstract: Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often limited to simple queries under constrained noise conditions, making them inadequate for assessing the ro...

📖 Read original article


211. Egocentric Bias in Vision-Language Models ​

Author: Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.15892v2 Announce Type: replace-cross Abstract: Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. T...

📖 Read original article


212. Xray-Visual Models: Scaling Vision models on Industry Scale Data ​

Author: Shlok Mishra, Tsung-Yu Lin, Linda Wang, Hongli Xu, Yimin Liu, Michael Hsu, Chaitanya Ahuja, Hao Yuan, Jianpeng Cheng, Hong-You Chen, Haoyuan Xu, Chao Li, Sreya Dutta Roy, Abhijeet Awasthi, Jihye Moon, Don Husa, Michael Ge, Sumedha Singla, Arkabandhu Chowdhury, Phong Dingh, Satya Narayan Shukla, Yonghuan Yang, David Jacobs, Qi Guo, Jun Xiao, Xiangjun Fan, Aashu Singh
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.16918v2 Announce Type: replace-cross Abstract: We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 billion curated image-text pairs and 10 billion video-hashtag p...

📖 Read original article


213. 1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World ​

Author: Qiao Xu, Yipeng Yu, Chengxiao Feng, Xu Liu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2602.18548v3 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inconsistent datasets, toolchains, and evaluation protocols. We introduce 1D-Bench, a benchmark grounded...

📖 Read original article


214. Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast ​

Author: Mehmet Yigit Avci (and for the Alzheimer's Disease Neuroimaging Initiative), Akshit Achara (and for the Alzheimer's Disease Neuroimaging Initiative), Andrew King (and for the Alzheimer's Disease Neuroimaging Initiative), Jorge Cardoso (and for the Alzheimer's Disease Neuroimaging Initiative)
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.04113v2 Announce Type: replace-cross Abstract: Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisition characteristics have been shown to contribute substantially to this predictability. Whether the sam...

📖 Read original article


215. TADPO: Reinforcement Learning Goes Off-road ​

Author: Zhouchonghao Wu, Raymond Song, Vedant Mundheda, Luis E. Navarro-Serment, Christof Schoenborn, Jeff Schneider
Published: 7/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2603.05995v2 Announce Type: replace-cross Abstract: Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics. Addressing these challenges requires effective long-horizon planning and adaptable control. Reinforcem...

📖 Read original article


216. Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts ​

Author: Ali Safari, Sahar Babaei
Published: 7/15/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.IR

arXiv:2603.22510v3 Announce Type: replace-cross Abstract: Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by changes in research novelty. We examine how relative abstract-level semantic novelty in Information Syst...

📖 Read original article


217. Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory ​

Author: Jon-Paul Cacioli
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.25112v2 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We app...

📖 Read original article


218. Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering ​

Author: Yanjie Zhang, Yafei Li, Rui Sheng, Zixin Chen, Yanna Lin, Huamin Qu, Lei Chen, Yushi Sun
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2603.28583v2 Announce Type: replace-cross Abstract: Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designe...

📖 Read original article


219. Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces ​

Author: Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.11996v4 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a fundamental limitati...

📖 Read original article


220. Neuro-Symbolic ODE Discovery with Latent Grammar Flow ​

Author: Karin Yu, Eleni Chatzi, Georgios Kissas
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.SC

arXiv:2604.16232v2 Announce Type: replace-cross Abstract: Understanding natural and engineered systems often relies on symbolic formulations, such as differential equations, which provide interpretability and transferability beyond black-box models. We introduce Latent Grammar Flow (LGF), a neuro-sy...

📖 Read original article


221. The TIME Machine: On The Power of Motion for Efficient Perception ​

Author: Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2605.23045v3 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have p...

📖 Read original article


222. Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory ​

Author: Quanjiang Li, Zhiming Liu, Wei Luo, Tingjin Luo, Chenping Hou
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.24602v5 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated ...

📖 Read original article


223. dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ​

Author: Giuseppe Franco, Ian Colbert, Pablo Monteagudo-Lago, Felix Marty, Nicholas Fraser
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.04115v2 Announce Type: replace-cross Abstract: Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy. ...

📖 Read original article


224. The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport ​

Author: Peter Halmos, Boris Hanin
Published: 7/15/2026, 4:00:00 AM
Categories: math-ph, cs.AI, cs.LG, math.MP, physics.data-an

arXiv:2606.05217v3 Announce Type: replace-cross Abstract: We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family of Schr"odinger operators we call Score Hamiltonians, built from the learned score's quantum potentia...

📖 Read original article


225. Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles ​

Author: Upasana Chatterjee
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.06715v2 Announce Type: replace-cross Abstract: We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-ve...

📖 Read original article


226. Explaining Data Mixing Scaling Laws ​

Author: Rui Dai, Shuran Zheng
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.08167v2 Announce Type: replace-cross Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understanding of these model loss behaviors remains absent. In this work, we propose a unified framework...

📖 Read original article


227. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models ​

Author: Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.10829v2 Announce Type: replace-cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictio...

📖 Read original article


228. ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories ​

Author: Siyuan Luo, Nairong Zheng, Lin Zhou, Tiankuo Yao, Shengyou Yuan, Haojia Yu, Cong Pang, Jiapeng Luo, Lewei Lu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.11520v4 Announce Type: replace-cross Abstract: Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution--properties absent from existing datasets. We propose ISE (Intent -> Simulate -> Execute), ...

📖 Read original article


229. Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy ​

Author: Kai Standvoss, Miriam H"agele, Rosemarie Krupar, Julika Ribbat-Idel, Jennifer Altsch"uler, Gerrit Erdmann, Hans Pinckaers, Evelyn Ramberger, Madleen Drinkwitz, 'Ad'am N'arai, Alexander M"ollers, Katja Lingelbach, Sebastian Kons, Lukas H"onig, Recepcan Adig"uzel, Joana Bai~ao, Alberto Megina Gonzalo, Marius Teodorescu, Marie-Lisa Eich, Paolo Chetta, Shakil Merchant, Verena Aumiller, Simon Schallenberg, Andrew Norgan, Klaus-Robert M"uller, Lukas Ruff, Maximilian Alber, Frederick Klauschen
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2606.12346v2 Announce Type: replace-cross Abstract: Hematoxylin and eosin (H&E) staining is the cornerstone of histopathology, yet scalable, quantitative analysis of H&E whole-slide images (WSIs) remains a central challenge in computational pathology. We present Atlas H&E-TME, an AI-based syst...

📖 Read original article


230. Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents ​

Author: Tianyu Ding, Jianhong Xin, Juan Pablo De la Cruz Weinstein
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2606.12634v3 Announce Type: replace-cross Abstract: Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning, API, and answer tokens. Direct self-distillation can supply a denser signal, but in our experiments i...

📖 Read original article


231. Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens ​

Author: Yizhen Yao, Qinglin Zhu, Runcong Zhao, Xiangxiang Dai, Yanzheng Xiang, Yulan He, Lin Gui
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.16847v3 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens...

📖 Read original article


232. Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement ​

Author: You Li, Samuel Mandell, David Z. Pan
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2606.19387v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high stakes in chip design ...

📖 Read original article


233. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers ​

Author: Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang, Zhenbo Li, Jean-Yves Guillemaut, Wenwu Wang
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM

arXiv:2606.20101v3 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-bas...

📖 Read original article


234. Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic ​

Author: Hyunsang Hwang, Suhyun Bae, Donghun Lee
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.23044v2 Announce Type: replace-cross Abstract: Numbers have algebraic structure that standard neural embeddings often fail to expose. We introduce Prime Fourier Embeddings (PFE), which encode integers as prime-indexed (cos, sin) pairs derived from the harmonic analysis of Q, providing a p...

📖 Read original article


235. Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking ​

Author: Mohamed Nagy, Naoufel Werghi, Jorge Dias, Majid Khonji
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.23604v4 Announce Type: replace-cross Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. However, these descriptors are frame-independent, limiting their robustness as visual cues. S...

📖 Read original article


236. Pigeonholing: how bad prompts hurt models, causing collapse and mistakes ​

Author: Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.24267v2 Announce Type: replace-cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." Unintentionally bad contexts can happen w...

📖 Read original article


237. Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model ​

Author: Sergey Kurilenko
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.IR

arXiv:2606.26373v4 Announce Type: replace-cross Abstract: Semantic search creates an asymmetric disclosure problem: query embeddings may reveal user intent, while returning exact provider vectors distributes reusable representations. We evaluate a deliberately restricted hybrid design. A public corp...

📖 Read original article


238. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders ​

Author: Nathana"el Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.27321v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features. The Top-$k$ SAE, a now...

📖 Read original article


239. SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval ​

Author: Sergey Kurilenko
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.IR

arXiv:2606.27976v3 Announce Type: replace-cross Abstract: Dense retrieval systems expose document geometry when vector stores are compromised, and a global protective transform can often be aligned from known pairs. We study SHARD, which splits PCA coordinates into a short routing prefix and a resid...

📖 Read original article


240. RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources ​

Author: Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2606.29538v2 Announce Type: replace-cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial vid...

📖 Read original article


241. FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images ​

Author: Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.30347v2 Announce Type: replace-cross Abstract: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed numbe...

📖 Read original article


242. SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents ​

Author: Jingyuan Zheng, Dongjing Wang, Xin Zhang, Hao Chen, Youhuizi Li, Xudong Shen, Haiping Zhang, Butian Huang, Dongjin Yu, Guandong Xu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.SE

arXiv:2607.00011v2 Announce Type: replace-cross Abstract: Reusable agent skills are emerging as a service-oriented capability layer for Large Language Model (LLM) agents. Unlike plain retrieval items, a skill exposes functional capabilities, input-output assumptions, tool dependencies, context cost,...

📖 Read original article


243. Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation ​

Author: Fengchong Yao, Jianbing Li, Qing Liu, Kefeng Song, Haitao Li, Song Wang, Feixiang Wang
Published: 7/15/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.02567v4 Announce Type: replace-cross Abstract: Radio frequency fingerprint identification (RFFI) provides a critical physical-layer security mechanism for dynamic Internet of Things (IoT) and ad hoc networks. However, the decentralized and open nature of these networks imposes two strict ...

📖 Read original article


244. PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition ​

Author: Dana A Abdullah
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.03581v2 Announce Type: replace-cross Abstract: The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of conventional face recognition systems. Existing approaches relying on fixed cosine thresholds, non-adapti...

📖 Read original article


245. Git-Assistant: Planning-Based Support for Updating Git Repositories ​

Author: Alfredo Garrach'on Ruiz, Tom'as de la Rosa, Daniel Borrajo
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.09224v2 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Recent advances in Large Language Models (LLMs) offer promising capabilities for interpreting develope...

📖 Read original article


246. ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams ​

Author: Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09759v2 Announce Type: replace-cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attr...

📖 Read original article


247. Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting ​

Author: Fengchong Yao, Jianbing Li, Qing Liu, Qikun Liu, Kefeng Song, Haitao Li, Song Wang
Published: 7/15/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.09760v2 Announce Type: replace-cross Abstract: Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for Internet of Things (IoT) devices, but deep models often degrade across acquisition environments. In mult...

📖 Read original article


248. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization ​

Author: Ethan Smith
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09967v2 Announce Type: replace-cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps. Adaptive optimizers such as Adam norma...

📖 Read original article


249. Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices ​

Author: Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li
Published: 7/15/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG

arXiv:2607.10183v2 Announce Type: replace-cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory. Ex...

📖 Read original article


250. Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter ​

Author: Achyuthan Sivasankar
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10203v2 Announce Type: replace-cross Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively. In autoregressive rollouts, the first assumption requires dep...

📖 Read original article


251. ABot-N1: Toward a General Visual Language Navigation Foundation Model ​

Author: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Yang Cai, Jingjing Ma, Shihui Su, Zixiao Tang, Linbo Zheng, Zedong Chu, Xiaolong Wu, Ziqiao Li, Mu Xu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2607.10383v2 Announce Type: replace-cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map ...

📖 Read original article


252. AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP ​

Author: Aritra Mazumder, Nusrat jahan Lia
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.11098v3 Announce Type: replace-cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, a...

📖 Read original article


253. RepTran: Search-Based Repair of Transformer Models ​

Author: Yuta Ishimoto, Paolo Arcaini, Fuyuki Ishikawa, Masanari Kondo, Naoyasu Ubayashi, Yasutaka Kamei
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11193v2 Announce Type: replace-cross Abstract: To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and repaired. Among AI components, Transformer models are increasingly integrated into software systems, w...

📖 Read original article


254. An Empirical Study for Android-to-OpenHarmony GUI Test Migration ​

Author: Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye
Published: 7/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.11245v2 Announce Type: replace-cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing GUI test cases has become a critical problem. However, current research neither proposes solutions ta...

📖 Read original article


255. PRISM Edit: One Vector for All Temporal Answers ​

Author: Chen Huang, Qi Zheng, Ruiqin Zheng, Long Zeng, Yuantong Xu
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11327v2 Announce Type: replace-cross Abstract: Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer shou...

📖 Read original article


256. Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks ​

Author: Tiberiu Musat, Tiago Pimentel, Nicolas Zucchet, Thomas Hofmann
Published: 7/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11875v2 Announce Type: replace-cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a gene...

📖 Read original article