Skip to content

arXiv cs.AI - 2026-07-16 ​

228 items collected.


1. OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets ​

Author: Haolin Xue
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13037v1 Announce Type: new Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or dataset le...

📖 Read original article


2. SPINE: Bridging the Cyber-Physical Gap with Agentic AI ​

Author: Minkyu Ham, Dongho Kim, Chan Lee, Jiayi Wang, Min Jun Kim, Yixi Zhang, Guo Ye, Jihai Zhao, Soyeon Park, Han Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.13049v1 Announce Type: new Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still demands tedious, expert-driven calibration. This deployment gap, the robot's spinal cord, remains a pr...

📖 Read original article


3. Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution ​

Author: Hironao Nakamura
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LO

arXiv:2607.13069v1 Announce Type: new Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we interv...

📖 Read original article


4. Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL ​

Author: Zoran Majkic
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13073v1 Announce Type: new Abstract: Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like lack of interpretability and logical structure) with formal logical machinery for self-reference. In t...

📖 Read original article


5. Self-Improvements in Modern Agentic Systems: A Survey ​

Author: Zhe Ren, Yimeng Chen, Dandan Guo, Guowei Rong, Tonghui Li, R. B. Xiong, Qingfeng Lan, Wenyi Wang, Li Nanbo, Yibo Yang, Mingchen Zhuge, J"urgen Schmidhuber
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.13104v1 Announce Type: new Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as...

📖 Read original article


6. Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools ​

Author: Konstantinos Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.13115v1 Announce Type: new Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues. We propose a mod...

📖 Read original article


7. Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents ​

Author: Richmond Alake, Cesare Bernardis, Paul Cayet, Luca Engel, Damien Hilloulin, Sungpack Hong, Allen Hosler, Nickolas Kavantzas, Ingo Kossyk, Son Le, Rhicheek Patra, Kartik Talamadupula, Valentin Venzin
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2607.13157v1 Announce Type: new Abstract: Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumulation of procedural knowledge fr...

📖 Read original article


8. Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models ​

Author: Ilias Kazantzidis, Timothy J. Norman, Yali Du, Christopher T. Freeman
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.13172v1 Announce Type: new Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical environments, we cons...

📖 Read original article


9. CayleyR: Solving the TopSpin puzzle via cycle intersection ​

Author: Yuri Baramykov
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13219v1 Announce Type: new Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs. The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation seq...

📖 Read original article


10. Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science ​

Author: Sutanay Choudhury, Jeffrey J. Czajka, Lummy M. O. Monteiro, Erin Bredeweg, Jason McDermott, Katherine Wolf, Alex Beliaev, Josh Elmore, Paul Piehowski, Kylee Tate, Yuqian Gao, Aivett Bilbao, Kelly Stratton, Scott Baker, Jaydeep P. Bardhan, Kristin Burnum Johnson, Chris Oehmen, Robert Rallo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.HC

arXiv:2607.13220v2 Announce Type: new Abstract: Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems a...

📖 Read original article


11. AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation ​

Author: Quanyan Zhu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY

arXiv:2607.13230v1 Announce Type: new Abstract: Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwritin...

📖 Read original article


12. Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management ​

Author: Xi Cheng, Ke Liu, Siyuan Feng, Jane Lin, H. Oliver Gao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center (TMC) tasks such as anomaly detection, incident reporting, and traveler information. Deploying multi...

📖 Read original article


13. Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable ​

Author: Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.13285v1 Announce Type: new Abstract: The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness...

📖 Read original article


14. Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases ​

Author: Marcus J. Min, Mike He, Zhaoyu Li, Zixuan Yi, Sharad Malik, Aarti Gupta, Xujie Si, Osbert Bastani
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.PL

arXiv:2607.13292v1 Announce Type: new Abstract: Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, definition...

📖 Read original article


15. EZSMT Version 3, Matured ​

Author: Yuliya Lierler
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13344v1 Announce Type: new Abstract: Constraint Answer Set Programming (CASP) is a hybrid reasoning paradigm that combines Answer Set Programming (ASP) with Constraint Processing and Satisfiability Modulo Theories (SMT), enabling powerful declarative encodings of complex combinatorial sea...

📖 Read original article


16. Set-shifting Behavioral Test for Harnessed Agents ​

Author: Ziwei Ye
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE

arXiv:2607.13396v1 Announce Type: new Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our benchmark mounts tool-skill ...

📖 Read original article


17. LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning ​

Author: Qiang Zhu, Jiajun Wu, Longyi Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13501v2 Announce Type: new Abstract: Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method base...

📖 Read original article


18. How Far Can Root Cause Analysis Go on Real-World Telemetry Data? ​

Author: Athira Gopal, Ashwanth Krishnan
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13548v1 Announce Type: new Abstract: Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA datase...

📖 Read original article


19. Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling ​

Author: Xixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang, Guangyin Jin, Song Gao, Yuxuan Liang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13558v1 Announce Type: new Abstract: Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic assessment, and environmental monitoring. Existing methods typically formulate this task as multimodal representation ...

📖 Read original article


20. AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized ​

Author: Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC

arXiv:2607.13562v1 Announce Type: new Abstract: Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions ...

📖 Read original article


21. SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing ​

Author: Tianyu Chen, Chujia Hu, Wenjie Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two dist...

📖 Read original article


22. Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System ​

Author: David Krongauz, Arad Zulti, Eran Segal, Teddy Lazebnik
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, math.DS

arXiv:2607.13608v1 Announce Type: new Abstract: Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, moving computational systems beyond data-fitting tools toward the generation and refinement of mechanistic models o...

📖 Read original article


23. STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle ​

Author: Sagar Deb, Ashwanth Krishnan
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the final cost cannot say why an agent failed: it may have misread the world, or read it correctly and stil...

📖 Read original article


24. UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following ​

Author: Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic re...

📖 Read original article


25. Explaining Reinforcement Learning Agents via Inductive Logic Programming ​

Author: Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in safety-critical and human-centric scenarios. However, it is mostly based on user studies, thus targetin...

📖 Read original article


26. When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects ​

Author: Yongren Shi, Wenyi Gong
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.13679v1 Announce Type: new Abstract: AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code...

📖 Read original article


27. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities ​

Author: Kai Chen, Zichen Ding, Jiaye Ge, Shufan Jiang, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.13705v2 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing re...

📖 Read original article


28. CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems ​

Author: Zexun Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13716v1 Announce Type: new Abstract: Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateways, and workflow engines. A single operational act such as publishing code, changing identity state, ...

📖 Read original article


29. Experience Memory Graph: One-Shot Error Correction for Agents ​

Author: Wenjun Wang, Yuchen Fang, Fengrui Liu, Zibo Liang, Kai Zheng
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13884v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer from ...

📖 Read original article


30. AIMO Interpretability Challenge ​

Author: Michal \v{S}tef'anik, Philipp Mondorf, Andreas Waldis, Qianying Liu, Chuan Yang, Michal Spiegel, Josef Kucha\v{r}, Marek Kadl\v{c}'ik, Adam Vawda-Oomerjee, Chaoran Liu, Simon Frieder, Barbara Plank, Fazl Barez, Pontus Stenetorp
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13899v1 Announce Type: new Abstract: We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms. The challenge is motivated by a central limitation of stan...

📖 Read original article


31. A Self-Evolving Agent for Longitudinal Personal Health Management ​

Author: Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements a...

📖 Read original article


32. Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0 ​

Author: Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.14004v1 Announce Type: new Abstract: Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for d...

📖 Read original article


33. AI-accelerated End-to-End Framework for Rapid Professional Upskilling ​

Author: Tam Nguyen, Hung Nguyen, Robert Ogburn
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.14044v1 Announce Type: new Abstract: By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skills gap grew from roughly 3 days in 2014 to 36 days in 2018. Most current frameworks accelerate single stages of upskilling programs and...

📖 Read original article


34. Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education ​

Author: Xanthi Kokkinou, Chaido Mizeli, Nafsika Koulaxidou, Marina Delianidi, Konstantinos Diamantaras
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.14046v1 Announce Type: new Abstract: This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake p...

📖 Read original article


35. Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models ​

Author: Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu, Jie Li
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.14049v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating ano...

📖 Read original article


36. FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents ​

Author: Srihari Unnikrishnan, Jaskaran Singh Walia, Drishti Goel, Supriyo Ghosh
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SE

arXiv:2607.13035v1 Announce Type: cross Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage and outdated documen...

📖 Read original article


37. Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry ​

Author: Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarifica...

📖 Read original article


38. Designing Safety-Constrained LLM Systems for Public Health Information Access ​

Author: Ben Torkian, Jun Zhou
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13038v1 Announce Type: cross Abstract: We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on maternal and child health (MCH) resource navigation. While LLM based systems offer flexible and natur...

📖 Read original article


39. Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants ​

Author: Dipesh Tharu Mahato
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a deployment question: for a fixed base model, how does the access condition users actually see change ...

📖 Read original article


40. Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance ​

Author: Zexun Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13040v1 Announce Type: cross Abstract: This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two governance models. The first, frontier-provider sovereignty, assigns privileged authority to the provider of the mo...

📖 Read original article


41. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents ​

Author: Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG, cs.MM

arXiv:2607.13041v1 Announce Type: cross Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 64...

📖 Read original article


42. Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning ​

Author: Daniel Vila-Cruz, Laura Mor'an-Fern'andez, Ver'onica Bol'on-Canedo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.13043v1 Announce Type: cross Abstract: Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands. We propose a lightweight training strategy that adapts normalization layers of the model to the new domai...

📖 Read original article


43. The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI ​

Author: Anubhab Banerjee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13044v1 Announce Type: cross Abstract: The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted content under Article 83 and Rule 42, creating pressure to triage suspected AI-generated patent text. ...

📖 Read original article


44. Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges ​

Author: Masoume Gholizade, Fabrizio Ruffini, Pietro Ducange, Francesco Marcelloni
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13045v1 Announce Type: cross Abstract: Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data sources. By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolv...

📖 Read original article


45. Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems ​

Author: Zhaohui Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at substantial cost. The central question of when to invoke the LLM has received limited formal treatme...

📖 Read original article


46. Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review ​

Author: Sebastian Jouannet-Contreras, Carola Figueroa-Flores
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.13054v1 Announce Type: cross Abstract: Environmental monitoring with unmanned aerial vehicles (UAVs) requires route planning methods that maximize covered area while handling energy limits, operational constraints, and geometric complexity. This paper reports the protocol and preliminary ...

📖 Read original article


47. Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes ​

Author: Hiroki Tamba
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.13071v1 Announce Type: cross Abstract: Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper documents a failure mode in Claude Code where partial standard output from timed-out commands (exit code 1...

📖 Read original article


48. HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models ​

Author: Luyuan Jia, Yinfeng Yu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, eess.SP

arXiv:2607.13072v1 Announce Type: cross Abstract: Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-trained large models are u...

📖 Read original article


49. When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data ​

Author: Giansalvo Cirrincione, Filippo Grassia
Published: 7/16/2026, 4:00:00 AM
Categories: math.NA, cond-mat.mtrl-sci, cs.AI, cs.LG, cs.NA

arXiv:2607.13074v1 Announce Type: cross Abstract: This work studies the inverse problem of recovering the relative magnitudes of the tension, bending, and bearing loads acting on a crack from its stress-intensity-factor profile along the crack front, using the public SIFBench finite-element data. Th...

📖 Read original article


50. The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators ​

Author: Dominik Schwarz
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model f...

📖 Read original article


51. The Hitchhiker's Guide to Monoculture ​

Author: Gordon Burtch
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SE

arXiv:2607.13077v1 Announce Type: cross Abstract: Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interac...

📖 Read original article


52. Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows ​

Author: Keyur Gabani
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational ...

📖 Read original article


53. Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs ​

Author: Sheng-Wei Peng, Yi-Hsun Lin, Yi-Pei Lee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.13080v1 Announce Type: cross Abstract: Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost -- and on-premise quantized open-weights models, which promise low-marginal-cost scaling and data sovereignty ...

📖 Read original article


54. SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification ​

Author: SingGuard Team
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG

arXiv:2607.13081v1 Announce Type: cross Abstract: We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. We first int...

📖 Read original article


55. Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing ​

Author: Ananda Dhakal, Krish Neupane, Aarjan Chaudhary
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13085v1 Announce Type: cross Abstract: Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier LLMs. Because these systems often change both architecture and backbone model, it is difficult to tell how much ...

📖 Read original article


56. Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework ​

Author: Aditya Aggarwal, Nahid Farhady Ghalaty
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.13091v1 Announce Type: cross Abstract: LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feedback. We present a closed-loop framework in which every accepted review comment is codified as a per...

📖 Read original article


57. Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ​

Author: Yi Li, Chen Li, Jiexiong Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13093v1 Announce Type: cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeas...

📖 Read original article


58. Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates ​

Author: Lynn Vonderhaar, Juan Couder, Siri Siqveland, Omar Ochoa, James Pembridge
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13094v1 Announce Type: cross Abstract: The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive tasks can be completed much more efficiently with the use of AI. This work uses AI techniques to analyze and r...

📖 Read original article


59. Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit ​

Author: Xiaomi MiMo Team, Anqi Liu, Aoxin Ma, Bo Chen, Bo Yang, Chen Wang, Chen Zhang, Chengda Tang, Chengwei Wang, Chiheng Lou, Depeng Yan, Fuli Luo, Gang Wang, Hailin Zhang, Jiale Sun, Kang Zhou, Rui Huang, Shaohui Liu, Shen Huang, Shijie Cao, Shuaishuai Fan, Tianling Zhou, Xiangwei Deng, Xueyang Xie, Xuli Wang, Yingchun Lai, Yu Yang, Yuan Zhang, Zhen Tang, Zhonghua Deng, Zihan Jiang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.13095v1 Announce Type: cross Abstract: We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attentio...

📖 Read original article


60. WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency ​

Author: Z Sun, Q Jiang, S Sheng, L Xiang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13099v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques. However, these techniques have rarely been adopted in practice mai...

📖 Read original article


61. TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling ​

Author: Songru Yang, Zili Liu, Tao Han, Ben Fei, Fenghua Ling, Lei Bai, Chang Liu, Xiangyang Ji, Zhenwei Shi, Zhengxia Zou
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13101v1 Announce Type: cross Abstract: Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error ...

📖 Read original article


62. Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing ​

Author: Duantengchuan Li, Yingqian Bi, Jinsong Chen, Rui Zhang, Mingwen Tong
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13103v1 Announce Type: cross Abstract: Knowledge tracing (KT) aims to predict students' future performance by modeling their evolving knowledge states from historical interactions. Existing KT methods usually treat the raw interaction sequence as a unified behavioral process, overlooking ...

📖 Read original article


63. STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting ​

Author: Sicong Lai, Yuehong Hu, Siru Zhong, Si Qiao, Yuxuan Liang, Guangyin Jin
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13108v1 Announce Type: cross Abstract: Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention...

📖 Read original article


64. A Hybrid Mamba for Audio-Visual Navigation ​

Author: Yi Wang, Yinfeng Yu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2607.13110v1 Announce Type: cross Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them ...

📖 Read original article


65. SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests ​

Author: Maha Ayub, Michael Konstantinou, Ahmed Khanfir, Nikolaos Tsantalis, Mike Papadakis
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.13111v1 Announce Type: cross Abstract: Distinguishing semantic-preserving commits from changing ones remains an open challenge in software repository mining. While existing approaches detect refactoring commits accurately, they cannot ensure that a commit is purely semantic-preserving, wi...

📖 Read original article


66. CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion ​

Author: Jiaze Song, Runhao Zhao, Minghao Xu, Bin Cui, Wentao Zhang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13120v1 Announce Type: cross Abstract: Inferring gene regulatory networks (GRNs) from single-cell transcriptomic data is crucial for biological discovery, yet existing approaches suffer from a fundamental misalignment with real-world needs. Researchers typically seek a small set of high-c...

📖 Read original article


67. AI in Cyberpsychology: A systematic literature review of Cybersecurity enhancement by using AI for analyzing psychology of Victims, Attackers, and Defenders ​

Author: Georg Thamer Francis, Malek Malkawi, Sevim Ey"upo\u{g}lu, Reda Alhajj, Selim Akyoku\c{s}
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13123v1 Announce Type: cross Abstract: Cybersecurity is the practice of protecting systems, networks, and data from digital attacks. Cyberpsychology (CPSY) is defined as the use of psychology to enhance cybersecurity applications. Since the early 2010s, the evolution of Artificial Intelli...

📖 Read original article


68. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation ​

Author: Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two obs...

📖 Read original article


69. Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation ​

Author: Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13125v1 Announce Type: cross Abstract: We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast infer...

📖 Read original article


70. Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach ​

Author: Tianyu Pang, Hongyu Li
Published: 7/16/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, eess.SP, math.IT

arXiv:2607.13160v1 Announce Type: cross Abstract: Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offl...

📖 Read original article


71. What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors ​

Author: Winston Zeng, Ali Emami, Jinho D. Choi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13162v2 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organiz...

📖 Read original article


72. SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy ​

Author: Yassine Chemingui, Chenhua Fan, Honghao Wei, Janardhan Rao Doppa
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate, a boundary-aware di...

📖 Read original article


73. RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar ​

Author: Marek \v{S}uppa, Vikt'oria Ondrejov'a, Lucia Ganajov'a, Gregor Karetka, Daniel Skala
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13189v1 Announce Type: cross Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chinese). RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pi...

📖 Read original article


74. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference ​

Author: Soumil Mandal
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.13205v1 Announce Type: cross Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-dense inpu...

📖 Read original article


75. Classifying daily activities needs posture, reconstructing them needs motion ​

Author: Arefeh Farahmandi, Gunnar Blohm
Published: 7/16/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.IR, cs.LG

arXiv:2607.13216v1 Announce Type: cross Abstract: Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to...

📖 Read original article


76. Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift ​

Author: Jayakumar Manoharan
Published: 7/16/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.LG, cs.SY

arXiv:2607.13221v1 Announce Type: cross Abstract: Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical guarantee and can silentl...

📖 Read original article


77. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System ​

Author: Ken Jon Miyachi, Dylan Uys
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13234v1 Announce Type: cross Abstract: Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations report AUC drops of 45-50% for state-of-the-art open-source models. We argue this gap is structural: static dete...

📖 Read original article


78. EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting ​

Author: Mingxing Xu, Rakesh Chowdary Machineni, Ke Liu, Xi Cheng, Chengqi Lu, Xin Hu, Lyuhao Chen, Xiangyu Li, Junwei You, Oliver Gao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13241v1 Announce Type: cross Abstract: Traffic forecasting is highly challenging due to complex and nonlinear spatial and temporal dependencies. Self-attention mechanisms have been widely adopted to model dynamic and long-range dependencies, achieving state-of-the-art performance, but suf...

📖 Read original article


79. Reassessing Muon for Matrix Factorization ​

Author: Ali Parviz, Gal Mishne, Alex Cloninger
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical...

📖 Read original article


80. Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance ​

Author: Stylianos Loukas Vasileiou, Olga Derendiaeva
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13260v1 Announce Type: cross Abstract: Policy documents shape governance outcomes, but their reasoning is often implicit. Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly. Existing computational a...

📖 Read original article


81. Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners ​

Author: Haseeb Shah, Lingwei Zhu, Adam White, Martha White
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13274v1 Announce Type: cross Abstract: Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited. Acto...

📖 Read original article


82. Faithful Autoformalization of Natural Language Assertions ​

Author: Hongyi Liu, Madhusudan Parthasarathy, Adithya Murali
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.13303v1 Announce Type: cross Abstract: Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language specificati...

📖 Read original article


83. Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks ​

Author: Jae Joong Lee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2607.13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency G...

📖 Read original article


84. Tabular Foundation Models for Discrete Choice Estimation ​

Author: Liu Liu, Dan Zhang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, econ.EM

arXiv:2607.13314v1 Announce Type: cross Abstract: Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFMs can be effectively applied to discrete choice, a central demand estimation framework in marketing ...

📖 Read original article


85. Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains ​

Author: Rwik Rana, Jesse Quattrociocchi, Christian Ellis, Nathan Tsoi, Garrett Warnell, Joydeep Biswas
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.13319v1 Announce Type: cross Abstract: High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains. Recent forward kinodynamic (FKD) prediction foundation models suggest a promising path, starting from a generalist...

📖 Read original article


86. Privacy Preserving Recommender Systems Balancing Personalization with Privacy ​

Author: Ranjeet K Jha, Venkata Suresh Gummadilli
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.13328v1 Announce Type: cross Abstract: Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed user interaction data, creating significant privacy and regulatory challenges. With increasing requi...

📖 Read original article


87. Efficient Text-to-Audio Generation via Pruning ​

Author: Arshdeep Singh, Yi Yuan, Yun Chen, Wenwu Wang, Mark D. Plumbley
Published: 7/16/2026, 4:00:00 AM
Categories: eess.AS, cs.AI

arXiv:2607.13330v1 Announce Type: cross Abstract: Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their practical deployment is hindered by the substantial computational cost of the U-Net denoising backbone. I...

📖 Read original article


88. The Refusal Residue: When Probes Catch Alignment Faking and When They Don't ​

Author: Aman Mehta
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2607.13346v1 Announce Type: cross Abstract: Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When no scratchpad is visible, behavior alone cannot distinguish strategic from genuine compliance. We ask...

📖 Read original article


89. Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition ​

Author: Donghwan Kim
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.13347v1 Announce Type: cross Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, us...

📖 Read original article


90. Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System ​

Author: Teri Rumble, Javad Zarrin, P. George Lovell, Ruth Falconer
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2607.13370v1 Announce Type: cross Abstract: This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledg...

📖 Read original article


91. The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle ​

Author: Augusto Camargo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.PF, cs.AI

arXiv:2607.13393v1 Announce Type: cross Abstract: A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself unable to when the incumbent's own output has quietly become the specification. This note offers that o...

📖 Read original article


92. Price of Fairness in Bandits: A Tight Minimax Characterization ​

Author: Dhruv Sarkar, Soumyadeep Dutta, Sayak Ray Chowdhury
Published: 7/16/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2607.13402v1 Announce Type: cross Abstract: In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ante losses in settings such as clinical trials. Recent work addresses this by evaluating the sequence ...

📖 Read original article


93. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models ​

Author: Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qinming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
Published: 7/16/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD

arXiv:2607.13408v1 Announce Type: cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or ...

📖 Read original article


94. Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification ​

Author: Matthew Steven P. Toledo, Justine Raphael H. Jacinto, Vivekjeet Singh Chambal, Rodolfo C. Camaclang III, Jamlech Iram N. Gojo Cruz, Reginald Neil C. Recario
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13413v1 Announce Type: cross Abstract: This study presents an empirical benchmarking comparison between Kolmogorov-Arnold Networks (KANs) and Multi-Layer Perceptrons (MLPs) on structured tabular classification tasks. Motivated by the growing interest in KANs as an alternative function-app...

📖 Read original article


95. Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents ​

Author: Jiwen Zhou, Xiang Liu, Mingming Li, Pengbo Mo, Jiao Dai, Honglei Lv, Jizhong Han, Songlin Hu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.13418v2 Announce Type: cross Abstract: Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, ...

📖 Read original article


96. ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding ​

Author: Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with prec...

📖 Read original article


97. Data-Efficient Adaptation of LLMs via Attention Head Reweighting ​

Author: Tuomas Oikarinen, Zixiao Chen, Charlotte Siska, Tsui-Wei Weng, Chandan Singh, Jianfeng Gao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.13425v1 Announce Type: cross Abstract: Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs) have demonstrated some capabilities for data-efficient learning, especially through parameter-efficient adapta...

📖 Read original article


98. Discrete Diffusion Models: A Unified Framework from Tokenization to Generation ​

Author: Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Philip S. Yu, Xue Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.13431v1 Announce Type: cross Abstract: Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, wh...

📖 Read original article


99. Learning Physics-Guided Residual Dynamics for Deformable Object Simulation ​

Author: Shivansh Patel, Kaifeng Zhang, Sanjay Pokkali, Svetlana Lazebnik, Yunzhu Li
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.13451v1 Announce Type: cross Abstract: Simulating deformable objects is essential for a wide range of robotic manipulation applications, yet accurately predicting their dynamics remains challenging. We propose Physics-Guided Residual Dynamics (PGRD), a hybrid simulation framework that com...

📖 Read original article


100. Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection ​

Author: Mingyue Zeng, De Cheng, Zhipeng Xu, Huaijie Wang, Nannan Wang, Xinbo Gao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13452v1 Announce Type: cross Abstract: Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learning perspective, separating feature spaces to sharpen decision boundari...

📖 Read original article


101. Adversarial Prompting Framework for AI Safety Assessment ​

Author: Yash Bhatnagar, Kunal Banerjee, Anirban Chatterjee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13453v1 Announce Type: cross Abstract: Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use of these models may also expose systems to new forms of cyberattacks by different malicious actors --...

📖 Read original article


102. GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding ​

Author: Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, Zhiyu Lin, Ye Yuan, Zhongjiang He, Yu Yu, Hao Sun
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13454v1 Announce Type: cross Abstract: Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack th...

📖 Read original article


103. DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments ​

Author: Huatao Li, Xinwei Geng, Yuheng Wang, Yutong Li, Runde Yang, Hantao Chen, Shu Yao, Jingru Fan, Xuhui Ren, Yuanyuan Zhao, Fei Huang, Chen Qian
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.MA, cs.SE

arXiv:2607.13465v1 Announce Type: cross Abstract: LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be proce...

📖 Read original article


104. Explainable Artificial Intelligence for Anomaly Detection in Banking Transactions: An Internal Audit Perspective ​

Author: Anupa Lodhi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13469v1 Announce Type: cross Abstract: The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-based approaches struggle with high false-positive rates and offer no justification for their outputs, limiting t...

📖 Read original article


105. DeepLoop: Depth Scaling for Looped Transformers ​

Author: Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13491v1 Announce Type: cross Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transfo...

📖 Read original article


106. ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level ​

Author: Chethan Reddy G. P
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13511v1 Announce Type: cross Abstract: We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \approx B \mathrm{diag}(D) C$ with ternary factors $B \in {-1,0,+1}^{m \times k}$, $C \in {...

📖 Read original article


107. Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning ​

Author: Shiqi Zhang, Marius Fai{\ss}, Ariana Strandburg-Peshkin, Tuomas Virtanen
Published: 7/16/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD

arXiv:2607.13555v1 Announce Type: cross Abstract: Bioacoustic call-type classification relies on costly expert annotation. Active learning can reduce this burden by selecting a small batch of segments for expert annotation and using the labeled segments for training the classifier. The setting is ha...

📖 Read original article


108. Grounded world models in biological organisms and future embodied AI ​

Author: Giovanni Pezzulo, Davide Nuzzi, Marco D'Alessandro, Riccardo Proietti, Roberto Bottini, Paul Cisek
Published: 7/16/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2607.13560v1 Announce Type: cross Abstract: Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data. However, the resulting systems remain largely based on passive training regimes where linguistic regularities create the scaffold ...

📖 Read original article


109. UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors ​

Author: Dima Galat, Marian-Andrei Rizoiu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2607.13565v1 Announce Type: cross Abstract: We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the top 5 positions on the ELOQUENT 2026 Voight-Kampff leaderboard. While adversarial fine-tuning trivially closes ...

📖 Read original article


110. Spectral-Informed Neural Networks Outperform Spectral Methods in High-dimensional PDEs ​

Author: Tianchi Yu, Ivan Oseledets
Published: 7/16/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.CE, cs.LG, cs.NA

arXiv:2607.13566v1 Announce Type: cross Abstract: For low-dimensional problems ($d\leq3$), spectral methods can achieve exceptionally high accuracy. For middle-dimensional problems ($4 \leq d \lesssim 10$), spectral methods remain feasible through specific techniques such as sparse grids or hyperbol...

📖 Read original article


111. GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning ​

Author: Kaicong Huang, Weiheng Oh, Ruimin Ke
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13569v1 Announce Type: cross Abstract: Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models (VLMs) ...

📖 Read original article


112. Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification ​

Author: Shiqi Zhang, Tuomas Virtanen
Published: 7/16/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD

arXiv:2607.13571v1 Announce Type: cross Abstract: Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selecting the audio segments whose labels help the classifier most. One of the prevailing acquisition strategies for th...

📖 Read original article


113. IMMNet: Hybrid Fusion of Model-based and Data-driven Approaches for Maneuvering Target Tracking ​

Author: Yixuan Zhao, Chaoqun Yang, Lin Gao, Yongxiao Tian, Ting Yuan
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.13573v1 Announce Type: cross Abstract: Maneuvering target tracking in three-dimensional space remains a challenging problem due to complex motion dynamics and model mismatch. To address this, this paper proposes a hybrid model/data-driven algorithm named IMMNet, which integrates the inter...

📖 Read original article


114. Agile perceptive multi-skill locomotion for quadrupedal robots in the wild ​

Author: Jun-Gill Kang, Jaehyun Park, Tae-Gyu Song, Joon-Ha Kim, Seungwoo Hong, Hae-Won Park
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.13579v1 Announce Type: cross Abstract: Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration of multiple motor skills, smooth transitions between gaits, and high-speed perceptive locomotion using only on...

📖 Read original article


115. From Prediction to Collaboration: Interactive Symbolic Music Analysis ​

Author: Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth, Gerhard Widmer
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.13587v1 Announce Type: cross Abstract: Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such as full-score prediction, and therefore do not match the broader range of operations that arise in analysis wo...

📖 Read original article


116. Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents ​

Author: Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.13591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-d...

📖 Read original article


117. Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities ​

Author: Eunna Lee, Jungpyo Nam, Sunjun Hwang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it ...

📖 Read original article


118. Semantic Anchoring for Robotic Action Representations ​

Author: Yuan Xu, Youheng Shi, Chengyang Li, Wentao Zhu, Yizhou Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.13597v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization. A fundamental question therefore ...

📖 Read original article


119. The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models ​

Author: Fabio Arnez, Alexandra Gomez-Villa
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13612v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performance rather than a normative principle. We show that the choice of anti-collapse regulariser determin...

📖 Read original article


120. From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception ​

Author: Jose Mart'inez-Fajardo, Pablo Pueyo, Fernando Caballero, Luis Merino
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.13624v1 Announce Type: cross Abstract: Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming user requests into executable navigation actions remains a challenging task, requiring the integration of languag...

📖 Read original article


121. OvisOCR2 Technical Report ​

Author: Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen, An-Yang Ji, Jun-Peng Jiang, Qing-Guo Chen, Jianshan Zhao, En Lin, Haijun Li, Cheng Qin, Zhao Xu, Weihua Luo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13639v1 Announce Type: cross Abstract: We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. We ...

📖 Read original article


122. Consensus as Privileged Context for Label-Free Self-Distillation ​

Author: John Gkountouras, Josip Juki'c, Ivan Titov
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.13643v1 Announce Type: cross Abstract: Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large language models without labels, and a growing family of methods converts this consensus signal into training supe...

📖 Read original article


123. Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction ​

Author: Tianshun Han, Ziyu Shi, Lijian Liu, Ajian Liu, Benjia Zhou, Hugo Jair Escalante, Yanyan Liang, Sergio Escalera, Zhen Lei, Jun Wan
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13646v1 Announce Type: cross Abstract: Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-world scenarios. They often produce unstable geometry, inaccurate limb articulation and unreliable predictions un...

📖 Read original article


124. Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models ​

Author: Ayan Igali, Pakizar Shamoi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13647v1 Announce Type: cross Abstract: Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces such as CIELAB or with discrete color labels. These references capture perceptual distance or category membership,...

📖 Read original article


125. Barnamala: Parameter-Efficient Handwritten Devanagari Recognition at Benchmark Saturation ​

Author: Ashish Thapa, Samrat Karki
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13689v1 Announce Type: cross Abstract: We built a compact convolutional network (1.11 M parameters) for 46-class DHCD Devanagari recognition and reached 99.73%, the highest reported at 15.6x smaller than prior state-of-the-art. We have effectively reached the saturation point: every model...

📖 Read original article


126. Social Simulations: from Agent-Based Modeling to Digital Twins ​

Author: Erica Cau, Andrea Failla, Valentina Pansanella, Giulio Rossetti
Published: 7/16/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.SI, physics.soc-ph

arXiv:2607.13693v1 Announce Type: cross Abstract: This book chapter covers the evolution of social simulation from classical agent-based models, in which agents interact according to explicitly defined behavioral rules, to AI-enhanced simulations based on Large Language Models and, ultimately, Socia...

📖 Read original article


127. Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs ​

Author: Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Chunxiao Liu, Dongming Zhang, Zhendong Mao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.MM

arXiv:2607.13712v1 Announce Type: cross Abstract: Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and p...

📖 Read original article


128. How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement ​

Author: Alexandra E. Michael, Franziska Roesner
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.13718v1 Announce Type: cross Abstract: As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as hallucination, can cause agents to leak private information to third parties. As autonomous systems, agents also prese...

📖 Read original article


129. Anatomically Faithful but Temporally Blind: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography ​

Author: Hyunkyung Han, Min Jung Kim
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13738v1 Announce Type: cross Abstract: Background and Objective: Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attribution (Chefer relevance for transformers, Grad-CAM for CNNs) is increasingly used to cert...

📖 Read original article


130. MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model ​

Author: Charilaos Papaioannou, Ioannis Tsantilas, Dimitris Giannakakos, Vasilis Michalakopoulos, Sotiris Pelekis, Vangelis Marinakis, Arsam Aryandoust, Antonello Monti, Ricardo J. Bessa, Perdo P. Vergara, Jochen Cremer, Elissaios Sarmas
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13763v1 Announce Type: cross Abstract: Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting: the ten...

📖 Read original article


131. Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations ​

Author: Wenxuan Miao, Haosong Liu, Weiming Hu, Zihan Liu, Aiyue Chen, Jianlin Yu, Yiwu Yao, Yiming Gan, Jieru Zhao, Jingwen Leng, Minyi Guo, Yu Feng
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.13770v1 Announce Type: cross Abstract: Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timesteps and self attention computation. As diffusion timesteps are reduced, the computation cost of self attention ...

📖 Read original article


132. CAS I: A Geometric Coding Theorem ​

Author: Romie Banerjee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.CT, math.GR, math.IT

arXiv:2607.13796v1 Announce Type: cross Abstract: This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijections on the set of binary strings, called symmetries and define the symmetry prior of a string as the probability...

📖 Read original article


133. Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection ​

Author: Zhenpeng Li
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.13801v1 Announce Type: cross Abstract: Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness against feasible traffic manipulation remains largely empirical. We present Traffic-Aware Randomized Smoothing (...

📖 Read original article


134. Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning ​

Author: Vincent Ochs, Christoph Kuemmerli, Florentin Bieder, Julia Wolleb, Joel L. Lavanchy, Julia Ruppel, Jan Liechti, Stephanie Taha-Mehlitz, Christian Andreas Nebiker, Beat Mueller, Giuseppe Kito Fusai, Joerg-Matthias Pollok, Anas Taha, Philippe C. Cattin, Sebastian Staubli
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR, cs.LG

arXiv:2607.13826v1 Announce Type: cross Abstract: Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability. We introduce a fu...

📖 Read original article


135. NodeImport: Imbalanced Node Classification with Node Importance Assessment ​

Author: Nan Chen, Zemin Liu, Bryan Hooi, Bingsheng He, Jun Hu, Jia Chen
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13837v1 Announce Type: cross Abstract: In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend t...

📖 Read original article


136. AI-Augmented Human Resource Management? Insights from German companies ​

Author: Yannick Kalff, Katharina Simbeck
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13839v1 Announce Type: cross Abstract: This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \enquote{augmenting} human resource management. Organisations employ generative AI or predictive analytics to tr...

📖 Read original article


137. Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild ​

Author: Ting Lei, Jialin Liu, Zhu Xu, Yuxin Peng, Yang Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13881v1 Announce Type: cross Abstract: Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangl...

📖 Read original article


138. Verifying formulas for interventional distributions ​

Author: Francesco Freni, Leonard Henckel, Sebastian Weichwald
Published: 7/16/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG, stat.ML

arXiv:2607.13883v1 Announce Type: cross Abstract: We formalize verification in causal graphical models: deciding whether a given observational formula identifies a target interventional distribution. This opens a problem complementary to identification, asking not whether any identifying formula exi...

📖 Read original article


139. Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings ​

Author: Jiangang Han
Published: 7/16/2026, 4:00:00 AM
Categories: math.ST, cs.AI, cs.LG, stat.TH

arXiv:2607.13918v1 Announce Type: cross Abstract: Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $k$ verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.15712) shows that posterior l...

📖 Read original article


140. Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code ​

Author: Niels M"undler-Sasahara, Hristo Venev, Dawn Song, Martin Vechev, Jingxuan He
Published: 7/16/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.LG

arXiv:2607.13921v1 Announce Type: cross Abstract: Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not guide inter...

📖 Read original article


141. Music-to-Dance Generation via Atomic Movements ​

Author: Xinhao Cai, Yixuan Sun, Minghang Zheng, Qingchao Chen, Xin Jin, Song-chun Zhu, Yang Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13978v1 Announce Type: cross Abstract: Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continu...

📖 Read original article


142. The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce ​

Author: Sai Srikanth Madugula, Peplluis Esteva de la Rosa, Daya Shankar
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SI, cs.AI, cs.GT, cs.MA

arXiv:2607.13998v1 Announce Type: cross Abstract: The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive recommendation algorithms to autonomous, goal-directed agents capable of executing purchasing decisio...

📖 Read original article


143. Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation ​

Author: Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.14006v1 Announce Type: cross Abstract: Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational controls to achieve security-relevant compromise. This paradigm remains necessary for AI-enabled system...

📖 Read original article


144. Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth ​

Author: Katie Everett
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, ...

📖 Read original article


145. Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study ​

Author: Daniel Grillmeyer, Marius Hadry, Michael Stenger, Vanessa Borst, Veronika Lesch, Samuel Kounev
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.14024v1 Announce Type: cross Abstract: With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global energy mix continues to grow. Unlike conventional power generation, the output of renewable energy sources cannot ...

📖 Read original article


146. Early Adoption of Agentic Coding Tools by GitHub Projects ​

Author: Maliha Noushin Raida, Daqing Hou
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY, cs.LG

arXiv:2607.14037v2 Announce Type: cross Abstract: Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of ag...

📖 Read original article


147. Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study ​

Author: Zhan Chen, Jiqiao Ma, Chih-wen Kuo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.14041v1 Announce Type: cross Abstract: Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cursive chancery hand used in palace memorials, despite limited labeled data. We study a multi-expert system that ...

📖 Read original article


148. A Survey on Hypergame Theory: Modelling Misaligned Perceptions and Nested Beliefs for Multi-Agent Systems ​

Author: Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2507.19593v3 Announce Type: replace Abstract: Classical game-theoretic models typically assume rational agents, complete information, and common knowledge of payoffs - assumptions that are often violated in real-world MAS characterized by uncertainty, misaligned perceptions, and nested beliefs...

📖 Read original article


149. Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate ​

Author: Pratik S. Sachdeva, Tom van Nuenen
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.10002v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in sensitive everyday contexts -- offering personal advice, mental health support, and moral guidance -- understanding their behavior in navigating complex moral reasoning is essential. Most...

📖 Read original article


150. Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution ​

Author: Zhengbo Jiao, Hongyu Xian, Qinglong Wang, Yunpu Ma, Zhebo Wang, Zifan Zhang, Dezhang Kong, Meng Han
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test-time scaling methods treat execution feedback merely as an external signal for filtering or rewriti...

📖 Read original article


151. When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents ​

Author: Aman Mehta
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training-free, black-box uncertainty signal that instantiates selective classification and distribution-free...

📖 Read original article


152. Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation ​

Author: Zeyu Chen, Huanjin Yao, Ziwang Zhao, Min Yang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2603.00546v2 Announce Type: replace Abstract: Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains. Evaluating the capability and reliability of MLLM-as-a-judge systems is therefo...

📖 Read original article


153. NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing ​

Author: Ming Yang, Zhi Zhou, Shi-Yu Tian, Kun-Yang Yu, Lan-Zhe Guo, Yu-Feng Li
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.16307v3 Announce Type: replace Abstract: Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex scenes and constraints and make reliable decisions. Current remote-sensing benchmarks mainly focus on evaluati...

📖 Read original article


154. How LLMs Might Think ​

Author: Joseph Gottlieb, Ethan Kemp, Matthew Trager
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.09674v2 Announce Type: replace Abstract: Do large language models (LLMs) think? Daniel Stoljar and Zhihe Vincent Zhang have recently developed an argument from rationality for the claim that LLMs do not think. We contend, however, that the argument from rationality not only falters, but l...

📖 Read original article


155. Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation ​

Author: Sum Kyun Song, Bong Gyun Shin, Jae Yong Lee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.NE, cs.SC

arXiv:2605.07323v2 Announce Type: replace Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regression approaches rely primarily on quantitative metrics; however, real-world differential equatio...

📖 Read original article


156. From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents ​

Author: Patrick Wilhelm, Odej Kao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.06223v2 Announce Type: replace Abstract: Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting i...

📖 Read original article


157. A Causal Model of Theory of Mind in Conflict for Artificial Intelligence ​

Author: Nikolos Gurney
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2606.16944v2 Announce Type: replace Abstract: Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely assumed to be essential for effective human-machine integration. Existing AI-ToM models address \emph{how} to me...

📖 Read original article


158. OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3 ​

Author: David Courtis, Wenhao Li, Scott Sanner
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.01531v2 Announce Type: replace Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but data-hungry and transfer poorly beyond their training distribution. Program-...

📖 Read original article


159. Infinity-Parser2 Technical Report ​

Author: Zuming Huang, Jun Huang, Kexuan Ren, Baode Wang, Weizhen Li, Jianming Feng, Yu Wang, Yichen Yao, Shijun Lin, Yige Tang, Cheng Peng, Weidi Xu, Wei Chu, Yinghui Xu, Yuan Qi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.07836v3 Announce Type: replace Abstract: We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing cor...

📖 Read original article


160. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation ​

Author: Runhan Shi, Quan Zhou, Yuqian Xu, Shuai Yang, Xin Wu, Zitong Zhou, Hui Liu, Bin Zha, Zheming Wang, Liya Li, Wei Wei, Jinru Ding, Wenrao Pang, Mouxiao Bian, Haoyuan Hu, Jun Xu, Jie Xu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2607.09142v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medi...

📖 Read original article


161. Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes ​

Author: Yossi Eliaz
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, math.PR, math.ST, stat.TH

arXiv:2607.09689v2 Announce Type: replace Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs--Boltzmann measure $\exp{-\beta E(\theta)}$ whose inverse temperature is the sample size, $\beta=n$. Three conseque...

📖 Read original article


162. ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ​

Author: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Zixiao Tang, Zhining Gu, Yang Cai, Linbo Zheng, Jingjing Ma, Mingyang Yin, Zedong Chu, Ziqiao Li, Mu Xu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.10350v2 Announce Type: replace Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot...

📖 Read original article


163. Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans ​

Author: Jasmin Thelen, Oliver Wilhelm
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.11263v3 Announce Type: replace Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measureme...

📖 Read original article


164. Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation ​

Author: Congren Dai, Danni Zhao, Enyang Liu, Michael Ching Yam, Zhancheng Guo, Siyi Gu, Wentao Yang, Bo Dai, Xiaobing Li, Maosong Sun
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.11334v2 Announce Type: replace Abstract: Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verific...

📖 Read original article


165. Optimization Is Not All You Need ​

Author: Minh Hua, Rita Raley
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.11977v2 Announce Type: replace Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it ins...

📖 Read original article


166. From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery ​

Author: Ingmar Posner, Anson Lei, Bernhard Sch"olkopf
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12474v2 Announce Type: replace Abstract: Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from protein folding to weather forecasting. Yet prediction alone does not constitute scientific discov...

📖 Read original article


167. Evidence-Grounded AI for Musculoskeletal Care ​

Author: Wenjie Li, Yujie Zhang, Fanrui Zhang, Haoran Sun, Renhao Yang, Junjun He, Weiran Huang, Yuanfeng Ji, Chenrun Wang, Kailing Wang, Hongcheng Gao, Kaipeng Zhang, Hanyu Wang, Angela Lin Wang, Xingqi He, Yilin Huang, Shiyi Yao, Lilong Wang, Yankai Jiang, Yirong Chen, Chenglong Ma, Jiyao Liu, Ming Hu, Gen Li, Yidong Xu, Chengyu Zhuang, Jiawei Liu, Yin Zhang, Lequan Yu, Lu Chen, Yinpeng Dong, Lei Liu, Carlos Gutierrez Sanroman, Yu Qiao, Weijie Ma, Xiaosong Wang, Lei Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12527v2 Announce Type: replace Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery, remodelling and degeneration often unfold over months to years, musculoskeletal care requires lo...

📖 Read original article


168. Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes ​

Author: Jonas Ehrhardt, Ren'e Heesch, Oliver Niggemann
Published: 7/16/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12924v2 Announce Type: replace Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determ...

📖 Read original article


169. Koopman-driven grip force prediction through EMG sensing ​

Author: Tomislav Bazina, Ervin Kamenar, Maria Fonoberova, Igor Mezi'c
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, math.DS

arXiv:2409.17340v2 Announce Type: replace-cross Abstract: Loss of hand function due to conditions like stroke or multiple sclerosis significantly impacts daily activities. Robotic rehabilitation provides tools to restore hand function, while novel methods based on surface electromyography (sEMG) ena...

📖 Read original article


170. PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors ​

Author: Xinwei Liu, Xiaojun Jia, Yuan Xun, Hua Zhang, Xiaochun Cao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2502.16167v2 Announce Type: replace-cross Abstract: Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyright concerns. Malicious actors can misuse these models to generate unauthorized portraits or artistic s...

📖 Read original article


171. NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache ​

Author: Donghyun Son, Euntae Choi, Sungjoo Yoo
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache. Vector Quantization (VQ) is recently adopted to alleviate this...

📖 Read original article


172. Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials ​

Author: Kingsley Yeon, Steven B. Damelin
Published: 7/16/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.LG, cs.NA, stat.ML

arXiv:2506.21306v2 Announce Type: replace-cross Abstract: Functions that grow without bound on one side of the real line and decay to zero on the other cannot be approximated uniformly by ordinary polynomials on unbounded domains. Motivated by classical weighted polynomial approximation, we introduc...

📖 Read original article


173. Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery ​

Author: Yi-Shan Chu, Hsuan-Cheng Wei
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2507.16849v3 Announce Type: replace-cross Abstract: We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery, aiming to support and enhance the Emergent Value Added Product (EVAP) developed by the Taiwan Space...

📖 Read original article


174. Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping ​

Author: Xuhui Zhan, Tyler Derr
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2508.12466v2 Announce Type: replace-cross Abstract: Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting visual features into discrete text token spaces using large-scale image--text data. We revisit this de...

📖 Read original article


175. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models ​

Author: Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.22415v4 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token hidden state into t...

📖 Read original article


176. Representation-Based Exploration for Language Models: From Test-Time to Post-Training ​

Author: Jens Tuyls, Dylan J. Foster, Akshay Krishnamurthy, Jordan T. Ash
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the discovery of novel behaviors, or simply sharpen those already present in the base model. In this paper,...

📖 Read original article


177. Benefits and Limitations of Communication in Multi-Agent Reasoning ​

Author: Michael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi, Guillaume Rabusseau, Michael Hahn
Published: 7/16/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG

arXiv:2510.13903v2 Announce Type: replace-cross Abstract: Chain-of-thought prompting has popularized step-by-step reasoning in large language models, yet model performance still degrades as problem complexity and context length grow. By decomposing difficult tasks with long contexts into shorter, ma...

📖 Read original article


178. Column Generation with Domain-Independent Dynamic Programming ​

Author: Ryo Kuroiwa, Edward Lam
Published: 7/16/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2510.14317v2 Announce Type: replace-cross Abstract: Column generation and branch-and-price (B&P) are leading mathematical optimization methods for large-scale exact optimization, iterating between solving a master problem and a pricing problem. Due to the difficulty of discrete optimization, h...

📖 Read original article


179. Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals ​

Author: Shuntaro Suzuki, Shunya Nagashima, Komei Sugiura
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2510.15371v2 Announce Type: replace-cross Abstract: Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including communication assistance and rehabilitation support for patients with motor impairments. These signals re...

📖 Read original article


180. MASPRM: Multi-Agent System Process Reward Model ​

Author: Milad Yazdani, Mahdi Mostajabdaveh, Zirui Zhou, Ying Xiong
Published: 7/16/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2510.24803v3 Announce Type: replace-cross Abstract: Inference-time search over multi-agent systems (MAS) wastes compute when it cannot identify which agent's intermediate message advanced progress. We present the Multi-Agent System Process Reward Model (MASPRM), which scores routed transcripts...

📖 Read original article


181. Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs ​

Author: Amirali Ebrahimzadeh, Seyyed M. Salili
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctuates significantly depending on how information is distributed in real-world corpora. We investigate ...

📖 Read original article


182. Mind the Gap: Action Rebinding Attacks against Android GUI Agents ​

Author: Yi Qian, Kunwei Qian, Xingbang He, Ligeng Chen, Jikang Zhang, Tiantai Zhang, Haiyang Wei, Linzhang Wang, Hao Wu, Bing Mao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE

arXiv:2601.12349v3 Announce Type: replace-cross Abstract: Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen content and inject inputs across application boundaries. While these agents aim to automate complex tasks, we...

📖 Read original article


183. ELF: A Family of Encoder-Free ECG-Language Models ​

Author: William Han, Tony Chen, Chaojing Duan, Xiaoyu Song, Yihang Yao, Yuzhe Yang, Michael A. Rosenberg, Emerson Liu, Ding Zhao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.MM, cs.AI

arXiv:2601.18798v3 Announce Type: replace-cross Abstract: ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, most existing ELMs inherit Vision-Language Model (VLM) design choices and rely on pretrained ECG encoders,...

📖 Read original article


184. With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots ​

Author: Zeinab Sadat Taghavi, Ali Modarressi, Hinrich Schutze, Andreas Marfurt
Published: 7/16/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2602.09616v2 Announce Type: replace-cross Abstract: Reliable retrieval-augmented generation (RAG) systems depend fundamentally on the retriever's ability to find relevant information. We show that neural retrievers used in RAG systems have blind spots, which we define as the failure to retriev...

📖 Read original article


185. Left-right asymmetry in predicting brain activity from LLMs' representations emerges with their formal linguistic competence ​

Author: Laurent Bonnasse-Gahot, Christophe Pallier
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, q-bio.NC

arXiv:2602.12811v2 Announce Type: replace-cross Abstract: When humans and large language models (LLMs) process the same text, activations in the LLMs correlate with brain activity measured, e.g., with functional magnetic resonance imaging (fMRI). Moreover, it has been shown that, as the training of ...

📖 Read original article


186. 1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World ​

Author: Qiao Xu, Yipeng Yu, Chengxiao Feng, Xu Liu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2602.18548v3 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inconsistent datasets, toolchains, and evaluation protocols. We introduce 1D-Bench, a benchmark grounded...

📖 Read original article


187. A novel network for classification of cuneiform tablet metadata ​

Author: Frederik Hagelskj{\ae}r
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.03892v2 Announce Type: replace-cross Abstract: In this paper, we present a network structure for classifying metadata of cuneiform tablets. The problem is of practical importance, as the size of the existing corpus far exceeds the number of experts available to analyze it. But the task is...

📖 Read original article


188. When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech ​

Author: Akif Islam, Raufun Nahar, Md. Ekramul Hamid
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG

arXiv:2603.04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription. In this work, we examine whether this assumption holds for modern ...

📖 Read original article


189. PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner ​

Author: Eugene Ku, Yiwei Lyu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2603.10330v2 Announce Type: replace-cross Abstract: Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that learn behavior from expert demonstrations. Diffusion-based trajectory planners have recently shown stron...

📖 Read original article


190. RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset ​

Author: Yongzhong Wang, Keyu Zhu, Yong Zhong, Liqiong Wang, Jinyu Yang, Feng Zheng
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2603.11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms. To break this barr...

📖 Read original article


191. LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement ​

Author: Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang, Shao-Yi Chien, Yu Tsao, Fan-Gang Zeng
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2603.13952v3 Announce Type: replace-cross Abstract: In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often subo...

📖 Read original article


192. Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion ​

Author: Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, SoonYoung Lee, Wonbin Ahn
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.22372v3 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) forecasting. However, most existing methods provide limited gains, often improving performance only in ...

📖 Read original article


193. Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts ​

Author: Ali Safari, Sahar Babaei
Published: 7/16/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.IR

arXiv:2603.22510v3 Announce Type: replace-cross Abstract: Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by changes in research novelty. We examine how relative abstract-level semantic novelty in Information Syst...

📖 Read original article


194. Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies ​

Author: Zhanzhi Lou, Hui Chen, Yibo Li, Qian Wang, Bryan Hooi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.00830v3 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is an adaptation policy that updates the actor policy based on e...

📖 Read original article


195. Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems ​

Author: Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram, Ritwik Garg, Sagar Jha, Nimet Beyza Bozdag, Dilek Hakkani-Tur
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative mu...

📖 Read original article


196. Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces ​

Author: Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.11996v4 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a fundamental limitati...

📖 Read original article


197. Robust Explanations for User Trust in Enterprise NLP Systems ​

Author: Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu, Amine Anoun, Jerry Ting
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.12069v3 Announce Type: replace-cross Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and e...

📖 Read original article


198. Partially Observed Structural Causal Models ​

Author: Turan Orujlu, Jordan Matelsky, Martin V. Butz, Charley M. Wu, Konrad P. Kording
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ME, stat.ML

arXiv:2605.03268v2 Announce Type: replace-cross Abstract: Here we introduce Partially Observed Structural Causal Models (POSCMs) as an extension of structural causal models (SCMs) to settings where upstream contexts co-determine both the interaction structure and downstream mechanisms on observed va...

📖 Read original article


199. Stable Attention Response for Reliable Precipitation Nowcasting ​

Author: Penghui Wen, Zexin Hu, Sen Zhang, Patrick Filippi, Xiaogang Zhu, Allen Benter, Thomas Bishop, Zhiyong Wang, Kun Hu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.13181v2 Announce Type: replace-cross Abstract: Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although recent methods increasingly adopt attention-based architectures in both unimodal and multim...

📖 Read original article


200. TuxBot: Semantic-Aware Online OS Tuning with Large Language Models ​

Author: Georgios Liargkovas, Mihir Nitin Joshi, Hubertus Franke, Kostis Kaffes
Published: 7/16/2026, 4:00:00 AM
Categories: cs.OS, cs.AI, cs.PF

arXiv:2605.15026v2 Announce Type: replace-cross Abstract: Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power, memory, and I/O controls as black-box variables and optimize a scalar reward. This view ignores cross-...

📖 Read original article


201. Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents ​

Author: Ummara Mumtaz, Summaya Mumtaz
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2605.16281v2 Announce Type: replace-cross Abstract: Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident reporting, and impact assessment obligations are visible when AI systems fail. This study analyzes real-world...

📖 Read original article


202. Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs ​

Author: Yigui Feng (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Qinglin Wang (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China), Yang Liu (The Shien-Ming Wu School of Intelligent Engineering, South China University of Technology, Guangzhou, Guangdong, China), Jie Liu (The College of Computer Science, National University of Defense Technology, Changsha, Hunan, China)
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.16366v2 Announce Type: replace-cross Abstract: Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many spatial tokens, while capturing short-lived events requires dense temporal sampling. We propose \textbf...

📖 Read original article


203. DIVE: Embedding Compression via Self-Limiting Gradient Updates ​

Author: Dongfang Zhao
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG

arXiv:2605.20689v2 Announce Type: replace-cross Abstract: High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce. We present DIVE (Dimensionality reduction with Implicit View Ensembles), a residual compr...

📖 Read original article


204. The TIME Machine: On The Power of Motion for Efficient Perception ​

Author: Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2605.23045v3 Announce Type: replace-cross Abstract: Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have p...

📖 Read original article


205. Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory ​

Author: Quanjiang Li, Zhiming Liu, Wei Luo, Tingjin Luo, Chenping Hou
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.24602v5 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated ...

📖 Read original article


206. When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation ​

Author: Faizan Faisal
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2605.24902v2 Announce Type: replace-cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical documentation; we investigate this question using SOAP note generation from clinical dialogue ...

📖 Read original article


207. A Multi-Model Metric-based Selection Framework for Abstractive Text summarization ​

Author: Ahmed Alansary, Ali Hamdi
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.05494v4 Announce Type: replace-cross Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model Summarization Framework designed to improve the robustness and quality of abstractive tex...

📖 Read original article


208. Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents ​

Author: Ankita Samaddar, Sandeep Neema, Daniel Balasubramanian, Xenofon Koutsoukos
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2606.18223v2 Announce Type: replace-cross Abstract: With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained via Reinforcement Learning (RL). These agents employ neurosymbolic approaches such as behavior trees...

📖 Read original article


209. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling ​

Author: Yixuan Lai, Tianjia Shao, Kun Zhou, Weijia Dou, Siyu Zhu, Jingdong Wang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.20799v2 Announce Type: replace-cross Abstract: Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, causing entities that reappear across shots -- characters, objects, and locations -- to drift ...

📖 Read original article


210. MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation ​

Author: Teerath Kumar, Raja Vavekanand, Muhammad Turab
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.28419v2 Announce Type: replace-cross Abstract: Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventional augmentation can improve training diversity but may distort diagnostically informative structures,...

📖 Read original article


211. The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis ​

Author: Hari Prasad, Ritam Pal
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.29581v2 Announce Type: replace-cross Abstract: Modern LLM deployments often combine quantization with higher sampling temperatures to reduce cost, latency, or repetition, yet safety evaluations usually treat these as fixed implementation details. We test whether models that are safe at FP...

📖 Read original article


212. ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ​

Author: Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.31650v3 Announce Type: replace-cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion,...

📖 Read original article


213. Post-Training Pruning for Diffusion Transformers ​

Author: Chengzhi Hu, Xuewen Liu, Jing Zhang, Mengjuan Chen, Zhikai Li, Qingyi Gu
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.00927v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' u...

📖 Read original article


214. Token Geometry ​

Author: Kathan Shah
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.01455v3 Announce Type: replace-cross Abstract: Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface between them. We show that this interface has gradient geometry distinct from dense hidden weights which ...

📖 Read original article


215. Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity ​

Author: Vincent Granville
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.04166v3 Announce Type: replace-cross Abstract: I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this day. The problem statement is rudimentary: kids can understand it. However, despite decades of researc...

📖 Read original article


216. Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models ​

Author: Donna Vakalis
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.04464v2 Announce Type: replace-cross Abstract: World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, which can leave planning-relevant errors in the model's latent rollouts unmeasured. We introduce a compleme...

📖 Read original article


217. REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing ​

Author: Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu, Hung-yi Lee
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD

arXiv:2607.05364v2 Announce Type: replace-cross Abstract: Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-spee...

📖 Read original article


218. AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning ​

Author: Kyuan Oh, Bumsoo Kim
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.07033v3 Announce Type: replace-cross Abstract: Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and toke...

📖 Read original article


219. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology ​

Author: Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
Published: 7/16/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.LG

arXiv:2607.08803v2 Announce Type: replace-cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repo...

📖 Read original article


220. HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation ​

Author: Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Haoran Zhang, Fuxian Huang, Zhanhui Lin, Zijun Xu, Weinan Zhang
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09776v2 Announce Type: replace-cross Abstract: When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterati...

📖 Read original article


221. AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP ​

Author: Aritra Mazumder, Nusrat jahan Lia
Published: 7/16/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.11098v3 Announce Type: replace-cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, a...

📖 Read original article


222. An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory ​

Author: Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare
Published: 7/16/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CR, cs.HC

arXiv:2607.11707v2 Announce Type: replace-cross Abstract: Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive inform...

📖 Read original article


223. Introducing Human-Centeredness in AI-Assisted Lexicography ​

Author: Antonio San Martin, Catherine Trekker
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11808v2 Announce Type: replace-cross Abstract: This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role ...

📖 Read original article


224. Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite ​

Author: Yushi Hirose, Hiroo Irobe, Takafumi Kanamori
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11947v2 Announce Type: replace-cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted t...

📖 Read original article


225. Removable Defects: The Economics and Limits of Deliberate Deficiency ​

Author: Cheng Qian
Published: 7/16/2026, 4:00:00 AM
Categories: econ.EM, cs.AI, cs.LG, stat.ML

arXiv:2607.11983v2 Announce Type: replace-cross Abstract: A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a deficiency can be kept because it pays and removed on demand in the rare situation where it ...

📖 Read original article


226. Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs ​

Author: Nikita Kozodoi, Zainab Afolabi, Jack Butler
Published: 7/16/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.11997v2 Announce Type: replace-cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematical...

📖 Read original article


227. Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel ​

Author: Niccol`o Caselli, Francesco Massafra, Samuele Punzo, Salvatore Lo Sardo, Ippokratis Pantelidis, Sathya Kamesh Bhethanabhotla
Published: 7/16/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.12547v2 Announce Type: replace-cross Abstract: We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an extension that freezes the pretrained low-level LeWM and adds high-level planning over latent subgoals. We e...

📖 Read original article


228. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation ​

Author: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
Published: 7/16/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.12752v2 Announce Type: replace-cross Abstract: While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicat...

📖 Read original article