Skip to content

arXiv cs.AI - 2026-07-31 ​

282 items collected.


1. Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models ​

Author: Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, Aishwarya Balwani
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.26119v1 Announce Type: new Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We t...

📖 Read original article


2. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems ​

Author: Marylou Fauchard, Florian Carichon, Margarida Carvalho, Golnoosh Farnadi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misal...

📖 Read original article


3. ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science ​

Author: Yuan Zhu, Ethan B. Liu, Frank Nie, Jindong Han
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26155v1 Announce Type: new Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLI...

📖 Read original article


4. When benchmark inferences do not compose: Projectibility in AI evaluation ​

Author: Brett Reynolds
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2607.26159v1 Announce Type: new Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumpt...

📖 Read original article


5. GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning ​

Author: Lang Cao, Yuhao Shen, Tianyang Luo, Simo Du, Hao Peng, Yue Guo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-sp...

📖 Read original article


6. GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure ​

Author: Xin Xin, Jincheng Lou, Junhui Li, Jinglin Yan, Panda Xiao, Di Wu, Haixiao Li, Weicong Lu, Weijian Fan, Xinyu Qu, Yuxiang Zhao, Min Yu, Zhixiong Di, Yibo Lin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly respin. Recent large language models (LLMs) offer new opportunities to automate this process, y...

📖 Read original article


7. Position: Evaluation Scores Are Perishable Knowledge Claims ​

Author: Sankalp Gilda, Shlok Gilda
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SE

arXiv:2607.26191v1 Announce Type: new Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence...

📖 Read original article


8. TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning ​

Author: Rwaida Alssadi, Muntaser Syed, Balaji Kasula, Lamine Deen, Majed Alotaibi, Mohammed Alghamdi, Tyler Ton, Ali Alqarni, Marius Silaghi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation ...

📖 Read original article


9. Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings? ​

Author: Wanyu Zhao, Wanbing Zhao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to ...

📖 Read original article


10. CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games ​

Author: Zheng Zhang, Nanjie Yao, Jiarui He, Deheng Ye, Peilin Zhao, Hao Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26393v1 Announce Type: new Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven ...

📖 Read original article


11. CG-World: A Large-Scale World-State Dataset and Protocol for World Models ​

Author: Yiming Cai, Fangjie Yu, Meiqing Yu, Ziyue Shi, Pengfei Yuan, Yong Guo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.GR

arXiv:2607.26452v1 Announce Type: new Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and pro...

📖 Read original article


12. MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning ​

Author: Kawai Chung, Chunkit Chan, Yauwai Yim, Yuxuan Liu, Haochen Shi, Weiqi Wang, Qing Zong, Tianshi Zheng, Yixuan Fu, Kai Chung Wong, Hao Liang, Yifan Gao, Xi Yang, Janet Hui-wen Hsiao, Yangqiu Song
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly exam...

📖 Read original article


13. EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks ​

Author: Peng Yin, Kai Li, Yifan Zhang, Jian Cheng
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.NE

arXiv:2607.26490v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance heavily relies on the manual, trial-and-error engineering of neural representations, loss formulations...

📖 Read original article


14. Evidence-Ledger Adjudication for Claim-Evidence Traceability ​

Author: Gengyu Chen, Yongjie Yu, Weiling Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a suppo...

📖 Read original article


15. Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models ​

Author: Shaopeng Wei, Yufei Cheng, Wenxi Sun, Yepeng Ding, Yu Zhao, Gang Kou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26588v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterf...

📖 Read original article


16. Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants ​

Author: Zijian Xu, Wenshuo Zhang, Zisen Qin, Rui Sheng, Yushi Sun, Huamin Qu, Chuhan Shi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each a...

📖 Read original article


17. AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining ​

Author: Jingyang Yi, Jian Yang, Yifei Jin, Yuqi Li, Jian Li
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26642v1 Announce Type: new Abstract: Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, witho...

📖 Read original article


18. Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting ​

Author: Hongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.26643v1 Announce Type: new Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way a...

📖 Read original article


19. AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution ​

Author: Junhao Qiu, Zidong Wang, Yansong Sun, Zhitong Ma, Ping Guo, Qingfu Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26661v1 Announce Type: new Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally differen...

📖 Read original article


20. UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks ​

Author: Zhilun Zhou, Jianghao Yu, Yuming Lin, yongjun yang, Sun Yongquan, Depeng Jin, Yong Li
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that require discovering a...

📖 Read original article


21. Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM ​

Author: Huixiang Zhang, Mahzabeen Emu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information....

📖 Read original article


22. Property-driven Causal Abstractions for Markov Decision Processes ​

Author: Jule Schmidt, Maximilian Weininger, Clemens Dubslaff, David Parker, Nils Jansen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.26787v1 Announce Type: new Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs ...

📖 Read original article


23. From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence ​

Author: Jia Luo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.26903v1 Announce Type: new Abstract: The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduc...

📖 Read original article


24. What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation ​

Author: Vishisht Choudhary, Lukas Schmidt, Anne Zo"e Kenntner, Feras Skhab, Michel Osswald, Jens Ernstberger
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.26935v1 Announce Type: new Abstract: Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We presen...

📖 Read original article


25. Belief-Guided Decision Making with Uncertainty Gating in the Game of Go ​

Author: Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand, Ashkan Rezaei
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26946v1 Announce Type: new Abstract: Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical...

📖 Read original article


26. Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data ​

Author: Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal...

📖 Read original article


27. On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment ​

Author: Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.LG

arXiv:2607.27081v1 Announce Type: new Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills w...

📖 Read original article


28. AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching ​

Author: Yiping Song, Jiaoyan Chen, Renate Schmidt, Hui Yang, Wen Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.27130v1 Announce Type: new Abstract: Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption m...

📖 Read original article


29. Linguistic Monoculture in LLM-Assisted Language Use ​

Author: Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT

arXiv:2607.27134v1 Announce Type: new Abstract: Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance...

📖 Read original article


30. OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding ​

Author: Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC

arXiv:2607.27155v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce ...

📖 Read original article


31. Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork ​

Author: Peter Tisnikar, Maja Swieczkowska, Benteng Ma, Gerard Canal, Matteo Leonetti
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.MA

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their abili...

📖 Read original article


32. Can AI agents conduct open-ended AI research? Early evidence from two case studies ​

Author: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2607.27191v1 Announce Type: new Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended res...

📖 Read original article


33. A Methodology for Designing Knowledge-Driven Missions for Robots ​

Author: Guillermo GP-Lenza, Carmen DR. Pita-Romero, Miguel Fernandez-Cortizas, Pascual Campoy
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2601.20797v1 Announce Type: cross Abstract: This paper presents a comprehensive methodology for implementing knowledge graphs in ROS 2 systems, aiming to enhance the efficiency and intelligence of autonomous robotic missions. The methodology encompasses several key steps: defining initial and ...

📖 Read original article


34. Predict before you train: Scaling Laws for particle physics foundation models ​

Author: Jan-Lucas Uslu, Benjamin Nachman, Christopher Re
Published: 7/31/2026, 4:00:00 AM
Categories: hep-ex, cs.AI

arXiv:2607.23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been ...

📖 Read original article


35. Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact ​

Author: Mateusz Koz{\l}owski
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a ...

📖 Read original article


36. Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning ​

Author: Scott M. Norton
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2607.26059v1 Announce Type: cross Abstract: We report a striking phenomenon: deep reinforcement learning agents trained with frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, without any sparsity-inducing objective. In t...

📖 Read original article


37. Large-Scale ChatBot Validation Through Customer Digital Twin Simulations ​

Author: Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26060v1 Announce Type: cross Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation...

📖 Read original article


38. Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football ​

Author: Mouad Zemzoumi, Amine Abouaomar
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26061v1 Announce Type: cross Abstract: Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams. This paper presents Sim2Win, a team-agnostic, event-based pre-match t...

📖 Read original article


39. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities ​

Author: Karly V. Coffey, Gloria L. Krahn, John P. Hanley, Jacob E. Neely
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differe...

📖 Read original article


40. Archetypes or ability? Clustering for modelling student mathematical competence ​

Author: Benjamin Mawdsley, Tom Quilter, Richard Turner, Sarah Jackson, Paul Edwards
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2607.26063v1 Announce Type: cross Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. In this work, we e...

📖 Read original article


41. The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science ​

Author: Belinda Mo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.26064v1 Announce Type: cross Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output a...

📖 Read original article


42. Do Methods Support the Claims? Intra-Paper Verification for Peer Review ​

Author: Ranjitha Shivaprasad Ballakuraya, Arash Mahyari, Ashok Srinivasan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL

arXiv:2607.26066v1 Announce Type: cross Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper's claimed contributions against prior literatu...

📖 Read original article


43. The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty ​

Author: Amanda La Hadi, Muhammad Johan Alibasa, Guanliang Chen, A. Taufiq Asyhari
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.26067v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment...

📖 Read original article


44. The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem ​

Author: Sivasathivel Kandasamy
Published: 7/31/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.CY, q-fin.EC

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regu...

📖 Read original article


45. AI Security Priorities: A Field-Wide Agenda ​

Author: Gil Gekker, Rachel Steratore, Everett Smith, Asher Brass-Gershovich, Varun Gandhi, Nicole Nichols, Vijay Bolina, Buck Shlegeris, Lisa Einstein, Dan Lahav, Omer Nevo, Sella Nevo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CR

arXiv:2607.26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, ...

📖 Read original article


Author: Guanming Xiong, Penghui Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, prep...

📖 Read original article


47. GuidedRAG: Semantic Steering of Retrieval-Augmented Generation ​

Author: Matthijs Jansen op de Haar, Tobias St"ahle, Lorenzo Gatti
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26071v1 Announce Type: cross Abstract: In this work, we propose GuidedRAG, a novel extension to traditional Retrieval-Augmented Generation (RAG) that introduces a dedicated selection stage and semantic steering during retrieval. In contrast to current state-of-the-art RAG approaches, whic...

📖 Read original article


48. IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval ​

Author: Changyu Du, Alexander Vosseler, Filippo Mazza, Andr'e Borrmann
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reuse information from ...

📖 Read original article


49. IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations ​

Author: David Kaleko, Sergey Ivanov, Md Mofijul Islam
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domain specialists 20 to ...

📖 Read original article


50. FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents ​

Author: Lingteng Zeng, Yifan Jin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26076v1 Announce Type: cross Abstract: Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse can remove GPU-bound generation work, yet response caches require dependency consistency when filings...

📖 Read original article


51. Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate ​

Author: Fangnian Wang, Nicholas Tan Jerome, Thomas Jordan, Frank Simon
Published: 7/31/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.26078v1 Announce Type: cross Abstract: Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges due to hydrogen's low ignition energy and wide flammability range. Current monitoring systems are largely reac...

📖 Read original article


52. A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models ​

Author: Vivek Shukla, Varun Shukla, Atul, Divya Mishra, Mehul Kumar Das
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26102v1 Announce Type: cross Abstract: Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answ...

📖 Read original article


53. Weight and Height Estimation from a Single Human Image Captured in the Wild ​

Author: Hira Yaseen, Arif Mahmood, Waqas Sultani
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.26104v1 Announce Type: cross Abstract: A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the wei...

📖 Read original article


54. TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions ​

Author: Xinran Liu, Shouqian Shi, Yutong Chen, Ge Wang, Xin-Wei Yao, Sheng Zhong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26107v1 Announce Type: cross Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for th...

📖 Read original article


55. GPT-Red: Automated Red Teaming via Self-Play at Scale ​

Author: Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cer'on Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, w...

📖 Read original article


56. Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models ​

Author: Yuvraj Verma
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.26117v1 Announce Type: cross Abstract: Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at all. We argue that this...

📖 Read original article


57. Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels ​

Author: Xinyu Yang, Tianxing Chen, Honghao Su, Minxuan Wang, Chenze Yu, Zhangzheng Tu, Yue Chen, Yuxiao Huo, Lingfeng Zhang, Yan Huang, Yan Qin, Shaolong Zhu, Qiwei Liang, Hekun Tian, Shujia Liu, Guangyu Chen, Junhao Gong, Zixuan Li, Wenwei Lin, Zijian Lin, Wenxuan Zhu, Eric J Chen, Yue Yuan, Qize Yu, Jiaqi Liang, Haowen Yan, Hengfei Zhao, Weijie Wan, Zikun Xiao, Junyuan Tang, Baijun Chen, Kai-Chong Lei, Kaixuan Wang, Kailun Su, Zanxin Chen, Yao Mu, Renjing Xu, Chuqiao Lyu, Qi Xiong, Ping Luo, Wenbo Ding
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CY

arXiv:2607.26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustwo...

📖 Read original article


58. A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment ​

Author: Hua Qian, Manisha Kotha, Tuan Tran, Jennifer Shin, Haining Zheng
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.26170v1 Announce Type: cross Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background interference; a color-based ...

📖 Read original article


59. Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition ​

Author: Chandra Sripada, Richard Lewis
Published: 7/31/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CL

arXiv:2607.26179v1 Announce Type: cross Abstract: LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this fr...

📖 Read original article


60. (EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations ​

Author: Neta Kirmayer, David Tayouri, Andr'es Murillo, Motoyoshi Sekiya, Asaf Shabtai, Rami Puzis
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26201v1 Announce Type: cross Abstract: Security operations centers rely on anomaly detection systems to flag suspicious events. Feature-level explanations for anomaly detectors offer limited value for operational investigations. To effectively handle alerts, analysts need to know contextu...

📖 Read original article


61. Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges ​

Author: Quim Motger, Marc Oriol, Jordi Marco, Xavier Franch
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.26212v1 Announce Type: cross Abstract: Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iteratively converge ...

📖 Read original article


62. Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring ​

Author: Ahmed Ibrahim
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.26220v1 Announce Type: cross Abstract: Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: ...

📖 Read original article


63. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech ​

Author: Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avenuti, Felice Dell'Orletta, Stefano Cresci
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2607.26236v1 Announce Type: cross Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and ...

📖 Read original article


64. Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection ​

Author: Nicolas Gutowski, Fabien Chhel, Alexandre Letard, Sylvain Lamprier
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.26273v1 Announce Type: cross Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. We do not aim at identifying a single optimal arm; inste...

📖 Read original article


65. Entity Resolution in Practice: Lessons from a Self-Serve Pipeline ​

Author: Kaushik Pavani, Ganga Aluri, Pravin Jadhav, Neeraj Prasad, Kiran Sanka
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26298v1 Announce Type: cross Abstract: We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. (1) No single matching algorithm wins everywhere - a self-serve pip...

📖 Read original article


66. AgentGUI: An Interface for Observing and Steering Long-Running AI Agents ​

Author: Xuan Zhao, Jiwoong Sohn, Qinyue Zheng, Michael Moor
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2607.26300v1 Announce Type: cross Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduc...

📖 Read original article


67. SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation ​

Author: Gaston Besanson
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betra...

📖 Read original article


68. StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents ​

Author: Ads Dawson, Adrian Wood
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats ach...

📖 Read original article


69. Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach ​

Author: Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We prop...

📖 Read original article


70. Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks ​

Author: Xin Xu, Siru Tao, Kaizhen Tan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26344v1 Announce Type: cross Abstract: A gradient-based GNN explainer given a molecule with two chemically equivalent nitro groups assigns them attribution scores that are equal to the last bit. It cannot do otherwise: message passing is exactly permutation equivariant, so any automorphis...

📖 Read original article


71. When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses ​

Author: Zihan Chen, Di Zhu, Lei Nico Zheng
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC

arXiv:2607.26348v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answe...

📖 Read original article


72. Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research ​

Author: Jaber Daneshamooz, Eugene Vuong, Alagappan Ramanathan, Manni Moghimi, Haarika Manda, Satyam Kumar, Snithik Thode, Satyandra Guthula, Sylee Beltiukov, Dongsu Han, Tarun Mangla, Sangeetha Abdu Jyothi, Walter Willinger, Arpit Gupta
Published: 7/31/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.MA

arXiv:2607.26352v1 Announce Type: cross Abstract: Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a hypothesis) and generating the data that tests it. Consider a concrete case: does a bulk BBR downlo...

📖 Read original article


73. High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption ​

Author: Loong Kuan Lee, Ragavi Krishnamoorthy, Nico Piatkowski
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.26357v1 Announce Type: cross Abstract: The problem of learning the graphical Markov blanket (MB) of a variable from data has applications in many areas such as structure learning for Bayesian networks and Markov random fields, causal discovery, and feature selection. However, a common ass...

📖 Read original article


74. Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning ​

Author: Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT

arXiv:2607.26358v1 Announce Type: cross Abstract: Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objectiv...

📖 Read original article


75. Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text ​

Author: Aman Kumar, Lasitha Vidyaratne, Dipanjan D Ghosh, Arnab Chakrabarti, Ahmed K Farahat
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26368v1 Announce Type: cross Abstract: Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may als...

📖 Read original article


76. Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment ​

Author: Yitong Shen, Cheng Guo, Peiliang Wang, Jingzhe Zhang, Yi Sheng, Haopeng Zhang, Hongfei Xue, Yili Ren
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26381v1 Announce Type: cross Abstract: Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present...

📖 Read original article


77. Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction ​

Author: Xin Xu, Chengrui Wu, Jiayu Lu, Kaizhen Tan, Siru Tao, Hanzhe Hong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CR

arXiv:2607.26385v1 Announce Type: cross Abstract: Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribut...

📖 Read original article


78. Misalignment Has a Personality: A Big Five Account of Emergent Misalignment ​

Author: Hasibur Rahman, Smit Desai
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26389v1 Announce Type: cross Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corp...

📖 Read original article


79. Voice Memory for Agentic Speech Recognition ​

Author: Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko, Zhehuai Chen, Jagadeesh Balam, Boris Ginsburg
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD, eess.AS

arXiv:2607.26410v1 Announce Type: cross Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchr...

📖 Read original article


80. FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing ​

Author: Hongyang Wang, Yichen Shi, Hongrui Li, Yiru Huo, Jun Feng, Zitong Yu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26432v1 Announce Type: cross Abstract: Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent...

📖 Read original article


81. Reinforcement Learning on Cost-Constrained Quadrupedal Hardware ​

Author: Javier C. Weddington, Bence P. "Olveczky, Stephen A. Baccus
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.26434v2 Announce Type: cross Abstract: Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuat...

📖 Read original article


82. Mergeable Model-Side Aggregation States for Long-Context Language Models ​

Author: Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26448v1 Announce Type: cross Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped statistics, which ...

📖 Read original article


83. ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models ​

Author: Ruxi Gu, Zhenliang Zhang, Wei Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26455v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradig...

📖 Read original article


84. PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories ​

Author: Zekun Ren, Hongzhao Tan, Jiaen Yee, Kedar Hippalgaonkar
Published: 7/31/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2607.26464v1 Announce Type: cross Abstract: Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a human-centered graphical user interface (GUI) orchestration layer, PUDA creates a command-line runtime environme...

📖 Read original article


85. Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection ​

Author: Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei, Xinhai Chen, Xiangyang Wang, Yigui Feng, Shuai Li, Gencheng Liu, Keqi Yang, Qinglin Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.26472v1 Announce Type: cross Abstract: Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, an...

📖 Read original article


86. LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving ​

Author: Ming-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka, Tushar Krishna, Muhammed Ahosan Ul Karim, Shimeng Yu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.ET, cs.LG

arXiv:2607.26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. A key contributor to chip energy dissipation is dat...

📖 Read original article


87. Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning ​

Author: Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26509v1 Announce Type: cross Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationa...

📖 Read original article


88. HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models ​

Author: Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju Durai Raju
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26515v1 Announce Type: cross Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of ...

📖 Read original article


89. A Graph-Native Bitemporal Memory Store for Conversational AI Agents ​

Author: Alp Niksarli, Gopesh Baheti
Published: 7/31/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.IR

arXiv:2607.26520v1 Announce Type: cross Abstract: Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window, or delegating to a third-party memory service, either exhaust the model's context budget or send p...

📖 Read original article


90. The Art of Not Forgetting A Local Learning Architecture for Continual Learning ​

Author: Ashmith Atmuri, Yashaswini Rao Bhogarajula
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26523v1 Announce Type: cross Abstract: We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores them in a two-tier competitive memory, and learns through local updates without end-to-end backpropagation thr...

📖 Read original article


91. Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression ​

Author: Manuel Rodriguez
Published: 7/31/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CE

arXiv:2607.26528v1 Announce Type: cross Abstract: Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, where state variables are often coupled through shared physical parameters. Independent symbolic regression can ...

📖 Read original article


92. AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control ​

Author: Jingbo Cui, Jitao Zhao, Di Jin, Dongxiao He
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26533v1 Announce Type: cross Abstract: Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental source of relational semantics in graphs, the transferability of topological patterns has long been central to...

📖 Read original article


93. A Persona-based Rate Action Index ​

Author: Hayden Helm, Andrew Dassori
Published: 7/31/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG

arXiv:2607.26545v1 Announce Type: cross Abstract: We propose an index for predicting the U.S.\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based on how a collection of personas responds to current market conditions. To construct the index, we ...

📖 Read original article


94. ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform ​

Author: Xiaoxiao Jiang, Suyi Li, Sheng Yao, Tianyu Feng, Lingyun Yang, Dapeng Nie, Haoran Yang, Wei Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2607.26566v1 Announce Type: cross Abstract: Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning...

📖 Read original article


95. Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks ​

Author: Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text...

📖 Read original article


96. Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models ​

Author: Nand Lal Yadav, Rajesh Kumar, Satyendra Singh, Sudhakar Singh
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.26580v1 Announce Type: cross Abstract: With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. Convolutional neural networks have given promising results when used for diagnosing diseases using im...

📖 Read original article


97. One Run Is Not an Idea: The Implementation Lottery in Automated Research ​

Author: Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Chenyan Xiong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.26587v1 Announce Type: cross Abstract: Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the par...

📖 Read original article


98. A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents ​

Author: Zhoupeng Shou, Xiaodong Hong, Congjing Ren, Jingdai Wang, Yongrong Yang, Zuwei Liao
Published: 7/31/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2607.26594v1 Announce Type: cross Abstract: PID tuning for chemical processes commonly relies on identified process models, whereas plant engineers often retune loops iteratively by observing responses, diagnosing deficiencies, adjusting gains, and validating the result. This work formalizes t...

📖 Read original article


99. Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution ​

Author: Mingkuan Feng, Zhengqi Wen, Jianhua Tao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26596v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction t...

📖 Read original article


100. Living-Harness Is an Interactive-Agent Evolver ​

Author: Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL

arXiv:2607.26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactio...

📖 Read original article


101. WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models ​

Author: Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou, Kuo Cai, Sheng Yu, Qiang Luo, Jian Liang, Ruiming Tang, Fei Pan, Peng Jiang, Wenwu Ou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26621v2 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recommendation models (FRMs). Existing approaches typically enhance recommendation with explicit Chain-of-Thought (CoT...

📖 Read original article


102. Understanding Context Sampling in TabPFN on Small Tabular Datasets ​

Author: Mohammed Abdullah
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26628v1 Announce Type: cross Abstract: TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the...

📖 Read original article


103. Guarding Organizations Against Malware Risk: A Novel Graph-Based Malware Detection Method ​

Author: Yinan Gao, Jiarong Xu, Xiaohang Zhao, Xiao Fang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26634v1 Announce Type: cross Abstract: Organizational digitalization expands cybersecurity risks, making cybersecurity an increasingly important research area in Information Systems (IS). Among these risks, malware has become a pervasive and destructive threat. Byte-based machine learning...

📖 Read original article


104. Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability ​

Author: Sizhe Zhou, Sheldon Yu, Hui Wei, Junda Wu, Siru Ouyang, Yizhu Jiao, Shijia Pan, Julian McAuley, Yu Zhang, Tong Yu, Jiawei Han
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26637v1 Announce Type: cross Abstract: Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this medium: prior ...

📖 Read original article


105. Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses ​

Author: Haoyu Zhang, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Mohammad Zandsalimy, Shanu Sushmita
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individually harmles...

📖 Read original article


106. FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking ​

Author: Mu~noz-Haro Javier, Teruel Andres, Tolosana Ruben, DeAlcala Daniel, Vera-Rodriguez Ruben, Morales Aythami, Fierrez Julian
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26641v1 Announce Type: cross Abstract: Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypas...

📖 Read original article


107. FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows ​

Author: Wenzhe He, Meng Wang, JiaWei Qian, Jinfeng Xu, Ying Liu, Ruihui Li
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26645v1 Announce Type: cross Abstract: Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during inference, they are initialized...

📖 Read original article


108. Physically Real-time Infrared Attack against Optical Flow Estimation Networks ​

Author: Shen You, Wei Jiang, Jiarui Liu, Yijian Ye, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26651v1 Announce Type: cross Abstract: With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion detection have become increasingly mature and relevant to human lives. In particular, Optical Flow Es...

📖 Read original article


109. Constitutional Midtraining: Content Presence Drives Alignment Gains ​

Author: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG

arXiv:2607.26654v2 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitution...

📖 Read original article


110. Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection ​

Author: Yikun Li, Ting Zhang, Jiakun Liu, Jinfeng Jiang, Yuheng Yieh, Yixin Yang, Wen Bin Leow, Yide Yin, Yintong Huo, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE

arXiv:2607.26656v1 Announce Type: cross Abstract: Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the function to be cla...

📖 Read original article


111. Scientific Knowledge Discovery in the Age of Large Language Models ​

Author: Eleni Adamidi, Serafeim Chatzopoulos, Thanasis Vergoulis
Published: 7/31/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.CL, cs.IR

arXiv:2607.26670v1 Announce Type: cross Abstract: The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depend heavily on manually formulated queries and effortful manual inspection. Generative large language...

📖 Read original article


112. Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL ​

Author: Mingxuan Che, Tsung-Yuan Tseng, Theresa Eimer, Marius Lindauer, Alexander von Rohr
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26680v1 Announce Type: cross Abstract: Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. We propose...

📖 Read original article


113. MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation ​

Author: Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, Chih-Pin Tan, Fang-Duo Tsai, Yi-Hsuan Yang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2607.26698v1 Announce Type: cross Abstract: Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voiced/unvoiced (V/UV) tags for con...

📖 Read original article


114. Automated Multilabel Mpox Research Classification with Explainable Transformer Models ​

Author: Tanjim Taharat Aurpa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26700v1 Announce Type: cross Abstract: The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some regions. Research on Mpox is vital for several reasons, including vaccine development, diagnostic improvement, viral ...

📖 Read original article


115. FARI: Robust One-Step Inversion for Watermarking in Diffusion Models ​

Author: Jindong Yang, Han Fang, Weiming Zhang, Nenghai Yu, Kejiang Chen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.26723v1 Announce Type: cross Abstract: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that is both slow and error-prone. While the primary challenge in the watermarking setting is robustness a...

📖 Read original article


116. Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives ​

Author: Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng, Dayong Ye, Tianqing Zhu, Huan Huo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2607.26735v1 Announce Type: cross Abstract: Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion methods suffer from significant ...

📖 Read original article


117. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model ​

Author: Carlos Mu~noz-Romero, Jose A. Gonzalez-Lopez
Published: 7/31/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD

arXiv:2607.26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. In this work, we propose a Fac...

📖 Read original article


118. Multimodal fusion of visual and morphometric features for avian bone classification ​

Author: Nevio Dubbini, Lisa Yeomans, Marco Pavia, Ramazan Parmaksiz, Ayse Atas Hooglugt, Gabriele Gattiglia, Beatrice Demarchi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26743v1 Announce Type: cross Abstract: Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, particularly for the identification of avian skeletal remains. This study presents a proof-of-concept multimodal ...

📖 Read original article


119. An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI ​

Author: Harshiddhi Pathak, Gowtham Reddy N, Mrinal Acharya, Manjunatha Mahadevappa
Published: 7/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.HC, cs.LG, eess.SP

arXiv:2607.26746v1 Announce Type: cross Abstract: Accurate identification of Alzheimers disease (AD) using resting-state functional magnetic resonance imaging (rs-fMRI) remains challenging due to the high dimensionality, noise, and complex inter-regional dependencies inherent in functional brain con...

📖 Read original article


120. Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text ​

Author: Lucas Zamora Vera, Jose A. Gonzalez-Lopez
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, eess.SP

arXiv:2607.26751v1 Announce Type: cross Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the o...

📖 Read original article


121. Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification ​

Author: Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich, Elena Kozachok, Egor Ushakov, Oleg Samovarov
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26765v1 Announce Type: cross Abstract: Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture artifacts. We study how data augmentation improves the robustness of a binary malignant-versus-non-mal...

📖 Read original article


122. MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities ​

Author: Francesco Tosoni
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.SE

arXiv:2607.26766v1 Announce Type: cross Abstract: Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the trade-off between the low latency of traditional Information Retrieval (IR) and the precision of Deep Learning (...

📖 Read original article


123. See2Think: Do Multimodal Models Really Use Intermediate Visual States? ​

Author: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26769v1 Announce Type: cross Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections ...

📖 Read original article


124. Journey Operators for Structured Multi-Axis Composition ​

Author: Mahesh Godavarti
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26775v1 Announce Type: cross Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog....

📖 Read original article


125. SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution ​

Author: Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang, Zhuowen Han, Zi-Han Wang, Ziang Ye, Qi Gu, Xunliang Cai, Weiwen Liu, Yongliang Shen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26784v1 Announce Type: cross Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus o...

📖 Read original article


126. SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response ​

Author: Lehan Wang, Boli Chen, Ruixue Ding, Pengjun Xie, Jinwei Huang, Zhendong Liu, Shuo Wang, Tao Lei, Xin Ouyang, Xiaomeng Li
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybe...

📖 Read original article


127. Crossing-Free Probabilistic K-Line Forecasts Without Retraining ​

Author: Runyao Yu, Yuchen Tao, Yujie Chen, Wentao Wang, Derek W. Bunn
Published: 7/31/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.CE, cs.LG, q-fin.CP

arXiv:2607.26792v1 Announce Type: cross Abstract: Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). However, it introduces two consistency problems: quantile crossing and K-line crossing. Quantile crossing occurs when a higher-...

📖 Read original article


128. FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning ​

Author: Zhaoyang Ma, Zhihao Wu, Xin Gao, Lipo Wang, Youfang Lin, Jing Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26801v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global know...

📖 Read original article


129. A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities ​

Author: Wenhao Yang, Runzhi He, Minghui Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.26819v1 Announce Type: cross Abstract: Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-off...

📖 Read original article


130. AI as Friction for Reflection Support in Ideation ​

Author: Janin Koch, Xiaohan Liao, G'ery Casiez
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.26827v1 Announce Type: cross Abstract: Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output translate into more value for the designer. We argue, however, that this framing leaves out so...

📖 Read original article


131. Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility ​

Author: Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Chen Ma, Yiyan Qi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26828v2 Announce Type: cross Abstract: Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, re...

📖 Read original article


132. From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs ​

Author: Ruikang Zhang, Shuo Wang, Qi Su
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26853v1 Announce Type: cross Abstract: Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-relate...

📖 Read original article


133. ReCo: Reweighting GRPO Against Distributional Concentration ​

Author: Junoh Park, Junseo Hwang, Wonguk Cho, Taesup Kim
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26862v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large,...

📖 Read original article


134. Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents ​

Author: Amirmohammad Farzaneh, Osvaldo Simeone
Published: 7/31/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.IT, cs.LG, math.IT

arXiv:2607.26865v1 Announce Type: cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their re...

📖 Read original article


135. Hearsay: Vision-Language Medical Diagnoses Without an Image ​

Author: Siddharth Vohra
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.CY

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across che...

📖 Read original article


136. Human diversity fuels collective creativity that large language models cannot simulate or sustain ​

Author: Mengchen Dong, Hiromu Yakura
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2607.26899v1 Announce Type: cross Abstract: Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogethe...

📖 Read original article


137. Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing ​

Author: Brandon Gower-Winter, Georg Krempl
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26908v1 Announce Type: cross Abstract: In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. This phenomena is known as Outcome Performativity. This paper formalises an approach for detecting Outco...

📖 Read original article


138. BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories ​

Author: Zhe Liu, Quan Lu, Zhaohui Du, Zhe Wang, Huanbo Jin, Jiaming Gu, Qi Wang, Ting Xiao, Minting Pan, Dongzhan Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.26914v1 Announce Type: cross Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position...

📖 Read original article


139. Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning ​

Author: Hongliang Zhang, Zhongyuan Yu, Guijuan Wang, Tianqing He, Wenshuo Ma, Xiaosong Zhang, Jiguo Yu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.26933v1 Announce Type: cross Abstract: Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical h...

📖 Read original article


140. Progressive Multimodal Alignment for Continual Instruction Tuning ​

Author: Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26947v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visua...

📖 Read original article


141. SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception ​

Author: Gabe Everett, Brice Gunter, Ryan Vander Stelt, Cleiver Ruiz-Martinez, Blake Hull, Juan Rojas
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.26985v1 Announce Type: cross Abstract: Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-s...

📖 Read original article


142. BayesAME: Bayesian Active Model Evaluation ​

Author: Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.27023v1 Announce Type: cross Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Cu...

📖 Read original article


143. CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation ​

Author: Fengming Yu, Haiwei Pan, Kejia Zhang, Chunling Chen, Jian Guan, Baoying Ma
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.27054v1 Announce Type: cross Abstract: Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous se...

📖 Read original article


144. ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection ​

Author: Paul Julius K"uhn, Saptarshi Neil Sinha, Tiago Kleist, Richard Hoffmann, Arjan kuijper, Michael Weinmann
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27065v2 Announce Type: cross Abstract: While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline th...

📖 Read original article


145. SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence ​

Author: Chuanzhi Xu, Zihan Deng, Huiqi Liang, Chengkun Yue, Zhanlin Cui, Pengfei Ye, Weidong Cai
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27066v1 Announce Type: cross Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, i...

📖 Read original article


146. Visual Credit Audit for Multimodal Spatial Reasoning ​

Author: Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27069v2 Announce Type: cross Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives...

📖 Read original article


147. Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise ​

Author: Vaneet Aggarwal
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC

arXiv:2607.27073v1 Announce Type: cross Abstract: We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. While static regret is well-understood, achiev...

📖 Read original article


148. MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair ​

Author: Xuanze Chen, Xukang Xie, Wentao Fu, Jiajun Zhou, Shanqing Yu, Qi Xuan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.27080v1 Announce Type: cross Abstract: Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly sh...

📖 Read original article


149. Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents ​

Author: Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.27083v1 Announce Type: cross Abstract: As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds...

📖 Read original article


150. SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context ​

Author: Zihan Deng, Chuanzhi Xu, Huiqi Liang, Haoyang Li, Xiaozhen Zhong, Lequan Yu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27084v1 Announce Type: cross Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly des...

📖 Read original article


151. MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning ​

Author: Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performa...

📖 Read original article


152. DLAM: Distributional Latent Actions with Temporal Constraints ​

Author: Zuojin Tang, Feifan Luo, Haoyun Liu, Botai Yuan, Dekang Qi, Ronghan Chen, Yandan Yang, Tong Lin, Xinyuan Chang, Mu Xu, Bin Liu, De Ma, Zhiheng Ma
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.27138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may ...

📖 Read original article


153. Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark ​

Author: Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.27143v1 Announce Type: cross Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provid...

📖 Read original article


154. Anatomy Contextualized Adaption of CT Foundation Models ​

Author: Roshan Kenia, Stephanie L McNamara, William Lotter
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27154v1 Announce Type: cross Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training a...

📖 Read original article


Author: Ji Xin, Xiao Xiao, Ishan Bhatt, Vinesh Gudla, Trace Levinson, Raochuan Fan, Shishir Kumar Prasad, Prakash Putta, Tejaswi Tenneti
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.27172v1 Announce Type: cross Abstract: Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outc...

📖 Read original article


156. The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making ​

Author: Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos, Mohammad Amin Samadi, Luise Mehner, Seehee Park, Spencer JaQuay
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making us...

📖 Read original article


157. APEX-Accounting ​

Author: Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2607.27189v2 Announce Type: cross Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing repo...

📖 Read original article


158. Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility ​

Author: Chaoxiong Ma, Yan Liang, Huixia Zhang, Hao Sun
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2504.04128v3 Announce Type: replace Abstract: In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the majority may be either critical evidence supporting the correct decision or anomalous evidence supporting an ...

📖 Read original article


159. Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning ​

Author: Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang, Na Su, Tao Zhou, Geng Chen, Wen Fan, Yi Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2508.16129v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them focus exclu...

📖 Read original article


160. HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring ​

Author: Xin Wang, Ting Dang, Xinyu Zhang, Vassilis Kostakos, Michael J. Witbrock, Hong Jia
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG

arXiv:2509.07260v5 Announce Type: replace Abstract: Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life. Previous studies on large language models (LLMs) have highl...

📖 Read original article


161. Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests ​

Author: Alexandra Yost, Shreyans Jain, Shivam Raval, Grant Corser, Allen Roush, Nina Xu, Jacqueline Hammack, Ravid Shwartz-Ziv, Amirali Abdullah
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.22170v3 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies using si...

📖 Read original article


162. Balancing Centralized Learning and Distributed Self-Organization: A Hybrid Model for Embodied Morphogenesis ​

Author: Takehiro Ishikawa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.10101v2 Announce Type: replace Abstract: Background: both embodied intelligence and developmental morphogenesis depend on a division of labour between centralized guidance and distributed material dynamics, but the amount of top-down control needed to steer self-organization remains uncle...

📖 Read original article


163. BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models ​

Author: Yujie Lin, Jiayao Ma, Qingguo Hu, Wenbo Li, Genji Li, Derek Wong, Jinsong Su
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-unaware perspective that enforces uniform treatment across demograph...

📖 Read original article


164. How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series ​

Author: Mathieu Cherpitel, Janne Luijten, Thomas B"ack, Camiel Verhamme, Martijn Tannemaat, Anna V. Kononova
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.10191v2 Announce Type: replace Abstract: Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs), yet the signals' high and heterogeneous sampling rates pose substantial computational challenges for featu...

📖 Read original article


165. AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing ​

Author: Zhenhua Xu, Dongsheng Chen, Shuo Wang, Jian Li, Chengjie Wang, Meng Han, Yabiao Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2601.11007v2 Announce Type: replace Abstract: LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion and adaptability. They typically under-model dynamic environmental information and assume largely static scene...

📖 Read original article


166. TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech ​

Author: Girish A. Koushik, Helen Treharne, Diptesh Kanojia
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MM, cs.SI

arXiv:2601.11178v3 Announce Type: replace Abstract: Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with high accur...

📖 Read original article


167. How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm ​

Author: Taisei Hishiki, Takaya Arita, Reiji Suzuki
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT, cs.MA

arXiv:2604.12250v2 Announce Type: replace Abstract: This study examines how memory shapes the collective and cooperative dynamics of Large Language Model (LLM) agents in a multi-agent system. To this end, we extend the Social Particle Swarm (SPS) model, in which agents move in a two-dimensional spac...

📖 Read original article


168. MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval ​

Author: Shaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton, Navid Safaei, Sultan Albarakati, William T. Freeman, Antonio Torralba
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.DL, cs.IR, cs.LG

arXiv:2604.18584v2 Announce Type: replace Abstract: Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a high-quality, large-scale, mu...

📖 Read original article


169. On the Hybrid Nature of ABPMS Process Frames and its Implications on Automated Process Discovery ​

Author: Anti Alman, Izack Cohen, Avigdor Gal, Fabrizio Maria Maggi, Marco Montali
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.22455v2 Announce Type: replace Abstract: A core component of any AI-Augmented Business Process Management System (ABPMS) is the process frame, which gives the system process-awareness and defines its maximal behavioral boundaries. Compared to traditional process models, the process frame ...

📖 Read original article


170. The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure ​

Author: Qiqi Liu, Runhan Song, Shilin Ye
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.17480v3 Announce Type: replace Abstract: Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates new attack surfaces. We identify semantic hijacking, an attack in which harmful requests are conce...

📖 Read original article


171. Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries ​

Author: Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE

arXiv:2605.19576v3 Announce Type: replace Abstract: Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval degradation, false-positive injections, and performance stagnation. Rec...

📖 Read original article


172. Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents ​

Author: Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2605.22148v2 Announce Type: replace Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation shows that LLM-authored skills deliver $+0.0$pp over no-skill baselines while human-curated ones ...

📖 Read original article


173. RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention ​

Author: Yang Liu, Zhaokai Luo, Huayi Jin, Zhiyong Wang, Ruozhou He, Boyu Wang, Guanjie Chen, Tao Xie, Junhao Hu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.06256v3 Announce Type: replace Abstract: As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serving concurrency, cache reuse, and distributed scalability. Multiple ...

📖 Read original article


174. BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation ​

Author: Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2606.19651v2 Announce Type: replace Abstract: Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. Latent diffusion has ...

📖 Read original article


175. Intent-Governed Tool Authorization for AI Agents ​

Author: Genliang Zhu, Chu Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.22916v3 Announce Type: replace Abstract: Tool-using AI agents commonly operate under integration credentials whose static permissions exceed a user's current request. We present Intent-Governed Access Control (IGAC), a server-side authorization layer that converts a trusted request into a...

📖 Read original article


176. Matilda: Engine-Agnostic Search with Human Policy Guidance ​

Author: Jason Carlson
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.25176v3 Announce Type: replace Abstract: Chess engines have evolved from search-based systems optimized for strength to neural policies optimized for predicting human decisions. Existing approaches largely separate these goals: search engines achieve superhuman strength but poorly model h...

📖 Read original article


177. ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks ​

Author: Qitai Tan, Zefang Zong, Mo Li, Yipeng Shi, Yang Li, Peng Chen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.27814v4 Announce Type: replace Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but ...

📖 Read original article


178. Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing ​

Author: Dvir Alsheich, Adar Peleg, Ben Hagag, Rom Himelstein, Amit Levi, Avi Mendelson
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2606.30555v3 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effective orchestration in these environments requires robust routing me...

📖 Read original article


179. FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents ​

Author: Yufeng Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.05682v2 Announce Type: replace Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exp...

📖 Read original article


180. When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals ​

Author: Kaihua Ding
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.08065v2 Announce Type: replace Abstract: LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a k...

📖 Read original article


181. Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting? ​

Author: Qihang Zhang, Siyao Zhang, Letao Kang, Wenzhe Liang, Miao Zhang, Zhao Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12462v2 Announce Type: replace Abstract: Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each node and all nodes across the traffic...

📖 Read original article


182. Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents ​

Author: Lujia Zhang, Xingzhou Chen, Hongwei Feng
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.15715v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workfl...

📖 Read original article


183. SkillSight: Calibrating Generic Content Bias for Skill Retrieval ​

Author: Jinying Xiao, Bin Li, Xiaopeng Li, Jianling Li, Jiacheng Jie, Xiaodong Liu, Ma Jun, Chao Wang, Nyima Tashi, Jie Yu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.18785v2 Announce Type: replace Abstract: As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, over...

📖 Read original article


184. ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D ​

Author: Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG

arXiv:2607.19321v2 Announce Type: replace Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potenti...

📖 Read original article


185. Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models ​

Author: Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22663v2 Announce Type: replace Abstract: Block diffusion is the dominant approach for scaling discrete diffusion language models (dLLMs), as fixed-size blocks preserve parallel decoding while keeping quadratic attention costs tractable. Yet blockwise generation creates a structural weakne...

📖 Read original article


186. The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards ​

Author: Keyu Li, Jin Gao, Dequan Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24063v2 Announce Type: replace Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation an...

📖 Read original article


187. Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness ​

Author: Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.24162v2 Announce Type: replace Abstract: Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimiza...

📖 Read original article


188. TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs ​

Author: Federico Valletta, Giacomo Longo, Enrico Russo, Alessio Merlo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.24563v2 Announce Type: replace Abstract: Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to d...

📖 Read original article


189. Do Models Fake Alignment Without Clear Consequences? ​

Author: Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24758v2 Announce Type: replace Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignme...

📖 Read original article


190. PLATO: Pointer Learner for Agent and Task Openness ​

Author: Alireza Saleh Abadi, Leen-Kiat Soh, Daniel Alan Redder, Adam Eck, Prashant Doshi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.25082v2 Announce Type: replace Abstract: Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to mul...

📖 Read original article


191. Observing sycophantic AI validate others reduces its appeal but not its persuasiveness ​

Author: Meryl Ye, Robert Kraut, Steve Rathje
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.25166v2 Announce Type: replace Abstract: AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindness"). We tested whether incre...

📖 Read original article


192. The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape ​

Author: Deyao Hong, Kehan Zheng, Qian Li, Jun Zhang, Jie Jiang, Hongning Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2607.25253v2 Announce Type: replace Abstract: Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need befo...

📖 Read original article


193. Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales ​

Author: Genliang Zhu (Accentrust, Georgia Institute of Technology), Chu Wang (Accentrust, University of Illinois Urbana-Champaign)
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.25364v2 Announce Type: replace Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that conv...

📖 Read original article


194. A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes ​

Author: Mingkang Liu, Huize Yu, Yanbin Gao, Nan Yao, Xiang Chen, Lei Shen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.25597v2 Announce Type: replace Abstract: The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion participation. This interplay operates through the redistribution of electron density across donor, anion, and c...

📖 Read original article


195. Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification ​

Author: Chenrui Shi, Yuwei Wu, Yang Liu, Ruining Feng, Zirui Shang, Zhi Gao, Lifeng Fan, Che Sun
Published: 7/31/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.25904v2 Announce Type: replace Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals...

📖 Read original article


196. Pushing the Frontier on Approximate EFX Allocations ​

Author: Georgios Amanatidis, Aris Filos-Ratsikas, Alkmini Sgouritsa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.DM

arXiv:2406.12413v3 Announce Type: replace-cross Abstract: We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve approximate envy-freeness up to any good ($\alpha$-EFX). The state-of-the-art results on the problem include...

📖 Read original article


197. One-Frame Calibration with Siamese Network in Facial Action Unit Recognition ​

Author: Shuangquan Feng, Virginia R. de Sa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2409.00240v2 Announce Type: replace-cross Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis. Most existing AU recognition systems aim for cross-participant non-calibrated generalization (NCG) to unseen faces without further calibration. Howeve...

📖 Read original article


198. MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications ​

Author: Praveenkumar Kanithi, Cl'ement Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2409.07314v4 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To ...

📖 Read original article


199. The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness ​

Author: Krishna Subedi
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC

arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and ...

📖 Read original article


200. Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills ​

Author: Benned Hedegaard, Yichen Wei, Ziyi Yang, Ahmed Jaafar, Stefanie Tellex, George Konidaris, Naman Shah
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2504.17901v3 Announce Type: replace-cross Abstract: Task and motion planning (TAMP) is a well-established approach for solving long-horizon robot planning problems. Although TAMP methods have historically assumed that each task-level robot action, or skill, can be reduced to kinematic motion p...

📖 Read original article


201. When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents ​

Author: Prothit Sen, Sai Mihir Jakkaraju
Published: 7/31/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.HC

arXiv:2504.20903v4 Announce Type: replace-cross Abstract: How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint sequential adaptation in which two agents differ in a single, precisely specified way: the memory regim...

📖 Read original article


202. Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities ​

Author: Zhiwei Hao, Jianyuan Guo, Li Shen, Yong Luo, Han Hu, Guoxia Wang, Dianhai Yu, Yonggang Wen, Dacheng Tao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.01043v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. To mitigate this ch...

📖 Read original article


203. AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages ​

Author: Muzhe Wu, Yanzhi Zhao, Shuyi Han, Michael Xieyang Liu, Hong Shen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through social-technical lenses. However, in cross-functional industry teams, this work is often stalled by...

📖 Read original article


204. Equivariant Eikonal Neural Networks: Grid-Free, Scalable Travel-Time Prediction on Homogeneous Spaces ​

Author: Alejandro Garc'ia-Castellanos, David R. Wessels, Nicky J. van den Berg, Remco Duits, Dani"el M. Pelt, Erik J. Bekkers
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.16035v3 Announce Type: replace-cross Abstract: We introduce Equivariant Neural Eikonal Solvers, a novel framework that integrates Equivariant Neural Fields (ENFs) with Neural Eikonal Solvers. Our approach employs a single neural field where a unified shared backbone is conditioned on sign...

📖 Read original article


205. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent ​

Author: Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, Hao Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2507.02259v2 Announce Type: replace-cross Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity without performance degradation during extrapolation remains the ultimate challenge in long-text p...

📖 Read original article


206. FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing ​

Author: Shida Wang, Chaohu Liu, Yubo Wang, Linli Xu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable intellectual assets. Nevertheless, these AI assets remain vulnerable to unauthorized redistribution ...

📖 Read original article


207. Balancing Privacy and Efficiency: Music Information Retrieval via Additive Homomorphic Encryption ​

Author: William Zerong Wang, Dongfang Zhao
Published: 7/31/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CR

arXiv:2508.07044v2 Announce Type: replace-cross Abstract: Modern music retrieval runs on vector embeddings, and once these embeddings are shared for search or matching they can be copied, probed, or used to train generative models. Fully homomorphic encryption can compute on them but is impractical ...

📖 Read original article


208. GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning ​

Author: Jizhuo Chen, Diwen Liu, Jiaming Wang, Harold Soh
Published: 7/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2509.11594v3 Announce Type: replace-cross Abstract: GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapshot. The method uses a two stage curriculum: (1) a simple distance-visibility rule auto-labels a large dataset at low cost; and (2) a sm...

📖 Read original article


209. Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation ​

Author: Xin Yu, Cong Xie, Xunmei Liu, Tiantian Fan, Lingzhou Xue, Zhi Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.00192v3 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its representational capacity often lags behind full fine-tuning. Within the context of LoRA, a key open question i...

📖 Read original article


210. VideoNorms: Benchmarking Cultural Awareness of Video Language Models ​

Author: Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.CY

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce VideoNorms, a dataset of cu...

📖 Read original article


211. ARC-Encoder: learning compressed text representations for large language models ​

Author: Hippolyte Pilchen, Edouard Grave, Patrick P'erez
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.20535v2 Announce Type: replace-cross Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective approaches require...

📖 Read original article


212. $\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials ​

Author: Trisha Das, Mandis Beigi, Jacob Aptekar, Jimeng Sun
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2601.06300v2 Announce Type: replace-cross Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly amended component. We introduce \textit{eligibility criteria amendment prediction}, a novel N...

📖 Read original article


213. How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs ​

Author: Shivam Adarsh, Maria Maistro, Christina Lioma
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.06599v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change when context is intro...

📖 Read original article


214. DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English ​

Author: Jio Oh, Paul Vicinanza, Thomas Butler, Steven Euijong Whang, Dezhi Hong, Amani Namboori
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.22888v4 Announce Type: replace-cross Abstract: More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects and generate stereotyped responses for their speakers. We introduce DialectLLM, the first large-...

📖 Read original article


215. Structurally Separated Uncertainty in Supervised Latent Variable Models ​

Author: Tanmoy Mukherjee, Marius Kloft, Pierre Marquis, Zied Bouraoui
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.11219v2 Announce Type: replace-cross Abstract: Predictive uncertainty is commonly decomposed into epistemic and aleatoric components, but standard decompositions often produce strongly correlated estimates because both quantities are derived from the same predictive distribution. We study...

📖 Read original article


216. Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments ​

Author: Zehao Xu, Tony Paek, Kevin O'Sullivan, Attila Dobi
Published: 7/31/2026, 4:00:00 AM
Categories: stat.AP, cs.AI

arXiv:2602.16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments. LLM-based labeling provides a high-fidelity reference measurement, but is c...

📖 Read original article


217. PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging ​

Author: Jitindra Fartiyal, Pedro Freire, Sergei K. Turitsyn, Sergei G. Solovski
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.21987v3 Announce Type: replace-cross Abstract: Low-dose CT images are essential for reducing radiation exposure in cancer screening, pediatric imaging, and longitudinal monitoring protocols, but their quality is often degraded by noise from low-dose acquisition, patient motion, or scanner...

📖 Read original article


218. Ask don't tell: Reducing sycophancy in large language models ​

Author: Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While prior work has documen...

📖 Read original article


219. The Rise of AI in Weather and Climate Information and its Impact on Global Inequality ​

Author: Amirpasha Mozaffari, Amanda Duarte, Lina Teckentrup, Stefano Materia, Gina E. C. Charnley, Lluis Palma, Eulalia Baulenas Serra, Dragana Bojovic, Paula Checchia, Aude Carreric, Francisco Doblas-Reyes
Published: 7/31/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.LG

arXiv:2603.05710v2 Announce Type: replace-cross Abstract: AI development's current trajectory risks automating and amplifying the North-South divide in the global climate information system. Frontier models are built almost exclusively in the Global North, and this inequality continues through input...

📖 Read original article


220. Making Implicit Premises Explicit in Logical Understanding of Enthymemes ​

Author: Xuyao Feng, Anthony Hunter
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.06114v2 Announce Type: replace-cross Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language processing (NLP) methods for handling enthymemes can potentially identify enthymemes in text but the...

📖 Read original article


221. MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue ​

Author: Naifan Zhang, Ruihan Sun, Jinwei Su, Hengjie Yang, Zhengyuan Pan, Zhaohan Chen, Xiaofan Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.06194v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn interaction remains challenging due to sparse rewards and poor per-turn credit assignment. In emoti...

📖 Read original article


222. Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge ​

Author: Shuai Lu, Meng Wang, Jia Guo, Jiawei Du, Bo Liu, Shengzhu Yang, Weihang Zhang, Huazhu Fu, Huiqi Li
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.07131v4 Announce Type: replace-cross Abstract: Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is severely hindered by lacking domain-specific knowledge. In this work, we identify two structural deficiencie...

📖 Read original article


223. Gated Adaptation for Continual Learning in Human Activity Recognition ​

Author: Reza Rahimi Azghan, Gautham Krishna Gudur, Mohit Malu, Edison Thomaz, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.10046v2 Announce Type: replace-cross Abstract: Wearable sensors in Internet of Things (IoT) ecosystems increasingly support applications such as remote health monitoring, elderly care, and smart home automation, all of which rely on robust human activity recognition (HAR). Continual learn...

📖 Read original article


224. State-Dependent Safety Failures in Multi-Turn Language Model Interaction ​

Author: Pengcheng Li, Jie Zhang, Tianwei Zhang, Han Qiu, Zhang kejun, Weiming Zhang, Nenghai Yu, Wenbo Zhou
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2603.15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure rem...

📖 Read original article


225. Adaptively Robust LLM Monitoring via Activation Watermarking ​

Author: Toluwani Aremu, Daniil Ognev, Samuele Poppi, Nils Lukas
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY, cs.LG

arXiv:2603.23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and often openly available, so $\emph{adaptive}$ attackers with a local copy can search offline for prompts tha...

📖 Read original article


226. Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory ​

Author: Jon-Paul Cacioli
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.25112v3 Announce Type: replace-cross Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows (Type-1 accuracy) and how well its confidence signal tracks that knowledge (Type-2 metacognitive sensi...

📖 Read original article


227. GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring ​

Author: Xinyi Duan, Yuanrong Tang, Jiangtao Gong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2603.26807v2 Announce Type: replace-cross Abstract: The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) address these issues by incorporating externa...

📖 Read original article


228. REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage ​

Author: Smriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali, Shubham Ugare, Satish Chandra
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2604.01527v4 Announce Type: replace-cross Abstract: Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks and risks user experience, shadow deployment yields signals t...

📖 Read original article


229. Shot-based quantum encoding: a data-loading paradigm for quantum neural networks ​

Author: Basil Kyriacou, Viktoria Patapovich, Maniraman Periyasamy, Alexey Melnikov
Published: 7/31/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2604.06135v2 Announce Type: replace-cross Abstract: Efficient data loading remains a bottleneck for near-term quantum machine learning. Existing schemes (angle, amplitude, and basis encoding) either underuse the exponential Hilbert-space capacity or require circuit depths that exceed the coher...

📖 Read original article


230. The Fast Lane Hypothesis: Von Economo Neurons Implement a Biological Speed-Accuracy Tradeoff ​

Author: Esila Keskin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, q-bio.NC

arXiv:2604.09229v2 Announce Type: replace-cross Abstract: von Economo neurons (VENs) are large bipolar projection neurons found exclusively in the anterior cingulate cortex (ACC) and frontal insula of species with complex social cognition, including humans, great apes, cetaceans, and elephants. Thei...

📖 Read original article


231. Facial-Expression-Aware Prompting for Empathetic LLM Tutoring ​

Author: Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2604.15336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and...

📖 Read original article


232. BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels ​

Author: Mengfei Lan, Lecheng Zheng, Halil Kilicoglu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2604.15591v2 Announce Type: replace-cross Abstract: Effective biomedical information retrieval requires modeling domain semantics and hierarchical relationships among biomedical texts. Existing biomedical generative retrievers build on coarse binary relevance signals, limiting their ability to...

📖 Read original article


233. The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers ​

Author: Benjamin Minhao Chen, Xinyu Xie
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2604.24155v4 Announce Type: replace-cross Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in ...

📖 Read original article


234. Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation ​

Author: Yang Xiao, Huiyuan Chen, Kaiyuan Deng, Chao Jiang, Zinan Ling, Ruimeng Ye, Fei Wang, Xiaolong Ma, Bo Hui
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.08810v2 Announce Type: replace-cross Abstract: We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. CVA first summarizes frozen VFM frame embeddings into a semantic-consensus ancho...

📖 Read original article


235. Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval ​

Author: Jeffrey Flynt
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2605.11325v4 Announce Type: replace-cross Abstract: Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief store can achieve perfect recall and mask severe precision failures. We show this evaluation gap persist...

📖 Read original article


236. Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents ​

Author: Jingxing Wang, Chenyu Zhou, Zhihui Fu, Jun Wang, Weiwen Liu, Weinan Zhang, Jianghao Lin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.16986v2 Announce Type: replace-cross Abstract: Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. We call this challenge test-time compute-to-capability conversi...

📖 Read original article


237. Exact Symmetry as Algebra: A Machine-Verified Tensor Calculus that Enforces Physical Selection Rules ​

Author: Paulina Hoyos, Shashanka Ubaru, Dongsung Huh, Vasileios Kalantzis, Kenneth L. Clarkson, Misha Kilmer, Haim Avron, Lior Horesh
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.RA

arXiv:2605.20440v2 Announce Type: replace-cross Abstract: Symmetry is central to the physical sciences, yet machine learning usually captures it only approximately, leaving a residual per-step equivariance error $\varepsilon$ that compounds with depth $M$ as $M\varepsilon$, whereas exact equivarianc...

📖 Read original article


238. VistaHop: Benchmarking Long-Horizon Visual DeepSearch ​

Author: Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan, Ting Su, Haiying Sun, Jiajun Chai, Xiaohan Wang, Guojun Yin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2606.03273v2 Announce Type: replace-cross Abstract: Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps...

📖 Read original article


239. An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers ​

Author: Ossi Lehtinen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.04752v3 Announce Type: replace-cross Abstract: Transformers consuming multi-channel scalar signals must embed $C$ simultaneous values into one $d_{\text{model}}$-dimensional vector per time step. We audit eight input encoders -- a shared-scalar baseline, per-channel linear projections, an...

📖 Read original article


240. The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport ​

Author: Peter Halmos, Boris Hanin
Published: 7/31/2026, 4:00:00 AM
Categories: math-ph, cs.AI, cs.LG, math.MP, physics.data-an

arXiv:2606.05217v4 Announce Type: replace-cross Abstract: We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family of Schr"odinger operators we call Score Hamiltonians, built from the learned score's quantum potentia...

📖 Read original article


241. TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation ​

Author: Eric Spencer, Arslan Bisharat, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG, cs.LO

arXiv:2606.06133v4 Announce Type: replace-cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produce TLA+ specifications that fail the TLC model checker for semantic reasons. Across 25 LLMs,...

📖 Read original article


242. Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent ​

Author: Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
Published: 7/31/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2606.06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning. We establish quantitative bounds show...

📖 Read original article


243. SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning for Open-Set ECG Classification ​

Author: Hongkyu Koh, Ikbeom Jang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.08037v2 Announce Type: replace-cross Abstract: Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive strategy for reducing annotation costs. In clinical settings, however, unlabeled pools frequently conta...

📖 Read original article


244. Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning ​

Author: Enhan Zhao, Wei Wu, Yuanrui Zhang, Xueliang Zhao, Di He
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.11719v2 Announce Type: replace-cross Abstract: Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-scale, statically curated datasets, where all training samples are treated uniformly regardless of the mo...

📖 Read original article


245. Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents ​

Author: Saehun Chun, Wonje Choi, Sera Choi, Sanghyun Ahn, Honguk Woo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.PL, cs.AI

arXiv:2606.13097v2 Announce Type: replace-cross Abstract: Code-writing large language models (CodeLLMs) generate executable code policies for embodied agents by translating natural language goals and environmental constraints into structured control programs. However, policy generation in open-domai...

📖 Read original article


246. Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers ​

Author: Xin Gao, Xingming Xu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.21848v2 Announce Type: replace-cross Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability limitations of the standard QKV attention mechanism. The Key-Value (KV) cache is a major bottlene...

📖 Read original article


Author: Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Min Zhang, Shaoping Ma, Qingyao Ai
Published: 7/31/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2606.24894v4 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit summarization-oriented metrics, using lexical or semantic sim...

📖 Read original article


248. MLVC: Multi-platform Learned Video Codec for Real-World Deployment ​

Author: Tanel P"arnamaa, Martin Lumiste, Ardi Loot, Evgenii Indenbom, Andrei Znobishchev, Ando Saabas
Published: 7/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG

arXiv:2606.28027v2 Announce Type: replace-cross Abstract: Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost. Existing quantization-based solutions fail to produce determinis...

📖 Read original article


249. Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control ​

Author: Brett Reynolds
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SE

arXiv:2607.01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic ...

📖 Read original article


250. Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring ​

Author: Dekun Yang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.01240v2 Announce Type: replace-cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench,...

📖 Read original article


251. Revealing Hidden Model Behaviors with Task-Specific Self-Reports ​

Author: Taras Kutsyk, Bartosz Zieli'nski
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.03640v2 Announce Type: replace-cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a li...

📖 Read original article


252. Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding ​

Author: Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.04383v4 Announce Type: replace-cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of the...

📖 Read original article


253. G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement ​

Author: Meng Du, Hongchang Chen, Ran Li, Junjie Zhang, Qi Ouyang, Shibo Zhang, Shuxin Liu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.04607v2 Announce Type: replace-cross Abstract: Rapid advances in AI video generation pose increasing security risks and call for reliable detectors with strong cross-domain generalization. Although existing methods perform well under in-domain evaluation, their performance degrades substa...

📖 Read original article


254. VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection ​

Author: Sharayu N. Deshmukh, Md Rashidunnabi, Nelton Tiago Gemo, Kurundkar G. D., Mahamune M. R., Nilesh K. Deshmukh
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.06254v2 Announce Type: replace-cross Abstract: Deepfake image detection is served by three fundamentally different paradigms - commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors - that are rarely evaluated under a common protocol, making direct comparison ...

📖 Read original article


255. Introducing Human-Centeredness in AI-Assisted Lexicography ​

Author: Antonio San Martin, Catherine Trekker
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.11808v3 Announce Type: replace-cross Abstract: This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role ...

📖 Read original article


256. Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning ​

Author: Zhen Yin, Wenkang An, Hao Wang, Keran You
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.14117v2 Announce Type: replace-cross Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous elements such as ...

📖 Read original article


257. Fantastic Adaptive Taxonomies and How to Use Them ​

Author: Mert Cemri, Andrei Cojocaru, Melissa Pan, Shu Liu, Shubham Agarwal, Alexander Krentsel, Jay Tang, Kannan Ramchandran, Joseph E. Gonzalez, Matei Zaharia, Alex Dimakis, Ion Stoica
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.16387v2 Announce Type: replace-cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory selection, prompt and workflow optimization, runtime monitoring) read these traces for feedback. Yet ...

📖 Read original article


258. Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer ​

Author: Jingjie Ning, Xiaochuan Li, Shanshan Zhong, Ji Zeng, Guolin Ke
Published: 7/31/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.SE

arXiv:2607.17100v2 Announce Type: replace-cross Abstract: Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal score cannot reveal which technical decision produced a gain or...

📖 Read original article


259. Towards an Automated Test of LLM Security Knowledge ​

Author: Shufan Chai, Liangliang Sun, Jessica Staddon
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.HC

arXiv:2607.18496v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on ident...

📖 Read original article


260. REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning ​

Author: Yunjie Chen, Xiaoxin Chen, Fang Wang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.19450v2 Announce Type: replace-cross Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domai...

📖 Read original article


261. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering ​

Author: Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Biao Tian, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2607.20253v3 Announce Type: replace-cross Abstract: In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation,...

📖 Read original article


262. Adaptive Multi-Horizon Reinforcement Learning ​

Author: Manoosh Samiei, Doina Precup, Paul Masset
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.20656v3 Announce Type: replace-cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a s...

📖 Read original article


263. On the Depth Scalability of Logic Gate Networks ​

Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO

arXiv:2607.21633v2 Announce Type: replace-cross Abstract: Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased depth. We identify two causes: optimization collapse and topology-induced degradation of output-specific c...

📖 Read original article


264. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​

Author: Jie Zhang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.21636v2 Announce Type: replace-cross Abstract: Synthetic tabular data is prized for preserving not just each column's marginal distribution but the dependencies between columns - structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fr...

📖 Read original article


265. LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning ​

Author: Chen Wang, Boming Kang, Qinghua Cui
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22777v2 Announce Type: replace-cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do not explicitly constrain the model to learn three-dime...

📖 Read original article


266. Directional Influence Function: Estimating Training Data Influence in Constrained Learning ​

Author: Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23388v3 Announce Type: replace-cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in...

📖 Read original article


267. An Unofficial FastLAS Tutorial: A Programmer's Guide ​

Author: Fabio Aurelio D'Asaro
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.LG

arXiv:2607.23557v2 Announce Type: replace-cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and it searches for a set of logic program rules (a hypothesis) that explains the examples. The...

📖 Read original article


268. Where Is the Cost of Third-Party API Routers in Agentic Software Development? ​

Author: Donghao Fu, Jingxin Li, Xue Jiang, Yihong Dong
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.23624v2 Announce Type: replace-cross Abstract: Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows, high-autonomy operation is widely adopted because it reduces interaction overhead. As a result, a thi...

📖 Read original article


269. Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents ​

Author: Aayush Kumar, Avik Dutta, Sumit Gulwani, Gustavo Soares, Advait Sarkar, Emerson Murphy-Hill
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SE

arXiv:2607.23670v2 Announce Type: replace-cross Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this...

📖 Read original article


270. Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature ​

Author: Tanjin He, Aikaterini Vriza, Logan Ward, Xu Huang, Yiming Chen, Anubhav Jain, Gerbrand Ceder, Rajeev S. Assary, Ian T. Foster, Maria K. Y. Chan
Published: 7/31/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.CL, cs.DL, cs.IR

arXiv:2607.23886v2 Announce Type: replace-cross Abstract: X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described ...

📖 Read original article


271. Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging ​

Author: Parth G. Dangi, Yogesh Kumar Meena
Published: 7/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.ET

arXiv:2607.24081v2 Announce Type: replace-cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in developing BMIs is expanding their usability and control, which can be achieved by accurately dec...

📖 Read original article


272. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation ​

Author: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24241v2 Announce Type: replace-cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal m...

📖 Read original article


273. Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist ​

Author: Mohsen Aliabadi, Keith Driscoll, Elliot Krop, Petar Sirkovic, Everett Sullivan, Elahe Vedadi
Published: 7/31/2026, 4:00:00 AM
Categories: math.NT, cs.AI, math.GR

arXiv:2607.24847v2 Announce Type: replace-cross Abstract: We introduce an extremal invariant associated with Chowla-type order conditions in finite groups. A nonempty subset $S$ of a finite group $G$ is called a Chowla set if every element of $S$ has order greater than $|S|$, and we write $C(G)$ for...

📖 Read original article


274. Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation ​

Author: Chandan Kumar Sah, Li Zhang, Xiaoli Lian
Published: 7/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG

arXiv:2607.24884v2 Announce Type: replace-cross Abstract: Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-code examples, repository context, and project-specific APIs may provide complementary inform...

📖 Read original article


275. Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation ​

Author: Cesare Spinoso-Di Piano, Verna Dankers, Marius Mosbach, Jackie Chi Kit Cheung
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.25094v2 Announce Type: replace-cross Abstract: Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize uns...

📖 Read original article


276. Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors ​

Author: Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni, Si Chen
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.25335v2 Announce Type: replace-cross Abstract: Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires ide...

📖 Read original article


277. I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models ​

Author: Yimao Guo, Zuomin Qu, Wei Lu
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.25522v2 Announce Type: replace-cross Abstract: The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain under...

📖 Read original article


278. The LAIA Dataset: Labelled Attention for Intelligent Automobiles ​

Author: A. Contreras, D. Porres, R. Abad, P. Cano, A. Levy, G. Villalonga, A. M. L'opez, A. Hern'andez-Sabat'e
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.SE

arXiv:2607.25570v2 Announce Type: replace-cross Abstract: The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large volumes of sensor data with ground-truth annotations. While modular architectures are widely used, end-t...

📖 Read original article


279. Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs ​

Author: Chandan Kumar Sah, Li Zhang, Xiaoli Lian
Published: 7/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.25600v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from black-box language m...

📖 Read original article


280. Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction ​

Author: Xinyi Hong, Pinjun Dong, Xinyang Yu, Binyan Jiang
Published: 7/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2607.25718v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small task-relevant subset from a library of thousands of tools before the agent acts, has therefore ...

📖 Read original article


281. Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs ​

Author: Fanfu Wei, Thibault Ehrhart, Rapha"el Troncy
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.25959v2 Announce Type: replace-cross Abstract: Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a practical question...

📖 Read original article


282. Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition ​

Author: Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy
Published: 7/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.25961v2 Announce Type: replace-cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognition of A/H at the video level is difficult, since the signal arises from disagreement across and within...

📖 Read original article