Skip to content

arXiv cs.AI - 2026-08-18 ​

647 items collected.


1. FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment ​

Author: Enrique Barba Roque, Lu'is Cruz
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.PF

arXiv:2608.14550v1 Announce Type: new Abstract: AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computation...

📖 Read original article


2. Large Language Models Show Metacognitive Sensitivity in Medical Reasoning ​

Author: Ahmad Nazzal
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14552v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical ...

📖 Read original article


3. The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning ​

Author: Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a ...

📖 Read original article


4. When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL ​

Author: Teoman Kaman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients ...

📖 Read original article


5. Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review ​

Author: Aasish Kumar Sharma, Dimitar Koysev, Christopher Anich, Roshni Kumari Ojha, Julian Kunkel
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14562v1 Announce Type: new Abstract: AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that ma...

📖 Read original article


6. Position: AI Lock-In Is in Progress, and We Must Be Prepared ​

Author: Jaeho Kim, Seokhyun Lee, Jieun Lee, Changhee Lee
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14565v1 Announce Type: new Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an eq...

📖 Read original article


7. Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture ​

Author: Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14566v1 Announce Type: new Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether ...

📖 Read original article


8. From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change ​

Author: Yuri Almeida, Arthur Casals
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2608.14567v1 Announce Type: new Abstract: This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation. Seeded by Doyle and London's foundational 1980 taxonomy, we trace the evolution of belief revision f...

📖 Read original article


9. Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws ​

Author: Azmine Toushik Wasi, Mst Rafia Islam, Mahfuz Ahmed Anik, Taki Hasan Rafi, Md Manjurul Ahsan, Dong-Kyu Chae
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.GT, cs.HC

arXiv:2608.14568v1 Announce Type: new Abstract: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary ...

📖 Read original article


10. Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration ​

Author: Shufeng Kong, Xiaochuan Zhang, Caihua Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14569v1 Announce Type: new Abstract: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confi...

📖 Read original article


11. Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System ​

Author: Shaochen Zhong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.DL

arXiv:2608.14571v1 Announce Type: new Abstract: With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presenc...

📖 Read original article


12. Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study ​

Author: Yixuan He, Jinni Su, Yun Kang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, stat.AP

arXiv:2608.14578v1 Announce Type: new Abstract: Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characteristics, longitudinal trajectories, and relational context remains unclear. Using data from approximately 11,860 part...

📖 Read original article


13. SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization ​

Author: Rui Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14579v1 Announce Type: new Abstract: Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods of...

📖 Read original article


14. OGX: An Open-Source, Vendor-Neutral Generative AI Application Server ​

Author: Francisco Javier Arceo, S'ebastien Han, Matthew Farrellee, Charlie Doern, Yuan Tang, Derek Higgins, Varsha Prasad Narsing, Gordon Sim, Sumanth Kamenani, Ben Browning, Raghotham Murthy
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.14580v1 Announce Type: new Abstract: OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers. Developers building agentic AI applications--such as retriev...

📖 Read original article


15. Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry ​

Author: Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14585v1 Announce Type: new Abstract: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities o...

📖 Read original article


16. An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case ​

Author: Nicolas Turenne, Youcef Sklab, Eric Chenin, Jean-Daniel Zucker
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14587v1 Announce Type: new Abstract: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementar...

📖 Read original article


17. The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines ​

Author: Prabhjot Singh, Bhushan Pawar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2608.14588v1 Announce Type: new Abstract: Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: r...

📖 Read original article


18. Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement ​

Author: Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans the...

📖 Read original article


19. Position: Medical AI Neglects Real Treatment Outcomes ​

Author: Shiva Kaul, Anjum Khurshid
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially t...

📖 Read original article


Author: Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) ca...

📖 Read original article


21. Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP ​

Author: Wael Albayaydh, Rui Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols spe...

📖 Read original article


22. Large Language Models and their Awareness of Mechanics and Spatial Geometry ​

Author: Johannes Gerstmayr, Sebastian Weyrer, Tobias M"oltner, Peter Manzl, Michael Pieber
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systemat...

📖 Read original article


23. A Human-Centred Approach to Benchmarking LLMs for Parenting Advice ​

Author: Yunke Zhao, Isobel Voysey, Alastair van Heerden, Rob Hughes, Jun Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information qualit...

📖 Read original article


24. Learning Agent Execution for KV-Cache Management in Agentic Serving ​

Author: Rui Zhang, Chaeeun Kim, Shaoting Feng, Kuntai Du, Yuhan Liu, Yi Zhong, Cheng-Wei Ching, Junchen Jiang, Liting Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of...

📖 Read original article


25. Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study ​

Author: Amelia Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14631v1 Announce Type: new Abstract: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic ch...

📖 Read original article


26. Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks ​

Author: Kiran N. Kumar, Santhosh K. Saminathan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14641v1 Announce Type: new Abstract: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hy...

📖 Read original article


27. Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance ​

Author: Anuridhi Gupta, Samara Mansoor, Hemant Purohit
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers...

📖 Read original article


28. When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation ​

Author: Pranav Rakasi, Maanas Lalwani, Arnav Srivastava, Arya Palanivel, Tinuade Adeleke, Ruizhe Li, Sean Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE

arXiv:2608.14659v1 Announce Type: new Abstract: Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can...

📖 Read original article


29. Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring ​

Author: Dhiraj Neupane, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14666v1 Announce Type: new Abstract: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling faults, where the physical relationship between sensor groups breaks while margi...

📖 Read original article


30. Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems ​

Author: Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.14667v1 Announce Type: new Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork...

📖 Read original article


31. Beyond Correctness: Toward Automated Novelty Verification with Lean 4 ​

Author: Ayrton Porto
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14669v1 Announce Type: new Abstract: Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that rec...

📖 Read original article


32. Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition ​

Author: Miko{\l}aj Sienicki, Krzysztof Sienicki
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, quant-ph

arXiv:2608.14673v1 Announce Type: new Abstract: Chapter 6 of OpenAI's Ten Advances in Mathematics and Theoretical Computer Science claims an exponential parallel-repetition theorem for all finite two-player, one-round entangled games. Early in the proof, the chapter uses a quantitative greedy cond...

📖 Read original article


33. When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry ​

Author: Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fa...

📖 Read original article


34. A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks ​

Author: Naveed Khan, Besan Al Sbeihi, Maryam Alshehhi, Nasir Saeed
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.NI, eess.SP

arXiv:2608.14694v1 Announce Type: new Abstract: Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning ...

📖 Read original article


35. Synchronized Logit Steering: Real-world Steganography ​

Author: Andrew Rufail, Aadi Dash, Onir Narahari, Ethan Mui, Mahi Gajare, Prakhar Tiwari, Shrija Makapothula, Nick Cui
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed...

📖 Read original article


36. Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems ​

Author: John Knowlton, Aritra Guha, Risto Miikkulainen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14707v1 Announce Type: new Abstract: As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge. Existing orchestration strategies typically rely on fixed interaction patterns and often lack me...

📖 Read original article


37. Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation ​

Author: Jiajun Jiang, Sharon Zheng, Natan Vidra, Spurthi Setty
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attempts, conflati...

📖 Read original article


38. Advanced modelling and data analytics in aviation ​

Author: Aziida Nanyonga
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures. Despite the vast accumulation of aviation safety data over time, its full potential in predicting and pr...

📖 Read original article


39. Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs ​

Author: Hadi Fadlallah
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2608.14765v1 Announce Type: new Abstract: Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evi...

📖 Read original article


40. From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving ​

Author: Dipankar Sarkar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.LO, cs.PL, cs.SC, math.OC

arXiv:2608.14771v1 Announce Type: new Abstract: Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful tran...

📖 Read original article


41. Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations ​

Author: Qian Yin, Xinwei Wang, Guohua Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14789v1 Announce Type: new Abstract: Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically dispersed ground targets, but emergency requests may arrive after committed routine-plan execution has begun. The resulting dynamic emergency obse...

📖 Read original article


42. CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs ​

Author: Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narr...

📖 Read original article


43. Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous ​

Author: Adam M. Oberman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.GT

arXiv:2608.14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fr...

📖 Read original article


44. Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning ​

Author: Augusto Bernardo Pissarra, Victor Lorena de Farias Souza
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is c...

📖 Read original article


45. Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking ​

Author: Yepeng Huang, Jiawen Zhang, Michelle Dai, Xiaorui Su, Shanghua Gao, Zi Wang, Marinka Zitnik
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.14808v1 Announce Type: new Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. We formalize multi-turn informat...

📖 Read original article


46. MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment ​

Author: Tony Tu, Sayan Chakraborty, Ruomeng Xu, Tony Qin, Austin Tian
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a suppo...

📖 Read original article


47. What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering ​

Author: Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee, Carl Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14841v1 Announce Type: new Abstract: Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. In our setting, the bottleneck shifts from retrieval recall to reranker-s...

📖 Read original article


48. Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning ​

Author: Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, Chris Piech
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.14851v1 Announce Type: new Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to t...

📖 Read original article


49. JarvisBench: Always-on Intelligence Between Humans and Agents ​

Author: Chen Chen, Zhehuai Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate access to an agent while work continues in the background, whereas agents may...

📖 Read original article


50. Personalized Auto-Research: Towards a True AI Co-Scientist ​

Author: Bo Ni, Franck Dernoncourt, Hongjie Chen, Yu Wang, Nesreen K. Ahmed, Zhengzhong Tu, Tyler Derr, Ryan A. Rossi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.14881v1 Announce Type: new Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnos...

📖 Read original article


51. Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence ​

Author: Fabricio F Costa
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections b...

📖 Read original article


52. LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks ​

Author: Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang, Vikram Vasudevan, Huihuo Zheng, Venkatram Vishwanath, Rajeev Thakur
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while...

📖 Read original article


53. Small Models Scout Bottleneck Order for Large-Model Data Control ​

Author: Seungmin Choi, Jiwon Sung, Muhammad Umer, Abhiram Rao Gorle, Guijin Son, Youngjae Yu, John M. Cioffi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14936v1 Announce Type: new Abstract: Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate ...

📖 Read original article


54. When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation ​

Author: Avyay M. Casheekar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.14940v1 Announce Type: new Abstract: Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would require two conditions that the endpoint does not itself necessarily establis...

📖 Read original article


55. Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid ​

Author: Hironobu Nakasuji
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14943v1 Announce Type: new Abstract: Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full, Skill Block, Reference, and Hybrid. Across SearchQA, SpreadsheetBench, ALFWorld, ScienceWorld, and SynthProc, we ...

📖 Read original article


56. Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL ​

Author: Qizhen Lan, Xi Xiao, Xiangchen Guan, Mengchen Fan, Moule Lin, Jung Im Choi, Lijing Zhu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.14945v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whethe...

📖 Read original article


57. RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder ​

Author: Yi Xiao, Harshit Sharma, Dessa Bergen-Cico, Asif Salekin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14947v1 Announce Type: new Abstract: Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventions for individuals with opioid use disorder (OUD). This challenge is especially pronounced under subject-independ...

📖 Read original article


58. T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework ​

Author: Zahra Fazel, Sunanda Gamage, Shayan Shirahmad Gale Bagi, Amir H. Ashouri, Tomasz S. Czajkowski, Bryan Chan, Reza Azimi, Yaoqing Gao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.PF, cs.PL

arXiv:2608.14953v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since emerged as one of the most fundamental tasks for LLMs to perform; however, at present...

📖 Read original article


59. Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design ​

Author: Hossein Z. Saghazadeh, Yonas Ayalew, Reza Ahmari, Parham Kebria, Abdollah Homaifar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2608.14974v1 Announce Type: new Abstract: This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that links vertiport siting, fleet simulation, and door-to-door travel-time feasibility. Demand is estimated from commuter and passenger activity data, ...

📖 Read original article


60. Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 ​

Author: Justin Bronder
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.14992v1 Announce Type: new Abstract: Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model give...

📖 Read original article


61. S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices ​

Author: Haochen Huang, Shengxuan Qiu, Meng Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15018v1 Announce Type: new Abstract: Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively comb...

📖 Read original article


62. Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form ​

Author: Parsa Mazaheri
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.15022v1 Announce Type: new Abstract: Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an a...

📖 Read original article


63. LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning ​

Author: Changhong He, Jinda Gao, Xinkuan Liu, Le Zhang, Xizi Luo, Yu Mei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15041v1 Announce Type: new Abstract: Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-ba...

📖 Read original article


64. SCOPE: Score-Isolated Agentic Optimization for Video World Models ​

Author: Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15043v1 Announce Type: new Abstract: Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation problem: prompts, samplers, verifiers, and selectors may evolve together, making it di...

📖 Read original article


65. Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research ​

Author: Zi'an Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, math.OC

arXiv:2608.15052v1 Announce Type: new Abstract: Andy is an autonomous mathematical research agent that solves and verifies submitted problems, formulates new research problems, and constructs rigorous proofs. It separates proof generation from correctness evaluation and supports knowledge acquisitio...

📖 Read original article


66. TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning ​

Author: David Yoon Suk Kang, JungHyun Kim, Juhyun Jeon, Sang-Wook Kim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15055v1 Announce Type: new Abstract: Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on ...

📖 Read original article


67. GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG ​

Author: Zafar Ali, Asad Khan, Aalia Malik, Pavlos Kefalas
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15056v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs, which can introduce noisy evidence, weaken multi-hop reasoning, and increase unsupported generation. We present G...

📖 Read original article


68. LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents ​

Author: Yuefeng Zou, Yichen Lu, Jingxiao Yang, Bingtao Fu, Gaoyang Zhang, Xiongfei Bai, Tian Chen, Xiang Qi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15064v1 Announce Type: new Abstract: Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also ...

📖 Read original article


69. Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning ​

Author: Chanhee Park, Sungbin Han, Jeongho Yoon, Seongtae Hong, Heuiseok Lim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15065v1 Announce Type: new Abstract: Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de f...

📖 Read original article


70. Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents ​

Author: Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing Lu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in reali...

📖 Read original article


71. Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems ​

Author: Aashna Sofat, Balwinder Sodhi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15082v1 Announce Type: new Abstract: Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not decision-making agents: thresholds trigger alerts, but nothing relates violations to cumulative product degradation or converts degradation signa...

📖 Read original article


72. StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling ​

Author: Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scal...

📖 Read original article


73. Validation-Frontier Representation Selection under Constrained Observation ​

Author: Wesley Shu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15095v1 Announce Type: new Abstract: AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state repre...

📖 Read original article


74. Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation ​

Author: Wesley Shu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15101v1 Announce Type: new Abstract: Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy changes the system it enters: actors adapt, enforcement capacity shifts, burdens move, and new equilibria form arou...

📖 Read original article


75. Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents ​

Author: Jianxing Zhao, Mao Guan, Dongyu Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15109v1 Announce Type: new Abstract: Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three compleme...

📖 Read original article


76. Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads ​

Author: Anubhab Banerjee
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG

arXiv:2608.15117v1 Announce Type: new Abstract: Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirically within a ...

📖 Read original article


77. Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark ​

Author: Wesley Shu, Peng Wei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.15131v1 Announce Type: new Abstract: Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requirements, appeal processes, and access policies. These interventions are rarely absorbed passively. Creators, sell...

📖 Read original article


78. ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits ​

Author: Zhiqiang He, Zhi Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15138v1 Announce Type: new Abstract: Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching or beating hand-built designs. But either way, the design fits only the world visible at its birth, and fails on...

📖 Read original article


79. Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy ​

Author: Tias Guns, Ignace Bleukx, Hendrik Bierlee, Jo Devriendt, Emilio Gamba, Orestis Lomis, Wout Piessens, Thomas Sergeys, Dimos Tsouros, Wout Vanroose, H'el`ene Verhaeghe
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15143v1 Announce Type: new Abstract: Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and decision variables, and a generic solver is used to find a solution. Several constra...

📖 Read original article


80. ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models ​

Author: Xinmei Huang, Jie Song, Peng Li, Fuxin Jiang, Jing Zhang, Tieying Zhang, Jianjun Chen, Chenming Liu, Tao Yang, Maoyin Liu, Wenda Li, Hong Chen, Cuiping Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality traini...

📖 Read original article


81. Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World ​

Author: Jiang Jiang (Persagy Science and Technology Co., Beijing, China), Yifu Sun (Persagy Science and Technology Co., Beijing, China), Qi Shen (Persagy Science and Technology Co., Beijing, China)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.15147v1 Announce Type: new Abstract: Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is ...

📖 Read original article


82. SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion ​

Author: Yu He, Weikai Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15165v1 Announce Type: new Abstract: Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgm...

📖 Read original article


83. LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures ​

Author: Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogang Xie, Dan Pei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.15242v1 Announce Type: new Abstract: When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize t...

📖 Read original article


84. Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts ​

Author: Diego Mardian, Frank Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.15254v1 Announce Type: new Abstract: Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question le...

📖 Read original article


85. Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark ​

Author: Vasiliy Seibert
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15255v1 Announce Type: new Abstract: Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain. Despite advancements in automated domain modeling, the absence of standardized benchmarks has hindered the c...

📖 Read original article


86. Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication ​

Author: Lin Yin, Tiejun Lv, Weicai Li, Xi Yu, Xiaoyu He
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.15256v1 Announce Type: new Abstract: Collaborative training in distributed semantic communication (DSC) networks typically relies on decentralized federated learning (DFL). However, pushing topology-agnostic aggregation into heterogeneous, multi-task environments creates a fundamental bot...

📖 Read original article


87. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? ​

Author: Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15265v1 Announce Type: new Abstract: Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user in...

📖 Read original article


88. $D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction ​

Author: Ninghan Fan, Qi Liu, Xunuo Zhu, Yukai Sun, Luyuan Chen, Xuheng Zhou, Yuetian Du, Ming Kong, Xiaojun Zhu, Jie Liu, Zhan Zhou, Qiang Zhu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15288v1 Announce Type: new Abstract: Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Existing approaches, however, typically predict an entire expression profile as a whole, leaving the order in which in...

📖 Read original article


89. ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning ​

Author: Ziyue Yang, Chaolin Xu, Yijing Wang, Tiankai Gu, Hui Yang, Yanhong Lin, Kaiyuan Liu, Fei Xiao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15291v1 Announce Type: new Abstract: Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking know...

📖 Read original article


90. Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis ​

Author: Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15303v1 Announce Type: new Abstract: Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase primitive consisting of...

📖 Read original article


91. Understanding Cognition-Induced Risks in Agentic AI Systems ​

Author: Guanchu Wang, Qinuo Li, Mengnan Du, Xia Hu, Bowen Zhou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15304v1 Announce Type: new Abstract: Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that re...

📖 Read original article


92. Physiological World Models for Human State Transitions ​

Author: Chongyang Zhang, Rendong Wang, Hao Zheng, Hanwen Zhang, Yang Liu, Xiaolong Wei, Bin Chong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15309v1 Announce Type: new Abstract: Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate r...

📖 Read original article


93. MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning ​

Author: Ankita Sharma, Bahar Farahani, Sanaz Rahimi Moosavi, Amir Rrahmani, Farshad Firouzi, Krishnendu Chakrabarty
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15311v1 Announce Type: new Abstract: Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing. Recent Mixture-of-Experts (MoE) LLMs are particularly attractive for federated learning because the...

📖 Read original article


94. Physics-informed VAE-EVT for Tail Aware Radio Map Prediction ​

Author: Amanda Sheron Gamage, Niloofar Mehrnia, James Gross
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15314v1 Announce Type: new Abstract: Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold. In this context, an outage refers to instances in which SNR falls below a specifi...

📖 Read original article


95. The Benchmark Trap: Structures of Power and Injustice in AI Evaluations ​

Author: Jason Branford, Angelie Kraft
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.15326v1 Announce Type: new Abstract: Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of ...

📖 Read original article


96. A concentration result for multilayer feedforward neural networks ​

Author: Vera Koponen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, math.LO, math.PR

arXiv:2608.15335v1 Announce Type: new Abstract: We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last laye...

📖 Read original article


97. Incoherent by Design? On the Moral Self-Consistency of LLMs ​

Author: Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15354v1 Announce Type: new Abstract: LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or refra...

📖 Read original article


98. UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms ​

Author: Phillip Jiang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.15372v1 Announce Type: new Abstract: We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.S. Air Force SBIR solicitation. We ...

📖 Read original article


99. FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA ​

Author: Juseok Jeon, Ramy E. Ali, Doyun Kwon, Myungbeom Her, Jinhwi Kim, Jinhyun So
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.15381v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor-wise ag...

📖 Read original article


100. Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot ​

Author: Ummara Mumtaz, Aimen Noor, Awais Ahmed
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.15382v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a repr...

📖 Read original article


101. Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL ​

Author: Changruo Zhao, Zujun Peng, Yu Tian, Yuting Liu, Yiyun Su, Huiying Zhu, Luyan Zhang, Heming Zeng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15389v1 Announce Type: new Abstract: LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the metrics authors themselves report ...

📖 Read original article


102. TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions ​

Author: Md Fazley Rafy
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.15391v1 Announce Type: new Abstract: Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but syntactic validity does not imply physical admissibility. This paper presents TwinGridShield, a model-independent runt...

📖 Read original article


103. Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish ​

Author: Madhusudhanan G
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15392v1 Announce Type: new Abstract: Chain-of-thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains uncertain. In a small case study of eight manually verified synthetic scenarios, one model, one annotator, and one d...

📖 Read original article


104. Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks ​

Author: Azmeer Akhtar, Md Fazley Rafy, Anurag K. Srivastava
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SY, eess.SY

arXiv:2608.15396v1 Announce Type: new Abstract: Battery energy storage systems (BESS) are increasingly used in distribution networks for voltage regulation and demand response, which increases the volume and complexity of operational telemetry available to grid operators. This paper presents an AI-e...

📖 Read original article


105. Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs ​

Author: Charles Courchaine, Ricky J. Sethi, Hefei Qiu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.15400v1 Announce Type: new Abstract: Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizing when problems exceed their expertise; such limitations inevitably undermine reliability and trust in LLMs. In ...

📖 Read original article


106. A survey of AI-generated voices and their detection ​

Author: Chengzhe Sun, Tianle Yang, Siwei Lyu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15411v1 Announce Type: new Abstract: The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual assistants and creative applications, but they also enable harmful uses, including ...

📖 Read original article


107. Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs ​

Author: Tadd Mao, Tianjun Zhong, Dhruva Arekar, Yuming Feng, One An, Jiani Huang, Xujie Si, Ziyang Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15432v1 Announce Type: new Abstract: In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated proof search can produce a formal proof that compiles, the generated proof does not necessarily reflect how the n...

📖 Read original article


108. OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks ​

Author: Farbod Tavakkoli, Roderic Paulk, Jorden Terrazas, Kenneth Church, Mark Austin, Louis Powell, Gregory Diamos, Lina Bariah, Syed Ali Raza Zaidi, Maryam Hafeez, Ali Maatouk, Imtiaz Karim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.NI

arXiv:2608.15436v1 Announce Type: new Abstract: Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource with derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, plus 30 ...

📖 Read original article


109. Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization ​

Author: Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15445v1 Announce Type: new Abstract: When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoint accuracy on the training distribution cannot tell the two apart, becau...

📖 Read original article


110. Mental Model Management: An Operator-Based Framework for LLM Memory ​

Author: Oliver Kramer
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.NE

arXiv:2608.15451v1 Announce Type: new Abstract: Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowledge is represented ...

📖 Read original article


111. Dynamic Multi-Byte Prediction With Hierarchical Language Models ​

Author: Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15454v1 Announce Type: new Abstract: Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, w...

📖 Read original article


112. A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution ​

Author: Jie Wei, Yue Liu, Xiaochuan Tang, Biao Cai, Xiangtao Li, Yanmei Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15488v1 Announce Type: new Abstract: Effective public event forecasting is essential for intelligent service systems, enabling proactive risk management, adaptive resource allocation, and timely decision-making. In many real-world scenarios, the evolution of public events is driven by dyn...

📖 Read original article


113. EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints ​

Author: Ao Zhou, Bo Dai, Le Yu, Xingyu Liu, Zeyu Hao, Lingkun Long, Chunming Hu, Jianlei Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.15502v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute...

📖 Read original article


114. Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation ​

Author: Qinghao Fu, Yarong Wang, Shunlei Ning, Yilin Wang, Shunwen Bai, Xinda Wang, Jiaotuan Wang, Yinan Nie, Wei Zhou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15510v1 Announce Type: new Abstract: Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods either train visual and coding abilities separately, or fine-tune on chart-to-code...

📖 Read original article


115. From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation ​

Author: Albert Sadowski, Jaros{\l}aw A. Chudziak
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2608.15536v1 Announce Type: new Abstract: In value-based argumentation, an audience's ordering of values decides which attacks succeed as defeats. In many settings the deciding factor is not the audience but the circumstances: the same attack may succeed at one procedural stage, or under one r...

📖 Read original article


Author: Danial Yazdani, Mohammad Nabi Omidvar, Yuan Sun, Maksud Ibrahimov, Xiaodong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.NE

arXiv:2608.15546v1 Announce Type: new Abstract: Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organization and component interactions. We present ATLAS, an embedding-guided quality-diversity framework for scaffold-...

📖 Read original article


117. Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling ​

Author: Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label-free alternatives are unreliab...

📖 Read original article


118. From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM ​

Author: Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.15580v1 Announce Type: new Abstract: Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a un...

📖 Read original article


119. Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback ​

Author: Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing approaches address ...

📖 Read original article


120. When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction ​

Author: Feiyang Ren, Shengtao Wen, Lingbing Guo, Yu Tian, Yuanning Cui, Xiang Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15592v1 Announce Type: new Abstract: Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduc...

📖 Read original article


121. TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation ​

Author: Md Messal Monem Miah, Adrita Anika, Zhiyuan Yu, Ruihong Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15594v1 Announce Type: new Abstract: Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving ma...

📖 Read original article


122. VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation ​

Author: Mingyu Yuan, Shengtao Wen, Lingbing Guo, Zhen Bi, Xiang Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15600v1 Announce Type: new Abstract: The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extractio...

📖 Read original article


123. Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds ​

Author: Keito Inoshita
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15619v1 Announce Type: new Abstract: Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline. Our aim is not to pin this ceiling to a single number, but to quantify how far it depends on finite annotatio...

📖 Read original article


124. Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts ​

Author: Seungyeol Baek, Yoonbyung Chai, Yonghyeon Lee, Sungjoon Choi, Sungho Suh
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.15621v1 Announce Type: new Abstract: Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires reattaching inertial measurement units (IMUs) across sessions. In multi-IMU settings, this can induce independent ...

📖 Read original article


125. Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict ​

Author: Elisa Cavatorta, Antonio Rago
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15634v1 Announce Type: new Abstract: How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested narratives? Such acceptability is mediated not only by the clauses that agreements include or exclude, but crucial...

📖 Read original article


126. A Responsible Artificial Intelligence Framework for Groundwater Modeling ​

Author: Chong Chen, Yulu Zhang, Qingxi Guo, Yihan Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15657v1 Announce Type: new Abstract: The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsible AI systems in a manner aligned with human values and ethical standards. Compared to fields like healthcare, e...

📖 Read original article


127. THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts ​

Author: Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which ...

📖 Read original article


128. Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment ​

Author: Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.15693v1 Announce Type: new Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real ...

📖 Read original article


129. Adaptive Mixing of Policies from Searching and Policies from Learning ​

Author: Gavin B. Rens
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15700v1 Announce Type: new Abstract: Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can take exceedingly long. Objectives: Rather than perform search to the same depth every time (typically at a fixed per...

📖 Read original article


130. HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation ​

Author: XinQi Wang, Jinwei Xiao, Sijia Cui, Hongming Zhang, Yanna Wang, Qingyang Zhang, Bo Xu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15703v1 Announce Type: new Abstract: Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered over time. As interactions accumulate, detailed execution traces and intermediate outputs dominate the context, m...

📖 Read original article


131. PLeDO: Pain Level Detection for Osteoarthritis from EMR Data ​

Author: Yuhao Chen, Jiahao Cai, Nafiz Sadman, Farhana Zulkernine, John Queenan, David Barber
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.IR, cs.LG

arXiv:2608.15719v1 Announce Type: new Abstract: Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain se...

📖 Read original article


132. Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps ​

Author: Yonghe Sun, Zhenjia Liu, Hua Liao, Wenjia Xu, Nai Yang, Weihua Dong, Zhiwei Wei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15736v1 Announce Type: new Abstract: Foundation models (FMs) increasingly support multimodal and geospatial reasoning, yet it remains unclear whether cartographic principles designed for human perception are equally effective for machines. Focusing on sequential choropleth maps, we examin...

📖 Read original article


133. Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign ​

Author: Benjamin Icard, Elouan Vuichard, Louis Lefebvre, Lila Sainero, Thomas Girault, Alice Breton, Tanguy Launay, Gauvain Bourgne, Morgane Casanova, Guillaume Gadek, Victor Kl"otzer, Michel Le Nouy, Guillaume Gravier, Jean-Gabriel Ganascia, Paul 'Egr'e
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.15746v1 Announce Type: new Abstract: We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of 2,646 propagandist French articles from the Storm-1516/CopyCop campaign disclosed by VIGINUM and INSIKT GROUP in ...

📖 Read original article


134. Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents ​

Author: Meiling Tao, Yiling Tao, Peng Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15755v1 Announce Type: new Abstract: User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, missing information, and execution constraints. Existing context-management methods improve the use of past interact...

📖 Read original article


135. Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration ​

Author: Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15772v1 Announce Type: new Abstract: When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which y...

📖 Read original article


136. KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving ​

Author: Minsoo Cheong, Woosang Lim, Vincent-Daniel Yun, Sungjoo Yoo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.15797v1 Announce Type: new Abstract: KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where...

📖 Read original article


137. Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State ​

Author: Fanzhe Wei, Li Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15810v1 Announce Type: new Abstract: Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no soundness statement, and certified approaches budget request-level risk by a union bound over a pre-declared even...

📖 Read original article


138. RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning ​

Author: Shihong Huang, Shengjie Wang, Hong Ma, Zhou Xu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15817v1 Announce Type: new Abstract: The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradi...

📖 Read original article


139. The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale ​

Author: Parviz Shariff
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15832v1 Announce Type: new Abstract: As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action is insufficient: the system must also determine whether the action is authorised in its context. This paper intr...

📖 Read original article


140. Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs ​

Author: Marius Dragic, Ruben Ifrah, Alexandre Rio
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB

arXiv:2608.15834v1 Announce Type: new Abstract: Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, cat, grep). A knowledge graph admits the same interface: listing neighbours, reading node content and searching des...

📖 Read original article


141. RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration ​

Author: Yishun Wang, Wenjin Yi, Wenkai Li, Zongwei Li, Xiaoqi Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15857v1 Announce Type: new Abstract: Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly referred to as Gas, scale with the computational complexity of their functions. Smart contracts on Ethereum incur...

📖 Read original article


142. CoupVisor: Strategy Optimization by Round and Challenge Decision Support ​

Author: Cris Huynh
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG

arXiv:2608.15868v1 Announce Type: new Abstract: This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a player should do on each turn, and when a player should challenge an opponent's claim. The system is built around a s...

📖 Read original article


143. Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation ​

Author: Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Hong Yan, Shujian Bu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15877v1 Announce Type: new Abstract: Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politics} steer sub...

📖 Read original article


144. Bounded Agents: Delegation Security for Multi-Agent AI Systems ​

Author: Xabier Muruaga
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.15888v1 Announce Type: new Abstract: LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Wi...

📖 Read original article


145. Breaking and Defending LLM-Powered Social Media Bot Detection Systems ​

Author: Nof Orenstein, Yoni Birman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15893v1 Announce Type: new Abstract: The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but att...

📖 Read original article


146. Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning ​

Author: \v{S}imon Sukup, Ariyan Bighashdel, Pavol Jancura
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15929v1 Announce Type: new Abstract: Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations ...

📖 Read original article


147. UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations ​

Author: Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instr...

📖 Read original article


148. Augmenting Text to Increase Translation Difficulty ​

Author: William Kalikman, \v{S}imon Sukup, Michal Te\v{s}nar, Vil'em Zouhar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15932v1 Announce Type: new Abstract: As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by ...

📖 Read original article


149. Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces ​

Author: Shrey Shah, Levent Ozgur
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15956v1 Announce Type: new Abstract: Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces can adapt a deployed dense retriever to changing workflow distributions without new relevance labels, synthetic q...

📖 Read original article


150. Solvable Sokoban Without a Solver via Diffusion ​

Author: Sina Baghal
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG

arXiv:2608.15958v1 Announce Type: new Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no short certificate to check. Solvability is also a fragile property, since even a single misplaced wall can silently ...

📖 Read original article


151. ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction ​

Author: Eric Xie, Wenqian Ye, Aidong Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, t...

📖 Read original article


152. MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment ​

Author: Stefanos Gkikas, Eric Nichols, Christian Arzate Cruz, Randy Gomez
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15999v1 Announce Type: new Abstract: Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion. This paper presents MUPA\textsuperscript{2}E, a unifie...

📖 Read original article


153. Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency ​

Author: Parsa Mazaheri, Kasra Mazaheri
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench tra...

📖 Read original article


154. Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance ​

Author: Bowen Li, Guojun Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16055v1 Announce Type: new Abstract: Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench, a benchmark for the governability of financial agents---whether they escalate when obligated, abstain when requi...

📖 Read original article


155. Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics ​

Author: Conrad Ainslie, Pedram Hassanzadeh, Michael W. Mahoney, Ashesh Chattopadhyay
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, nlin.CD, physics.comp-ph

arXiv:2608.16084v1 Announce Type: new Abstract: Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instability and error growth remain poorly understood, leading to ad-hoc solutions. Here, we develop an eigenanalysis frame...

📖 Read original article


156. Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling ​

Author: Wengan He, Yongsheng Luo, Lihong Jiang, Wenhui Xu, Yu Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.16094v1 Announce Type: new Abstract: Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provides a basis for mechanistic interpretation. Recent advances in deep learning have transformed the field from mult...

📖 Read original article


157. Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity ​

Author: Silv`ere Gangloff
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, math.HO

arXiv:2608.16118v1 Announce Type: new Abstract: How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - re...

📖 Read original article


158. When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification ​

Author: Diyorbek Musaev
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a...

📖 Read original article


159. FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection ​

Author: Junxuan Li, Zhiqi Chen, Yuzhou Liu, Peng Zhang, Huaxiao Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16148v1 Announce Type: new Abstract: Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling...

📖 Read original article


160. TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents ​

Author: Huan Zhang, Mingju Chen, Dongxu Zhou, Can Lv, Heng Chang, Sen Cui, Faguo Wu, Shiji Zhou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur ann...

📖 Read original article


161. Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain ​

Author: Rocky Liu, Tengyu Liu, Baoxiong Jia, Fangwei Zhong, Xinyi Tong, Hongzhao Xie, Siyuan Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.16164v1 Announce Type: new Abstract: Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since unstructured terrain lacks explicit difficulty ordering for curriculum design, existing methods resort to heuristic ...

📖 Read original article


162. Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication ​

Author: Jia Guo, Xiaohan Zhao, Changwang Liu, Shuqing He, Chenyang Zhang, Bingchuan Zhao, Jinqi Zhu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver. However, existing token-selection criteria based on local uncertainty, importance, or diversity...

📖 Read original article


163. Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior ​

Author: Yifan Lu, Xiaopeng Yuan, Haohan Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.16196v1 Announce Type: new Abstract: Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plau...

📖 Read original article


164. Competing at Every Price Point with Agentic Evolution over a Menu of LLMs ​

Author: Andrew Borthwick
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16207v1 Announce Type: new Abstract: Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy elsewhere. Th...

📖 Read original article


165. BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics ​

Author: Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16211v1 Announce Type: new Abstract: Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structure...

📖 Read original article


166. Process-Constituted Intelligence: A Shared Criterion for Humans and Machines ​

Author: Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.ET

arXiv:2608.16213v1 Announce Type: new Abstract: Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \textit{traces} (textual and visual residues of human cognitive processes), reproducing sam...

📖 Read original article


167. AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment ​

Author: Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in inter...

📖 Read original article


168. DriveCache: Action-Aware Caching for Driving World Model Inference ​

Author: Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.16354v1 Announce Type: new Abstract: Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Diffusion-based driving generators repeatedly evaluate large backbones acr...

📖 Read original article


169. What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics ​

Author: Shuyu Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16370v1 Announce Type: new Abstract: Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's interaction cost by forcing it to reacquire dropped state while leaving completion statistically unchanged. We introduc...

📖 Read original article


170. AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems ​

Author: Zhenhang Nie (iFLYTEK Co., Ltd., Hefei, China), Gui Zheng (iFLYTEK Co., Ltd., Hefei, China), Xudong Sun (iFLYTEK Co., Ltd., Hefei, China), Tailong Zhu (iFLYTEK Co., Ltd., Hefei, China), Bin Zhang (iFLYTEK Co., Ltd., Hefei, China)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16381v1 Announce Type: new Abstract: Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages. We introduce a unified execution model that maintains a work item's persistent identi...

📖 Read original article


171. Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152 ​

Author: Vahid Zolfaghari, Nenad Petrovic, Andr'E Schamschurko, Alois Knoll
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.16394v1 Announce Type: new Abstract: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Gene...

📖 Read original article


172. A Policy Algebra for Trust-Preserving Agentic AI Execution ​

Author: Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar, Bhavesh Gadhe
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16402v1 Announce Type: new Abstract: Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a goal. Enterprise execution requires a stronger property. A successful result is no...

📖 Read original article


173. Reasoning-supported Robustness Validation of Automotive E/E Components ​

Author: Jan Novacek, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16421v1 Announce Type: new Abstract: This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components. The approach uses formalized knowledge from the RV process and stress, operating...

📖 Read original article


174. ParaTempo: Efficient Parallel Reasoning via Temporal Confidence ​

Author: Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16425v1 Announce Type: new Abstract: Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically...

📖 Read original article


175. Drive, Pack, Fly: The Travelling Thief Problem with Drone ​

Author: Kabir Murjani, Abhay Sobhanan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.NE, math.OC

arXiv:2608.16435v1 Announce Type: new Abstract: In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this penalty by retrieving outlying items, thereby shortening the makespan and increasing o...

📖 Read original article


176. The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach ​

Author: Rafael Pass
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CC, cs.IT, math.IT

arXiv:2608.16438v1 Announce Type: new Abstract: In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but what \emph{value} remains in the inputs (i.e., the prompts) we provide to it. Given a...

📖 Read original article


177. Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics ​

Author: Riccardo Andreoni, Andrei Buliga, Alessandro Daniele, Paolo Felli, Chiara Ghidini, Marco Montali, Massimiliano Ronzani
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16443v1 Announce Type: new Abstract: Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy approaches have targeted mainly symbolic reasoning in propositional and first-order logics, recent works have st...

📖 Read original article


178. HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents ​

Author: Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.16447v1 Announce Type: new Abstract: Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management methods such as ReC...

📖 Read original article


179. JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills ​

Author: Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16465v1 Announce Type: new Abstract: Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale. We introduce \textsc{Jailbre...

📖 Read original article


180. Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation ​

Author: Marc P'erez-Roig, David Fern'andez-Narro, Carlos S'aez
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.16482v1 Announce Type: new Abstract: The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned poli...

📖 Read original article


181. Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling ​

Author: Clemens Sch"achter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16507v1 Announce Type: new Abstract: Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited tim...

📖 Read original article


182. DeepInsight II: One Trace from Benchmark to Robot ​

Author: Siyi Li, Yuchen Kang, Wuliang Wang, Zhengjie Zhang, Jiangpin Liu, Jianhao Yao, Jie Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16556v1 Announce Type: new Abstract: Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific ...

📖 Read original article


183. CUBICS: Situation-aware performance estimation for safety-relevant ML components ​

Author: Benjamin Herd, Jessica Kelly, Mario Trapp
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16564v1 Announce Type: new Abstract: Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.g. by running ML components (M...

📖 Read original article


184. Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I) ​

Author: Robert Peharz
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.PR

arXiv:2608.16565v1 Announce Type: new Abstract: This cumulative habilitation thesis studies probabilistic circuits (PCs) as a powerful and tractable framework for reasoning and learning under uncertainty in artificial intelligence (AI). It first advocates for probability as a core language for AI, e...

📖 Read original article


185. Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents ​

Author: Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SI

arXiv:2608.16578v1 Announce Type: new Abstract: AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplif...

📖 Read original article


186. CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction ​

Author: Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li, Xiaomeng Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16594v1 Announce Type: new Abstract: Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while patient reports remain underexp...

📖 Read original article


187. Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate ​

Author: Yusuke Takahashi, Kyle Wild, Asako Uraki
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.IR

arXiv:2608.16621v1 Announce Type: new Abstract: Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is the...

📖 Read original article


188. A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory ​

Author: Zhihui Chen, Yize Sun, Yuhao Dong, Zeyu Xiao, Ray Y. Zhong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16626v1 Announce Type: new Abstract: Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, which, in turn, can be used to support dynamic shop floor production planning and scheduling. Within such an enviro...

📖 Read original article


189. Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement ​

Author: Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16628v1 Announce Type: new Abstract: Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, suc...

📖 Read original article


190. PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning ​

Author: Veit Laule, Jiangtao Shuai, Manfred Hauswirth, Sonja Schimmler
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16637v1 Announce Type: new Abstract: LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowing symbolic planners ...

📖 Read original article


191. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies ​

Author: Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2608.16645v1 Announce Type: new Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or fu...

📖 Read original article


192. Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents ​

Author: Amrapali Pednekar, Alvaro Garrido-Perez, Yara Khaluf, Pieter Simoens
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16666v1 Announce Type: new Abstract: This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspired by Overcooked, the suite comprises cooking scenarios that require temporal decision making. The tasks and rewar...

📖 Read original article


193. FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy ​

Author: Aniri, Chen Yilin, Jinhe Bi, Junfei Guo, Donglai Ran, Xu Bian, Zengjie Jin, Yujun Wang, Yijun Tian, Volker Tresp, Fei Shen, Tat-Seng Chua, Yunpu Ma
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16697v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability witho...

📖 Read original article


194. LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing ​

Author: Ruoqi Shu, Xuhui Wang, Isaac Wang, Yanming Mai, Bo Wan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16763v1 Announce Type: new Abstract: Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy, consistency, and reproducibility under strict enterprise constraints. In practice, documents arrive with heterog...

📖 Read original article


195. GRIP: Grounded Reasoning via Information-Restricted Premises ​

Author: Lirui Teng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16776v1 Announce Type: new Abstract: High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call this failure mode query dominance. To address it, we introduce \textbf{GRIP} (Ground...

📖 Read original article


196. When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding ​

Author: Giuseppe Destefanis, Tomaso Aste
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report whether the agents complete the task and how much the run costs, leaving the coordination inside the team largely unmeasured. We intro...

📖 Read original article


197. Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment ​

Author: Keren Artiaga (Victor), Yang Li (Victor), Ercan Engin Kuruoglu (Victor), Wai Kin (Victor), Chan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16804v1 Announce Type: new Abstract: Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100 distinct sign languages are severely lacking. In response, we present our work on sign language recognition usin...

📖 Read original article


198. Quipu: A Governed Bitemporal Knowledge Graph Store ​

Author: Steve Brown
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2608.16813v1 Announce Type: new Abstract: Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to d...

📖 Read original article


199. Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning ​

Author: Minh-Ha Nguyen, Cathy Shyr
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.16831v1 Announce Type: new Abstract: Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Huma...

📖 Read original article


200. What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models ​

Author: Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16852v1 Announce Type: new Abstract: Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such mo...

📖 Read original article


201. A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation ​

Author: Shide Zhou, Kailong Wang, Ling Shi, Haoyu Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.04784v1 Announce Type: cross Abstract: Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current benchmarks primarily rely on static datasets susceptible to data contamination or synthetic tasks lacking fine-grain...

📖 Read original article


202. Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale ​

Author: Nisha Sarwar, Lei Jiang, Fan Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.14557v1 Announce Type: cross Abstract: Low Earth Orbit (LEO) computing is emerging for low-latency, globally distributed AI services, enabled by advances in satellite constellations and reusable launch systems. However, its sustainability remains unclear. Prior work introduces ESpaS, a fr...

📖 Read original article


203. Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation) ​

Author: Rivaan Patil, Simon Dennis, Hao Guo, Kevin Shabahang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14563v1 Announce Type: cross Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard fine-tuning at ~40% less peak training memory, while leaving off-domain benchmarks within...

📖 Read original article


204. WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents ​

Author: Yuan Guo, Yilong Chen, Chao Hu, Xianghao Yu, Liang Hong, Jie Xu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.14573v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creating new opportunities for automating scientific and engineering research. To the best of our knowledge, this pap...

📖 Read original article


205. From Reactive to Autonomous: Evolution of AI Operations in Cloud Network Infrastructure ​

Author: Arun Malik
Published: 8/18/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.ET

arXiv:2608.14574v1 Announce Type: cross Abstract: The operational model for cloud network infrastructure has undergone a fundamental transformation over the past decade. What began as manual, human-driven troubleshooting has evolved through scripted automation, rule-based systems, and AI-assisted op...

📖 Read original article


206. HarmProfile: Characterizing Harmful Distributions in Frontier LLMs ​

Author: Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14577v1 Announce Type: cross Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partl...

📖 Read original article


207. Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework ​

Author: Hailong Yang, Jianqi Wang, Guanjin Wang, Zhaohong Deng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14584v1 Announce Type: cross Abstract: In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision making. Despite rec...

📖 Read original article


208. Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures ​

Author: Haibo HU, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressur...

📖 Read original article


209. Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach ​

Author: Zhenyu Tao, Yuxuan Li, Wei Xu, Yongming Huang, Xiaohu You
Published: 8/18/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.14599v1 Announce Type: cross Abstract: The placement of base station (BS) is a fundamental determinant of coverage and capacity of urban wireless networks. Yet large-scale BS deployment optimization remains challenging due to its dependency on site-specific radio propagation and user spat...

📖 Read original article


210. Extend the Safety Horizon for Intelligent Transportation Systems through Semantic-Aware Cooperative Perception ​

Author: Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CV, cs.IT, math.IT

arXiv:2608.14603v1 Announce Type: cross Abstract: Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots. While critical for autonomous driving and safet...

📖 Read original article


211. Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models ​

Author: Aryuemaan Kumar Chowdhury, Praveen Oosa, Vineesha Reddy
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14604v1 Announce Type: cross Abstract: Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and controlled scientific study, yet most of them reuse the standard transformer block without adaptation to the sm...

📖 Read original article


212. Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents ​

Author: Mantas Lukauskas, Viktorija \v{S}arkauskait.e
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, stat.AP

arXiv:2608.14606v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribut...

📖 Read original article


213. Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model ​

Author: Jaroslaw Grobelny, Mateusz Klakus, Kacper Szyma'nski, Teresa Chirkowska-Smolak
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.14609v1 Announce Type: cross Abstract: As artificial intelligence (AI) rapidly diffuses and concerns about job displacement intensify, the psychological mechanisms underlying AI job replacement anxiety remain insufficiently understood. Drawing on Integrated Fear Acquisition Theory and the...

📖 Read original article


214. DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs ​

Author: Zeyu Cao, Xuan Guo, Cheng Zhang, Cheuk Hang Lau, Ilia Shumailov, Yiren Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2608.14614v1 Announce Type: cross Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM in...

📖 Read original article


215. Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion ​

Author: Surya Saka
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.14617v1 Announce Type: cross Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline. ...

📖 Read original article


216. Explaining Reinforcement Learning Decisions in Self-adaptive Systems ​

Author: Jasmina Gajcin, Juan C. Rosero, Ivana Dusparic
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14620v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural networks, lack transparency and are difficult to understand. This can lead to diminished user trust, an...

📖 Read original article


Author: Lin Du, Jie Zhou, Yuxuan Cai, Kai Chen, Qin Chen, Xin Li, Bo Zhang, Wei Li, Liang He
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14621v1 Announce Type: cross Abstract: Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, how to store it, how to retrieve it, and how to manage it can vary substantially across tasks and backbone models...

📖 Read original article


218. Local AI pre-screening for human triple-blind peer review in health sciences ​

Author: Rodrigo Martins Boos
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.DL

arXiv:2608.14625v1 Announce Type: cross Abstract: Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received 12,107. This volume has outpaced the supply of qualified reviewers, and large language models (LLMs) are alread...

📖 Read original article


219. LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review ​

Author: Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdulmumin, Seid Muhie Yimam, Shamsuddeen Hassan Muhammad
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a System...

📖 Read original article


220. Inference-Time Mitigation of Adversarial Political Bias in Large Language Models ​

Author: Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14629v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to political bias is a critical step towards safer and more trustworthy Artificial Inte...

📖 Read original article


221. Characterizing Rhetorical Misalignment in Decision-Making with Language Models ​

Author: Zirui Cheng, Joey Chan, Simo Du, Chenhao Tan, Yue Guo, Hao Peng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14630v1 Announce Type: cross Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can am...

📖 Read original article


222. DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models ​

Author: Xin Zhang, Yili Wang, Yue Tan, Xin He, Yanyu Qian, Yixin Liu, Yi Chang, Shirui Pan, Xin Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14632v1 Announce Type: cross Abstract: Diffusion large language models (D-LLMs) have emerged as a promising paradigm for text generation. However, similar to autoregressive LLMs, D-LLMs remain vulnerable to hallucinations, where fluent outputs may contain factually incorrect or unsupporte...

📖 Read original article


223. Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network ​

Author: Sebastian Raubitzek, Georg Goldenits, Sebastian Schrittwieser, Philip K"onig, Kevin Mallinger
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14636v1 Announce Type: cross Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training. Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fr...

📖 Read original article


224. Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays ​

Author: Bhaskar Gurram
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.14639v1 Announce Type: cross Abstract: Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- is the trust contract document-extraction systems need, and the natural procedure silently violates it on real d...

📖 Read original article


225. BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials ​

Author: An Vuong, Chen Zhao, Jin Hu, Shui-Qing Yu, Xintao Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.mtrl-sci, cs.AI

arXiv:2608.14640v1 Announce Type: cross Abstract: Stacked bilayer materials exhibit rich stacking-dependent properties driven by the interplay between strong intra-layer bonding and weak inter-layer van der Waals interactions. The computational discovery of such materials is challenging because accu...

📖 Read original article


226. iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration ​

Author: Xiaowei Jiang, Daniel Leong, Beining Cao, Nan Zhou, Yingtao Ren, Yu-Cheng Chang, Thomas Do, Chin-Teng Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC

arXiv:2608.14646v1 Announce Type: cross Abstract: Interpretable representation learning remains a key challenge in modern neural computation, particularly when models are expected not only to perform but also to explain their reasoning. This paper introduces iFuzz-Meta, an interpretable fuzzy rule-b...

📖 Read original article


227. SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation ​

Author: Chenyang Jiang, Changhan Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14647v1 Announce Type: cross Abstract: Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, later turns are conditioned on that reply, and uniform distillation can spend loss on tokens that carry little co...

📖 Read original article


228. Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings ​

Author: Leonardo Kuffo, Peter Boncz
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG

arXiv:2608.14648v1 Announce Type: cross Abstract: In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clustering: dimensionality reduction, quantization, and dimension pruning. We propose an indexing pipeline in which...

📖 Read original article


229. Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling ​

Author: Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.14652v1 Announce Type: cross Abstract: The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as decades of reanalysis are only available at 0.25$^{\circ}$ resolution. While exis...

📖 Read original article


230. Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms ​

Author: Xianzong Wu, Xiaohong Li, Yuejun Guo, Xinyang Liu, Tianlin Li, Junjie Wang, Qiang Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14653v1 Announce Type: cross Abstract: Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation, data selection, and prediction rollback. Despite its demonstrated utility, the potential of uncertainty quant...

📖 Read original article


231. FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting ​

Author: Hai Anh Tran, Cuong Ta, Truong X. Tran
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14654v1 Announce Type: cross Abstract: Federated Learning (FL) is a collaborative paradigm that enables multiple devices to train a global model while preserving local data privacy. A major challenge in FL is the non-Independent and Identically Distributed (non-IID) nature of data across ...

📖 Read original article


232. P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval ​

Author: Zhongli Wu, Zhuangzhi Gao, He Zhao, Feixiang Zhou, Fu Wang, Jinru Ding, Yuankai Wang, Hongyi Qin, Gregory Y. H. Lip, Bil Kirmani, Yalin Zheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14656v1 Announce Type: cross Abstract: Photoplethysmography (PPG) is widely used in consumer wearables because of its low cost and ease of acquisition. However, unlike electrocardiography (ECG), PPG measures peripheral pulse dynamics rather than cardiac electrical activity, limiting its a...

📖 Read original article


233. pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier ​

Author: Gautam Kishore
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.IR

arXiv:2608.14658v1 Announce Type: cross Abstract: We introduce pico-type, a byte-level multi-head content classifier with approximately 1.5 million parameters that simultaneously predicts seven content properties from raw UTF-8 bytes in a single forward pass. Operating directly at the byte level -- ...

📖 Read original article


234. Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow ​

Author: Shun Nakayama, Takahiro Kanamori, Wanglin Yan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2608.14660v1 Announce Type: cross Abstract: This study proposes a ring-based SpatialTransformer to learn how building uses at different distances from a railway station interact to generate pedestrian flow. Concentric ring buffers at 100-meter intervals up to 800 meters were defined around 100...

📖 Read original article


235. Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning ​

Author: Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
Published: 8/18/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2608.14662v1 Announce Type: cross Abstract: Accurate recognition of pain using physiological signals remains a challenging problem due to pain's subjective nature and high inter-individual variability. In this study, we investigate self-supervised representation learning (SSL) methods applied ...

📖 Read original article


236. BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement ​

Author: Kaixiang Wang, Yidan Lin, Jiong Lou, Jie Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.14668v1 Announce Type: cross Abstract: LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate hallucinated or malicious outputs into system-level failures. Auditor agents mitigate these risks, yet existing ...

📖 Read original article


237. ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry ​

Author: Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li, Katia Obraczka
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.14670v1 Announce Type: cross Abstract: Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches rely largely on gait or activity cues and are rarely evaluated at scale. In this paper, we present \emph{Argus...

📖 Read original article


238. Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection ​

Author: Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven. In this work, we evalu...

📖 Read original article


239. Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation ​

Author: Lech Kalinowski, Artur Morys-Magiera, Piotr Mi{\l}kowski
Published: 8/18/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2608.14677v1 Announce Type: cross Abstract: Most mobile image-generation applications are thin clients over cloud services, leaving outputs hard to audit. We present an Android latent-diffusion application that runs entirely on-device and is driven by the ambient-light sensor rather than a tex...

📖 Read original article


240. Information-Theoretic Causal Modelling of Semiconductor Process Dynamics ​

Author: Daniel S{\o}rensen, Giorgio Melchiorre, Sudip Bandyopadhyay, Sandip Halder, Roel Wuyts, Bappaditya Dey
Published: 8/18/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.IT, math.IT

arXiv:2608.14678v1 Announce Type: cross Abstract: With the progress of the semiconductor industry toward increasingly complex compute devices and tighter process tolerances, advanced process control has become crucial. This work explores a novel framework to infer the underlying dynamics of semicond...

📖 Read original article


241. Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans ​

Author: Jinglei Ren, Yuyue Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming...

📖 Read original article


242. Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation ​

Author: Dingyao Yu, Tong Zhang, Yutao Mou, Yunxiao Zhang, Wei Ye, Shikun Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14684v1 Announce Type: cross Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. Evaluating all rubrics in a single pass is a natural alter...

📖 Read original article


243. Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level ​

Author: Hannah Cha
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.14692v1 Announce Type: cross Abstract: Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they o...

📖 Read original article


244. Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning ​

Author: Aravindhan Arunagiri, Ayaan Khan, Udayaadithya Avadhanam, SaiBarath Sundar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14693v1 Announce Type: cross Abstract: With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial component of data privacy and compliance frameworks. Traditional text sanitization solutions are majorly suitabl...

📖 Read original article


245. Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning ​

Author: Hansen Jin Lillemark, Alex Rojas, Zachary Novack, Runqian Wang, Yilun Du, Yian Ma, Taylor Berg-Kirkpatrick, Rose Yu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.14706v1 Announce Type: cross Abstract: Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. We introduce Equilibrium Forcing (EqF),...

📖 Read original article


246. Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics ​

Author: Ruochen Liu, Wei Lou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.14710v1 Announce Type: cross Abstract: Predicting spatial gene expression from hematoxylin and eosin (H&E)-stained images offers a cost-effective alternative to spatial transcriptomics (ST). However, existing methods treat H&E images as generic visual inputs and ignore their intrinsic b...

📖 Read original article


247. Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data ​

Author: Marios Papamichalis, Regina Ruane
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, math.ST, stat.TH

arXiv:2608.14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similar...

📖 Read original article


248. DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis ​

Author: Xiaoxiao Li, Xitong Ling, Jiawen Li, Weiming Chen, Zhenyang Cai, Xidong Wang, Tian Guan, Benyou Wang, Yonghong He
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14719v1 Announce Type: cross Abstract: Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distributions, MIL-based WSI analysis faces a nested dual long-tail: an inter-slide class long tail and an intra-slide ...

📖 Read original article


249. Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra ​

Author: Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin
Published: 8/18/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI

arXiv:2608.14720v1 Announce Type: cross Abstract: Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remains an outstanding challenge. Despite decades of computational efforts, no existing system achieved reliable rea...

📖 Read original article


250. Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering ​

Author: Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.14724v1 Announce Type: cross Abstract: The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur,...

📖 Read original article


251. Tail-Aware Top-$k$ On-Policy Distillation ​

Author: Huipeng Huang, Hongxin Wei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14728v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distribution with the teacher's along its own trajectories. To provide dense supe...

📖 Read original article


252. A Novel Fourier Feature Network for Solving Partial Differential Equations ​

Author: Qihong Yang, Zhijie Su, Yangtao Deng, Qiaolin He
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14733v1 Announce Type: cross Abstract: Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier features using $\cos$, $\sin$, or a combination of both. Similar to Extreme Learning Machines (ELMs), FENs empl...

📖 Read original article


253. Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks ​

Author: Kai Gu, Haizheng Zhong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.mtrl-sci, cs.AI

arXiv:2608.14734v1 Announce Type: cross Abstract: Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. However, their black-box nature hinders gaining deep insights into the underlying synthetic mechanisms. Here, we dev...

📖 Read original article


254. Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews ​

Author: Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Clayton Suguio Hida
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14737v1 Announce Type: cross Abstract: This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An experiment was conducted in five reviews, comparing individual and batch processing, with and without prevalence m...

📖 Read original article


255. PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models ​

Author: Siddharth Patel
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14741v1 Announce Type: cross Abstract: We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be ...

📖 Read original article


256. Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI ​

Author: Julia Huang, Camila Gonzalez, Rydham Goyal, Aja Zou, Sasha Alexander, Michael Moseley, Moss Y. Zhao, Gary K. Steinberg
Published: 8/18/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.14758v1 Announce Type: cross Abstract: For patients with Moyamoya disease, impaired cerebrovascular reserve (CVR) is an important hemodynamic criterion for recommending extracranial-to-intracranial bypass surgery. Standard CVR assessment in this cohort uses paired arterial spin labeling (...

📖 Read original article


257. Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening ​

Author: Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi
Published: 8/18/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG

arXiv:2608.14763v1 Announce Type: cross Abstract: Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement. Fetal brain MR...

📖 Read original article


258. NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving ​

Author: Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.RO

arXiv:2608.14767v1 Announce Type: cross Abstract: Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather ...

📖 Read original article


259. Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation ​

Author: Mark Nowak (Conflux Laboratory)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14770v1 Announce Type: cross Abstract: An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic household visits and interviews, systematically under-detect them. We present a real-time computer-vision pipel...

📖 Read original article


260. ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning ​

Author: Harshil Lodhiya
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14773v1 Announce Type: cross Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. We show that this choice conceals a large capability differ...

📖 Read original article


261. Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters ​

Author: Bernardo Modenesi, Jody Lin, Kimberly Kaphingst, Angela Zhu, Maya Wheeler, Peilu Zhang, Angela Fagerlin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.14792v1 Announce Type: cross Abstract: Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested ...

📖 Read original article


262. Handover Analysis for Vehicular Communication with Explainability on the Fly ​

Author: Ali Fuat Sahin, Semiha Tedik Ba\c{s}aran, Tufan Kumbasar
Published: 8/18/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2608.14820v1 Announce Type: cross Abstract: Handover (HO) management in vehicular networks requires fast and reliable decision-making under highly dynamic conditions. While machine learning (ML) approaches can improve HO detection by capturing complex relationships among various key performanc...

📖 Read original article


263. Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce ​

Author: Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single ...

📖 Read original article


264. Writing Style Similarity Reflects Academic Genealogy ​

Author: Cameron Manzo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14843v1 Announce Type: cross Abstract: As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems assume each author's style is their own. Researchers, however, ...

📖 Read original article


265. Evaluating Agentic Code Repair Capabilities in Distributed Systems ​

Author: Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.DC

arXiv:2608.14863v1 Announce Type: cross Abstract: LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes,...

📖 Read original article


266. Workspace Topology as an Attack Vector in Agentic Coding Assistants ​

Author: Alexandre G. R. Day, Pradeep Yadlapalli, Sriram Venkatapathy, Thomas Paniagua, Nick Raines, Sahil Wadhwa, Himanshu Kumar, Andy Luo, Sudeep Panyam, Rikhiya Ghosh, Pranab Mohanty, Giri Iyengar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG

arXiv:2608.14876v1 Announce Type: cross Abstract: Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code. This opens up a risk of malicious code being ingested as these coding tools operate with broad filesystem...

📖 Read original article


267. The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency ​

Author: Michael Glass
Published: 8/18/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA

arXiv:2608.14913v1 Announce Type: cross Abstract: We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each player's strategy is a natural-language document visible to all participants. The dictator's decision, to SHARE or TAKE an endowment, may depend...

📖 Read original article


268. Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task ​

Author: Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei, Laura Erhan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG

arXiv:2608.14916v1 Announce Type: cross Abstract: AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shifted, or otherwise edited. These edited versions form a difficult negative class: they are not generated by AI,...

📖 Read original article


269. SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable ​

Author: Philip H. Lee, Parth Padalkar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14922v1 Announce Type: cross Abstract: Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features. However, because ...

📖 Read original article


270. PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining ​

Author: Azim Dehghani Amirabad, Junchao Zhu, Pushpak Pati, Walid Abdelmoula, Tommaso Mansi, Rui Liao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14924v1 Announce Type: cross Abstract: Spatial transcriptomics (ST) links tissue morphology with molecular programs, motivating multimodal pretraining methods that align histology images with gene expression. However, existing approaches suffer from two key limitations: spatially informat...

📖 Read original article


271. Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment ​

Author: Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14942v1 Announce Type: cross Abstract: This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial image...

📖 Read original article


272. Generative data assimilation highlights fronts as key regulators of ocean energy cascade ​

Author: Scott A. Martin, Georgy E. Manucharyan, Patrice Klein
Published: 8/18/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI

arXiv:2608.14955v1 Announce Type: cross Abstract: Mesoscale eddies are fundamental to the ocean circulation, yet the extent to which submesoscale motions, a few kilometers across, influence mesoscale eddy energetics through a kinetic energy cascade remains uncertain. High-resolution simulations pred...

📖 Read original article


273. Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning ​

Author: Joanikij Chulev, Hendrik Baier
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC

arXiv:2608.14963v1 Announce Type: cross Abstract: Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-s...

📖 Read original article


274. Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games? ​

Author: Wenji Fu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.14982v1 Announce Type: cross Abstract: Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time. We ask whether geometry-aware positional encodings improve these capabilities, without claiming a new positional enco...

📖 Read original article


275. GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation ​

Author: Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.14986v1 Announce Type: cross Abstract: Long-horizon robotic manipulation fundamentally relies on persistent spatial memory. However, existing 3D memory systems function merely as passive recorders: they store observations using fixed, hand-crafted rules, treating every scene element--whet...

📖 Read original article


276. PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity ​

Author: Jindi Wu, Qun Li
Published: 8/18/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.DC

arXiv:2608.14995v1 Announce Type: cross Abstract: Quantum federated learning (QFL) lets multiple quantum clients collaboratively train quantum neural networks (QNNs) without sharing private local data. However, existing QFL methods commonly assume that all clients use the same ansatz, overlooking ho...

📖 Read original article


277. RamseyGadgets: A Graph Construction Dataset for LLMs ​

Author: Zohair Raza Hassan, Deepak Pandita
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration of relevant graphs and human ingenuity. Given the rise of generative AI usage in...

📖 Read original article


278. FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making ​

Author: Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15004v1 Announce Type: cross Abstract: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomic...

📖 Read original article


279. MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems ​

Author: Penghao Yin, Haomin Wang, Qihong Tang, Xiaoye Qu, Hongjie Zhang, Xiao-Ping Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2608.15006v1 Announce Type: cross Abstract: Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning. Some recent methods introduce intermediate visual states to facilitate reasoning, but they are often hindere...

📖 Read original article


280. SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system ​

Author: Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA

arXiv:2608.15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive. Despite substantial prior work across cy...

📖 Read original article


281. Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference ​

Author: Yiran Gao, Juntao Chen, Tao Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.15016v1 Announce Type: cross Abstract: Network incident response remains slow and labor-intensive as the defender must infer multi-stage attacks from partial observations and translate recovery decisions into reliable system commands. Decision-theoretic planners provide principled optimiz...

📖 Read original article


282. DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest ​

Author: Alibek Kamiluly, Milana Muratova, Yash Patel, Fan Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.15019v1 Announce Type: cross Abstract: Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast, irregular shapes, and boundaries that blend with surrounding breast tissue. To address this problem, we presen...

📖 Read original article


283. MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM ​

Author: Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.15024v1 Announce Type: cross Abstract: Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they attempt the ill-posed inverse problem of recovering sharp content from degraded observations. We present MotionGS-SLAM, which fundamentally re...

📖 Read original article


284. Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints ​

Author: Bruno Chicelli, Henrique Alves, Rodrigo Anselmo, Joshua Weinberg, Felipe Lemos, Jan Baryla
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2608.15032v1 Announce Type: cross Abstract: Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state. We ...

📖 Read original article


285. GATTA: Graph Active Learning with Test-Time Augmentation ​

Author: Zsombor B'anfi, Andr'as G'ezsi, Andr'as Formanek
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15084v1 Announce Type: cross Abstract: Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its application to graph-structured data remains largely unexplored. We introduce GATTA (Graph Active Learning with Te...

📖 Read original article


286. Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning ​

Author: Zihang Wang, Yishan Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15088v1 Announce Type: cross Abstract: Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior. We present a training method for this setting based on two components. First, an \emph{M...

📖 Read original article


287. WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing ​

Author: Zhiyu Zhang, Tingyue Wen, Senke Sun, Dengxiang Liang, Enhao Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE

arXiv:2608.15092v1 Announce Type: cross Abstract: In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks specify only functional objectives without explicit security requirements. WeSCE consists of 400 executable prog...

📖 Read original article


288. Beyond Direct Access: Resource Hijacking in LLM Agents ​

Author: Puyu Zeng, Qibing Ren
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.15108v1 Announce Type: cross Abstract: Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Existing agent security ...

📖 Read original article


289. CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation ​

Author: Peng Jia, Li Dai, Zhen Xiao, Xueliang Liu, Jia Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15110v1 Announce Type: cross Abstract: Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. The...

📖 Read original article


290. Fast Test-Time Refinement for Robust Learned Image Compression ​

Author: Jiaming Liang, Chi-Man Pun, Weisi Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.15113v1 Announce Type: cross Abstract: Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high representational capacity endowed by deep neural networks (DNNs) comes at the expense of increased adversarial vulnerab...

📖 Read original article


291. From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems ​

Author: Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.OS, cs.AI, cs.DC, cs.MA

arXiv:2608.15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottle...

📖 Read original article


292. Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models ​

Author: Varvara Arzt, Allan Hanbury, Terra Blevins
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15129v1 Announce Type: cross Abstract: We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither...

📖 Read original article


293. Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems ​

Author: Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng
Published: 8/18/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2608.15144v1 Announce Type: cross Abstract: Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generally intractable. We begin from an ideal one-parameter posterior SDE family in which a stochasticity parameter con...

📖 Read original article


294. Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models ​

Author: Yang Liu, Yuming Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15156v1 Announce Type: cross Abstract: World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether a small, directly addressable hidden-state change can place a learned world model on the intended counterfactu...

📖 Read original article


295. A Unified Backbone--Expert Framework with Relation-Token and Residual--Classifier Interfaces for Automatic Modulation Recognition ​

Author: Zhixiang Deng, Houbiao Li, Zongyong Cui
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15160v1 Announce Type: cross Abstract: Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model architecture often fails to excel. To address this, we propose a unified backbone-expert framework with a common ...

📖 Read original article


296. LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset ​

Author: Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti'errez Gait'an, Atefeh Hajijamali Arani, Rui Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making. Existing optimization-based, Machine ...

📖 Read original article


297. FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection ​

Author: Yixuan Chen, Hongyu Zhan, Jie Sheng, Weiyu Han, Shuai Chen, Tianyi Zhang, Xiao Tan, Jun Xia
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15177v1 Announce Type: cross Abstract: The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over interconnected financial entities. This shift has motivated graph-based fraud d...

📖 Read original article


298. The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests ​

Author: Douglas J. Leith
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.15188v1 Announce Type: cross Abstract: We evaluate the quality of Claude AI-written Python tests against human-written Python tests from two established open-source projects Django and Pandas. Hundreds of tests per corpus are scored under one identical protocol. Using one-sided non-inferi...

📖 Read original article


299. Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work ​

Author: Yuyang Zheng, Nan Li, Wenxia Deng, Lige Yan, Xiang Li, Si Chen
Published: 8/18/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2608.15193v1 Announce Type: cross Abstract: As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information retrieval and task continuity. However, most structured knowledge systems...

📖 Read original article


300. CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction ​

Author: Tran Xuan Hieu Le, Doanh C. Bui, Vu Trung Duong Le, Hoai Luan Pham, Khang Nguyen, Mai K. Nguyen, Tu Bao Ho, Yasuhiko Nakashima
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15246v1 Announce Type: cross Abstract: Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highly ill-posed and often produces severe streak artifacts. Existing deep reconstruction methods have achieved prom...

📖 Read original article


301. UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection ​

Author: Zhiqiang Hu, Shouren Huang, Masatoshi Ishikawa
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2608.15259v1 Announce Type: cross Abstract: Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to traffic surveillance. However, aerial video footage often suffers from severe motion blur due to rapid flight maneuvers, vibrations, and camera...

📖 Read original article


302. VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction ​

Author: Wei Zhang, Yihang Wu, Songhua Li, Qi Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15260v1 Announce Type: cross Abstract: Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment ...

📖 Read original article


303. VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments ​

Author: Haolin Yang, Yuxing Long, Zihan Yang, Hao Dong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.CV, cs.MM

arXiv:2608.15284v1 Announce Type: cross Abstract: Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction. Prior instruction generators assume discrete viewpoint graphs ...

📖 Read original article


304. PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies ​

Author: Yufei Guo, Yinan Wu, Haoran Duan, Guiguang Ding, Jungong Han
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15285v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of contin...

📖 Read original article


305. No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage ​

Author: Shiven Khurdi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15286v1 Announce Type: cross Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. Acro...

📖 Read original article


306. MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation ​

Author: Lie Li, Wen Li, Junxiao Shen, Gusheng Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15299v1 Announce Type: cross Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. We demonstrate that this uniformity is...

📖 Read original article


307. Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning ​

Author: Alexandre L. M. Levada
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ML

arXiv:2608.15313v1 Announce Type: cross Abstract: In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and dimensionality reduction that incorporates differential geometric information into the covariance structure of c...

📖 Read original article


308. Logical Embeddings for Argument Analysis ​

Author: Leander Heldring, Santiago Torres
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15325v1 Announce Type: cross Abstract: We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualized word embeddings used in most NLP tasks with logical embeddings, an alternative encoding that directly exploit...

📖 Read original article


309. When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text ​

Author: Shresth Shroff
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15338v1 Announce Type: cross Abstract: Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment clas...

📖 Read original article


310. ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution ​

Author: Duong M. Nguyen, Tuan Nghia Nguyen, Xuan Truong Nguyen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV

arXiv:2608.15349v1 Announce Type: cross Abstract: To accelerate single image super-resolution (SISR) networks on large images (2K-8K), many recent approaches decompose an image into small patches and dynamically determine an execution path according to its difficulty (referred to as a dynamic networ...

📖 Read original article


311. SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning ​

Author: Mohammad Aref Jafari-Raddani, Morteza Mohajjel Kafshdooz
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15360v1 Announce Type: cross Abstract: While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter s...

📖 Read original article


312. AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization ​

Author: Sahil Gangurde
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR

arXiv:2608.15369v1 Announce Type: cross Abstract: Lossy audio compression algorithms traditionally rely on psychoacoustic modeling and frequency-domain representations (e.g., MP3, AAC, and Opus) to discard information that is imperceptible to the human auditory system. While highly effective, these ...

📖 Read original article


313. Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die ​

Author: Botao Amber Hu, Fangting
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.15403v1 Announce Type: cross Abstract: You will die eventually. Your agents may not. An AI agent operating on decentralized blockchain infrastructure has no concept of death; it can only go bankrupt -- frozen when its wallet can no longer pay for its next transaction -- and revived the mo...

📖 Read original article


314. Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals ​

Author: Botao Amber Hu, Iris Long
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.15405v1 Announce Type: cross Abstract: Afterlife Delegation Protocol is a speculative design project that asks what death becomes when a will can act eternally. We design a speculative protocol through which a living person signs an agentic will: upon a verified death, a self-sovereign AI...

📖 Read original article


315. Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees ​

Author: Rohit Swami, Tushar Singh, Akash Warde, Sri Muthu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.NE

arXiv:2608.15407v1 Announce Type: cross Abstract: An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled adversary can confirm the presence of a deception environment within only a few diagnostic commands, limiting its intelligence value. High-...

📖 Read original article


316. FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge ​

Author: Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim, Sing-Yao Wu, Eli Bozorgzadeh, Nalini Venkatasubramanian, Nikil Dutt
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CV, cs.RO, cs.SY, eess.SY

arXiv:2608.15410v1 Announce Type: cross Abstract: Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interface for embodied agents. However, existing benchmarks largely focus on...

📖 Read original article


317. Invariant Pretraining for Robust Code Representations ​

Author: Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo, Zili Wang, Hao Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE

arXiv:2608.15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. Their robustness, however, is fragile: under invaria...

📖 Read original article


318. An Evaluation Framework for National AI Regulation ​

Author: Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory, Sreeparvathy Sajeev, Utkarsh Tomar, Avyay M Casheekar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.15417v1 Announce Type: cross Abstract: Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these national approaches is difficult. A binding rule and a detailed voluntary framework can address the same problem but c...

📖 Read original article


319. ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems ​

Author: Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey
Published: 8/18/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG

arXiv:2608.15424v1 Announce Type: cross Abstract: The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of thes...

📖 Read original article


320. NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models ​

Author: Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existin...

📖 Read original article


Author: Volodymyr Ovcharov
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2608.15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as reco...

📖 Read original article


322. Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off ​

Author: Aditya Singh
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15459v1 Announce Type: cross Abstract: Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose reasoning. This survey covers four connected threads: their formulation for sequence-to-sequence tasks, adaptat...

📖 Read original article


323. Optimal Lower Bounds for Networked Information Aggregation ​

Author: Ambar Pal
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.15472v1 Announce Type: cross Abstract: The problem of networked information aggregation, studied in Kearns et al. (2026), involves a group of learners situated on the vertices of a directed acyclic graph $G$, each learning a linear predictor $\widehat Y$ for a fixed random variable $Y$ gi...

📖 Read original article


324. Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability ​

Author: Yudong Gao, Linghan Chen, Wenhan Wu, Mia Zhou, Jiyao Wang, Kaiyan Ji, Mingyu Guo, Honglong Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.15475v1 Announce Type: cross Abstract: Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We present the first bit-flip attack on a VLA: a few gradient-selected flips reduce closed-loop success to $0%$, whil...

📖 Read original article


325. EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation ​

Author: Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan, Yang Yu, Huang Yan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15537v1 Announce Type: cross Abstract: Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts. From a signal-processing perspective, lesion boundaries repres...

📖 Read original article


326. Spectral Saliency for Machine Unlearning ​

Author: Cedar Site Bai, Amber Yijia Zheng, Raymond A. Yeh, Brian Bullins
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15548v1 Announce Type: cross Abstract: Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can be viewed as the inverse of learning, using gradient-based updates to reduce the influence of a forget-set by ...

📖 Read original article


327. MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration ​

Author: Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15549v1 Announce Type: cross Abstract: Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine reactive physical behaviors with stateful social behaviors, while existing interfaces often require developers to ...

📖 Read original article


328. Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors ​

Author: Iyad Assaad Nekka, Hamida Seba, Walid Khaled Hidouci, Karima Amrouche
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15559v1 Announce Type: cross Abstract: Anomaly detection in dynamic graphs underpins financial fraud analysis, intrusion detection, and platform integrity, where automated decisions require human-interpretable justifications. StrGNN, the strongest performer in recent benchmarks, produces ...

📖 Read original article


329. Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study ​

Author: Yogesh Kumar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15574v1 Announce Type: cross Abstract: Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupported by the cited frame. This deceptive hallucination arises because timestamps imply grounding without ensuring c...

📖 Read original article


330. ARENA: Automated Red-Teaming for Large Audio Language Models ​

Author: Jiaming He, Zhicong Huang, Tian Jin, Zhen Sun, Cheng Hong, Yi Yu, Wenbo Jiang, Xudong Jiang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.15578v1 Announce Type: cross Abstract: Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated ...

📖 Read original article


331. Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair ​

Author: Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata, Satoshi Nakashima, Yu Ishikawa, Kosuke Maeda, Nao Soma, Kenichi Kobayashi, Keisuke Miyazaki, Keizo Kato, Shigeki Fukuta, Tatsuo Kumano, Nobutaka Imamura, Kevin Musgrave, Shahbaz Abdul Khader, Kwun Ho Ngan, Joe Townsend, Fayas Asharindavida, Matthieu Parizy, Akira Sakai, Yuma Ichikawa, Yang Zhao, Michiaki Takizawa, Taku Fukui, Hiroki Ohtsuji, Wei-Peng Chen, Hiromichi Kobashi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.ET, cs.PL

arXiv:2608.15579v1 Announce Type: cross Abstract: Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operation adds long horizons, tool-use discipline, context persistence, heterogeneous clusters, and evaluation reuse. W...

📖 Read original article


332. GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix ​

Author: Jinhyun Jeon, Sungjoo Yoo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15584v1 Announce Type: cross Abstract: Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands...

📖 Read original article


333. FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy ​

Author: Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration pot...

📖 Read original article


334. EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input ​

Author: Matteo Stoiber, Niels Buus Lassen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15614v1 Announce Type: cross Abstract: The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and compute cost scale with the number of visual tokens, and high-resolution video quickly becomes expensive to tra...

📖 Read original article


335. Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis ​

Author: Alona Strugatski, Licol Zeinfeld, Giora Alexandron
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2608.15630v1 Announce Type: cross Abstract: The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skill...

📖 Read original article


336. Sparse Prototype Code Underlies Classification and Prediction Across Modalities ​

Author: Yehonatan Avidan, Daniel D. Lee, Haim Sompolinsky
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cs.AI, stat.ML

arXiv:2608.15632v1 Announce Type: cross Abstract: Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. We show that classification tasks give rise to a universal rep...

📖 Read original article


337. Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification ​

Author: Chunyu Qi, Zhuoran Song, Jian Weng, Haozhe Jiang, Xueyuan Liu, Naifeng Jing, Guanghui He, Xiaoyao Liang, Haibing Guan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15636v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator...

📖 Read original article


338. When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation ​

Author: Wenhao Yuan, Chenchen Lin, Wenhao Hu, Jian Chen, Jinfeng Xu, Shujie Li, Edith Cheuk Han Ngai
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG

arXiv:2608.15639v1 Announce Type: cross Abstract: \textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split strategy may be suboptimal because clients can differ ...

📖 Read original article


339. Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation ​

Author: Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15647v1 Announce Type: cross Abstract: Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balance...

📖 Read original article


340. When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations ​

Author: Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Ka Man Yan, Ying Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15654v1 Announce Type: cross Abstract: Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as t...

📖 Read original article


341. PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails ​

Author: Satchit Chatterji, Shihan Wang, Giovanni Sileno, Erman Acar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15673v1 Announce Type: cross Abstract: Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those facts imply under a given policy. Common approaches, including policy prom...

📖 Read original article


342. Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation ​

Author: Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.15680v1 Announce Type: cross Abstract: Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder explo...

📖 Read original article


343. Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media ​

Author: Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SI, cs.AI, cs.CL, cs.LG

arXiv:2608.15689v1 Announce Type: cross Abstract: This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We extend the classical Susceptible--Infected--Recovered (SIR) model to a six-compartment structure (SIRMMM), inco...

📖 Read original article


344. Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer ​

Author: Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG, cs.MM

arXiv:2608.15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding ...

📖 Read original article


345. RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation ​

Author: Kareem Hassani, Chaymaa Abbas, Hadi Al Mubasher, Mariette Awad
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15694v1 Announce Type: cross Abstract: Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with no opportunity to improve on it. Although trained to produce the best possible result in one step, such a mode...

📖 Read original article


346. Beyond Single Object: Learning 3D Relations with Large Language Models ​

Author: Kohsuke Ide, Ryousuke Yamada, Yue Qiu, Xianzheng Ma, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.15710v1 Announce Type: cross Abstract: We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for detailed object-level reasoning across multiple objects with three compon...

📖 Read original article


347. FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction ​

Author: Ali Boudaghi, Alireza Nemati, Hadi Zare
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.15727v1 Announce Type: cross Abstract: Diffusion models have recently shown strong potential for multivariate time-series anomaly detection by learning the distribution of normal data through iterative denoising. Existing diffusion-based approaches, however, typically perform anomaly dete...

📖 Read original article


Author: Haadia Amjad, Ronald Tetzlaff
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15731v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex i...

📖 Read original article


349. TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity ​

Author: Armin Steinhauser
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15767v1 Announce Type: cross Abstract: We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter...

📖 Read original article


350. Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection ​

Author: Xudong Chen, Shengbo Gong, Lu Cheng, Wei Jin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15768v1 Announce Type: cross Abstract: Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantification. In edge-level fraud detection on temporal interaction graphs, where false positives and false negatives both...

📖 Read original article


351. ALKEMIE Agent: an autonomous platform for computational materials design ​

Author: Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2608.15776v1 Announce Type: cross Abstract: Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge softwar...

📖 Read original article


352. Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay ​

Author: Di Bai, Feng Han, Zhenwei Tang, Jintao Liu, Luoshu Wang, Jialu Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.15780v1 Announce Type: cross Abstract: Stale recommendations are a pervasive challenge and a leading source of user complaints on large-scale content platforms. Items lose relevance through two primary mechanisms: supersession, where emerging updates render prior coverage stale, and relev...

📖 Read original article


353. Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation ​

Author: Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.15787v1 Announce Type: cross Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts. This creates a possible blind spot in same-weight self-distillation, where a demonstration-conditioned teacher supervises a query-o...

📖 Read original article


354. A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations ​

Author: Ana Naveriani, Jakob Suchan, Stefano Zoia, Mehul Bhatt, Antonio Lieto, Gian Luca Pozzato
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15828v1 Announce Type: cross Abstract: Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomp...

📖 Read original article


355. CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling ​

Author: Bo Zhao, Zheng Wu, Yiping Xie, Zitong YU
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15831v1 Announce Type: cross Abstract: Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable to illumination changes, motion artifacts, and skin-tone-dependent optical reflectance. We propose CardiacMamba,...

📖 Read original article


356. Characterising cardiac tissue properties with graph neural networks ​

Author: Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela
Published: 8/18/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, eess.SP

arXiv:2608.15843v1 Announce Type: cross Abstract: Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for localising ablation targets and improving arrhythmia treatment. We developed a gra...

📖 Read original article


357. Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning ​

Author: Yuxing Long, Lei Kang, Ziyan Yu, Yuzheng Gao, Bin Cheng, Jiyao Zhang, Xiaoqi Li, Haolin Yang, Dongjiang Li, Hui Shen, Hao Dong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.CV, cs.MM

arXiv:2608.15863v1 Announce Type: cross Abstract: Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no sufficiently diverse, task-oriented dataset exists to support such planning. To bridge this ...

📖 Read original article


358. Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes ​

Author: Farbod Abbasi, Zachary Patterson, Bilal Farooq
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15867v1 Announce Type: cross Abstract: Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. Small samples miss valid attribute combinations, known as sampling zeros, and genera...

📖 Read original article


359. Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning ​

Author: Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal, Zhongyu Jiang, Ian Fasel, Joerg Liebelt
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG, cs.MM

arXiv:2608.15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual fo...

📖 Read original article


360. Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive ​

Author: Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, Ren'e Schuster, Stanislav Frolov, Andreas Dengel
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15901v1 Announce Type: cross Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal...

📖 Read original article


361. Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology ​

Author: Hadi Hasan, Safaa Salman, Lama Sleem, Ralph Mouawad, Ali Chehab
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15915v1 Announce Type: cross Abstract: Lung cancer remains the leading cause of cancer-related mortality worldwide, while histopathological diagnosis is often affected by inter-observer variability and the substantial workload associated with manual slide examination. Although deep learni...

📖 Read original article


362. Pre-training Visual Dexterity in Simulation ​

Author: Sarthak Kamat, Adam Rashid, Satvik Sharma, Aseem Doriwala, Chelsea Finn, Phillip Isola, C. Karen Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.15917v1 Announce Type: cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers. Dexterous, multi-fingered hands remain comparativ...

📖 Read original article


363. Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery ​

Author: Nicola Cogotti
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.15919v1 Announce Type: cross Abstract: Retrieval-Augmented Generation over knowledge graphs (Graph-RAG) has emerged as a powerful paradigm for grounding large language models in domain-specific corpora. However, existing systems face persistent limitations: (1) static chunking fragments l...

📖 Read original article


364. Information Geometry of Message Passing ​

Author: Mykola Lukashchuk, Kyrylo Yemets, Alex Ledbetter, .{I}smail \c{S}en"oz
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15922v1 Announce Type: cross Abstract: We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We start from the Bethe free energy and constrain a selected edge marginal to an exponential family. At a stationar...

📖 Read original article


365. Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation ​

Author: Cedar Site Bai, Duanshun Li, Zhenyu Liao, Sheikh Sarwar, Huiyuan Chen, Yuan Chen, Changhe Yuan, Haiyang Zhang, Qilin Qi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.LG

arXiv:2608.15949v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences...

📖 Read original article


366. LLMs Get Smarter from Targeted Synthetic Multilingual Data ​

Author: Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic...

📖 Read original article


367. CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications ​

Author: Yubo Zhang, Yiyao Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.15972v1 Announce Type: cross Abstract: Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geomet...

📖 Read original article


368. A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency ​

Author: Ravi Satya Durga Prasad Yenugula
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CL, cs.LG

arXiv:2608.15975v1 Announce Type: cross Abstract: Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand-labeling every batch is not feasible, and two questions dominate: what label quality a teacher buys per dollar, and how t...

📖 Read original article


369. From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents ​

Author: Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step...

📖 Read original article


370. Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence ​

Author: Rajan Kadel, Bellal Hossain, Samar Shailendra, Bushra Naeem
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.16016v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of conventional assessments that rely on single-point submissions and product-only grading. This paper proposes a design ...

📖 Read original article


371. RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection ​

Author: Junxin Lu, Jing Zhao, Shiliang Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16018v1 Announce Type: cross Abstract: Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture ...

📖 Read original article


372. NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption ​

Author: Ziluowen Luo, Jun Yin, Ruochen Liu, Ming Cheng, Shirui Pan, Chengqi Zhang, Senzhang Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16038v1 Announce Type: cross Abstract: Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. However, such perturbations often introduce substantial distribution...

📖 Read original article


373. Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection ​

Author: Jiadao Zou, Hongyu Guo, Wei Xi
Published: 8/18/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.16039v1 Announce Type: cross Abstract: Brain parcellation and classification are typically evaluated in isolation, yet downstream AD detection performance depends on their interaction. We decouple these components and systematically benchmark fast deep learning parcellation methods (Synth...

📖 Read original article


374. Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents ​

Author: Yike Yuan, Virum Ranka, Tina Lasisi, Lin Ma
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.16045v1 Announce Type: cross Abstract: LLM-based data-analysis tools are increasingly used to help users analyze messy spreadsheets and workbooks, from answering questions over uploaded files to generating code, summaries, and visualizations. These systems are often evaluated by the corre...

📖 Read original article


375. CAPO: Constraint-Aware Prompt Optimization for LLM Agents ​

Author: Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16068v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise prompts and solution path...

📖 Read original article


376. OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation ​

Author: Wei Wu, Xiang Wang, Hongze Leng, Qingye Min, Junxing Zhu, Junqiang Song
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16070v1 Announce Type: cross Abstract: Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. Physics-based ocean forecasting models impose prohibitive computational costs, while existing deep learning approaches predomina...

📖 Read original article


377. Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization ​

Author: Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16072v1 Announce Type: cross Abstract: Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a...

📖 Read original article


378. Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner ​

Author: Simon Lam-Muir
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16085v1 Announce Type: cross Abstract: Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. These quantities need not identify the same event. We study a 30M-parameter recurrent-depth relat...

📖 Read original article


379. AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting ​

Author: Xiachong Lin, Du Yin, Hao Xue, Wen Hu, Imran Razzak, Arian Prabowo, Matthew Amos, Flora D. Salim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16098v1 Announce Type: cross Abstract: Multivariate time-series forecasting faces a structural dilemma: sharing one temporal predictor across variables is parameter-efficient but forces heterogeneous variables through an identical history-to-future map, whereas learning an independent pre...

📖 Read original article


380. RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction ​

Author: Mianzhi Liu, Fan Xiao, Zhiliang Yu, Huayang Huang, Yuke Li, Yi Yang, Wenbo Liu, Yu Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16111v1 Announce Type: cross Abstract: Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established ...

📖 Read original article


381. TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening ​

Author: Dong Chen, Kenneth M. C. Cheung
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.16122v1 Announce Type: cross Abstract: Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional exper...

📖 Read original article


382. Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System ​

Author: Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Guti'errez Gait'an, Eduardo Tovar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16142v1 Announce Type: cross Abstract: UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities. The latency of video transm...

📖 Read original article


383. A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis ​

Author: Unai Agirre, Imanol Jerico, Felipe Casta~no, Andrea Venturi, Francesco Zola
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2608.16158v1 Announce Type: cross Abstract: Phishing remains a persistent and evolving cybersecurity threat, with attack volumes reaching record levels. This growth is driven by the industrialization of phishing through widely available phishing kits and reusable templates, which enable cyberc...

📖 Read original article


384. Digital Twin Degradation: Detecting Cyber Physical Attacks via Temporal Inconsistencies ​

Author: Konstantinos E. Kampourakis, Vasileios Gkioulos, Sokratis Katsikas
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.16159v1 Announce Type: cross Abstract: Digital Twins (DTs) are increasingly used to monitor and analyze Cyber Physical Systems (CPS). However, in adversarial environments, the fidelity of a DT cannot be assumed. Communication delays, data manipulation, sensor degradation, or partial infor...

📖 Read original article


385. Domain-Specific Text Embedding Models for Entity Resolution ​

Author: Khajesh Sapram, Srivardhani Raju, Kishore Konda
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2608.16161v1 Announce Type: cross Abstract: General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same real-world business or person. This limitation affects applications such as entity resol...

📖 Read original article


386. QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents ​

Author: Heng Wang, Yifei Li, Lingling Zhang, Pengyu Li, Xinyu Che, Xinyu Zhang, Zesheng Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16168v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferences may be distributed across time, change with context, and conflict w...

📖 Read original article


387. Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm ​

Author: Hidayet Aksu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.16177v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent e...

📖 Read original article


388. Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations ​

Author: Jun He, Deying Yu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.16178v1 Announce Type: cross Abstract: Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant context across billions of log lines. As autonomous AI agents become primary operational consumers, feeding th...

📖 Read original article


389. MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems ​

Author: Wei-Hao Chen, Weixi Tong, Yuan Tian, Chenglong Wang, Tianyi Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.16181v1 Announce Type: cross Abstract: Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual effort, it ...

📖 Read original article


390. Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics ​

Author: Bozhou Chen, Yongyi Wang, Hanyu Liu, Xionghui Yang, Wenxin Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16182v1 Announce Type: cross Abstract: Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously unstable. Existing studies often attribute instability to isolated factors such as overestimation bias or represen...

📖 Read original article


391. LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents ​

Author: Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16185v1 Announce Type: cross Abstract: LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-mat...

📖 Read original article


392. Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline ​

Author: Mikhail Surikov
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE

arXiv:2608.16187v1 Announce Type: cross Abstract: AI-assisted development tools generate vulnerable code at significant rates, yet few automated mechanisms exist to detect, enrich, fix, and verify security issues at development velocity, particularly ones that ground remediation in real-world threat...

📖 Read original article


393. Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology ​

Author: Fabian Gr"oger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16198v1 Announce Type: cross Abstract: Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall ou...

📖 Read original article


394. HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction ​

Author: Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets remain fundamentally limited: internet-scale video data lack precise physical stat...

📖 Read original article


395. STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering ​

Author: Xinlong Dai, Jinchuan Zhang, Lei Gao, Xinzhe Hu, Yuefeng He, Hui Gao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16224v1 Announce Type: cross Abstract: By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro-symbolic systems continue to rely on LLMs for both semantic interpre...

📖 Read original article


396. A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation ​

Author: Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Guangyu Wu
Published: 8/18/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.16233v1 Announce Type: cross Abstract: Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable contrasts and restoring degraded acquisitions. Across ten completion tas...

📖 Read original article


397. Software Engineering for AI-driven Building Operation ​

Author: Philipp Zech, Sascha Hammes, Johannes Weninger, J"urgen Pannosch, Gernot Steidl
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.16237v1 Announce Type: cross Abstract: Building operations are energy-inefficient. Artificial Intelligence (AI)-driven control systems promise benefits through optimization and predictive control, but deploying them in real buildings reveals a significant software engineering (SE) challen...

📖 Read original article


398. CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills ​

Author: Mingxiao Liu, Zhoumian Jiang, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.16246v1 Announce Type: cross Abstract: Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fa...

📖 Read original article


399. Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection ​

Author: Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, Jianfu Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16259v1 Announce Type: cross Abstract: The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often gen...

📖 Read original article


400. Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic ​

Author: Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16273v1 Announce Type: cross Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effec...

📖 Read original article


401. Decoupled Temporal Encoding for Generative Recommendation ​

Author: Pengfei Jia, Jingjian Wang, Jingmao Li, Ge Zhang, Feng Shi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.16274v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences. Most positional encoding methods are inherited from natural language processing a...

📖 Read original article


402. Audio-Visual Segmentation via Depth-Guided Collaborative Modeling ​

Author: Zhaojin Fu, Yuyang Hong, Qi Yang, Zili Wang, Kun Ding, Shiming Xiang, Bin Fan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16285v1 Announce Type: cross Abstract: Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-comp...

📖 Read original article


403. Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps ​

Author: Zirui Song, Yuye Zhu, Yang Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.16309v1 Announce Type: cross Abstract: Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations...

📖 Read original article


404. Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning ​

Author: Ao Shen, Yongheng Zhang, Yinghui Li, Manning Wang, Di Yin, Xing Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.16316v1 Announce Type: cross Abstract: Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to smaller, ...

📖 Read original article


405. Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026 ​

Author: Marina Lepp, Joosep Kaimre
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.PF

arXiv:2608.16318v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code. However, their performance on authentic object-oriented programming (OOP) asses...

📖 Read original article


406. Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning ​

Author: Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16333v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially...

📖 Read original article


407. SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection ​

Author: Tiancheng Zhang, Mengmeng Wang, Yan Gao, Xiangjie Kong, Guojiang Shen, Jiaxin Du
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16338v1 Announce Type: cross Abstract: Video lane detection requires predictions that remain stable across frames, yet severe vehicle occlusions can break temporal cues. In streaming recurrent models, corrupted observations may enter the hidden state and produce errors that persist into l...

📖 Read original article


408. HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals ​

Author: Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16353v1 Announce Type: cross Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable truthfulness signals in their inte...

📖 Read original article


409. MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories ​

Author: Lauri Lov'en, Jaakko Sauvola, Jukka Riekki, Sasu Tarkoma
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.MA

arXiv:2608.16357v1 Announce Type: cross Abstract: Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased two ways, link related facts held apart, or reconcile contradictory knowledge wi...

📖 Read original article


410. OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations ​

Author: Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola, Elisa Carli, Diego Fernandez Prieto, Marie-Helene Rio
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.16373v1 Announce Type: cross Abstract: Despite comprising over 70% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere.Understanding the global ocean requires jointly observing its surface and subsurface structure, yet no standar...

📖 Read original article


411. Coverage-Maximizing Multinomial Subset Routing under Operational Constraints ​

Author: Quan Zhou, Yiyan Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16375v1 Announce Type: cross Abstract: We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial routing policy instead of a deterministic subset of experts. At each round, the learner samples $M$ experts i.i.d....

📖 Read original article


412. Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation ​

Author: Qinghui Liu, Jon Andr'e Ottesen, Atle Bj{\o}rnerud, Kyrre Eeg Emblem
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16377v1 Announce Type: cross Abstract: Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-level overlap. Still, most pipelines are trained and post-processed for voxel overlap alone. In particular, the ...

📖 Read original article


413. Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine ​

Author: Marta Sumyk, Oleksandr Kosovan, Iryna Voitsitska
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16380v1 Announce Type: cross Abstract: Monitoring war-induced damage to agricultural land in Ukraine is important for understanding threats to food security, environmental stability, and post-war recovery. However, the development of computer-vision systems for satellite-based damage anal...

📖 Read original article


414. Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics ​

Author: Luca Foppiano
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16390v1 Announce Type: cross Abstract: PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per document, and none decomposes its token total. The two units diverge sharply. On CC-MAIN-2021-31-PDF-UNTRUNCAT...

📖 Read original article


415. Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs ​

Author: Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.16391v1 Announce Type: cross Abstract: As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize host...

📖 Read original article


416. Towards Risk-free AI Agent Deployment ​

Author: Yintong Huo, Rangeet Pan, Abhik Roychoudhury
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.16411v1 Announce Type: cross Abstract: LLM-based agents are rapidly moving from research prototypes into the core business processes of organizations, but these agents pose deployment risks to security, compliance, and functionality. In this article, we argue that risk-free deployment mus...

📖 Read original article


417. PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data ​

Author: Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.QM

arXiv:2608.16419v1 Announce Type: cross Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, w...

📖 Read original article


418. Visualizing Uncertainty-to-Action Composition for Human Oversight ​

Author: Chisom Anyabolu, Akshat Dubey, Georges Hattab
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.16428v1 Announce Type: cross Abstract: Artificial intelligence systems often disclose uncertainty, yet they rarely make clear what response that uncertainty should trigger. Most uncertainty visualizations encode uncertainty in model outputs, leaving users to discern the most appropriate c...

📖 Read original article


419. Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos ​

Author: Mohamed Afham, Christoph Reich, Oliver Hahn, Daniel Cremers, Stefan Roth
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16457v1 Announce Type: cross Abstract: Procedure planning seeks to estimate a sequence of actions to transition from an observed initial state to a given goal state. Current procedure planning approaches directly predict action sequences from latent representations using feed-forward neur...

📖 Read original article


420. A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes ​

Author: Sourya Joyee De, Abdessamad Imine
Published: 8/18/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, cs.CE, cs.CY

arXiv:2608.16461v1 Announce Type: cross Abstract: Central Bank Digital Currency (CBDC)-based welfare schemes may be potentially privacy invasive as they process significant volumes of beneficiary personal data and lead to privacy harms such as surveillance, discrimination and stigmatization. Such we...

📖 Read original article


421. A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling ​

Author: Jingyi Chen, Chaofan Bu, Shibo Yan, Xuesong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.16470v1 Announce Type: cross Abstract: We examine the worldwide trend of mandatory labeling of generative artificial intelligence(GenAI) as a reactive, symbolic form of legislation triggered by technological panic and institutional responses. From a technical perspective, this study demon...

📖 Read original article


422. A Two-Stage Learning PINN Approach for Solving the Inverse Problem of the 1D Porous Medium Equation ​

Author: Noura Al Helwani, Sophie Moufawad, Nabil Nassif
Published: 8/18/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2608.16475v1 Announce Type: cross Abstract: The Porous Medium Equation (PME), given by $u_t = \Delta(u^m)$ for $m > 1$, is a degenerate nonlinear parabolic partial differential equation that arises in various physical applications such as fluid flow in porous media, heat transfer in plasmas, a...

📖 Read original article


423. RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning ​

Author: Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16480v1 Announce Type: cross Abstract: We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video id...

📖 Read original article


424. Graph Machine Learning: An Opportunity for Power Systems ​

Author: Martin Sadric, Sebastian P"utz, Christian Nauck, Veit Hagenmeyer, Frank Hellmann, Dirk Witthaut, Benjamin Sch"afer
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.SY, eess.SY

arXiv:2608.16494v1 Announce Type: cross Abstract: Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. Addressing these challenges traditionally ...

📖 Read original article


425. NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation ​

Author: Cong Zhao, Shuai Tian, Xu Zhang, Baocheng Ni, Xinguo Song, Xueying Sun, Shu Jiang, Shouchang Yang, Bo Tang, Jin Deng, Ge Zhu, YongCheng Wang, Jin Xu, Ri Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.16503v1 Announce Type: cross Abstract: Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that...

📖 Read original article


426. MLLM-Guided Semantic Correction for Text-to-Video Generation ​

Author: Junhao Chen, Zheqi Lv, Keting Yin, Shengyu Zhang, Zhou Zhao, Feiyang Chen, Xinyu Duan, Baoxing Huai, Fei Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16513v1 Announce Type: cross Abstract: Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actio...

📖 Read original article


427. Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans ​

Author: Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.HC, cs.MM

arXiv:2608.16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of ...

📖 Read original article


428. When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation ​

Author: Haolin Jin, Pengyue Yang, Huaming Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16515v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a source trust problem: retrieved context may be useful, irrelevant, or even misleading. Existing RAG systems oft...

📖 Read original article


429. Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization ​

Author: Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais, Philip J. B. Jackson, Jiankang Deng, Ismail Elezi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, eess.AS

arXiv:2608.16539v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio ...

📖 Read original article


430. VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience ​

Author: Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
Published: 8/18/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.16544v1 Announce Type: cross Abstract: Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolution methods primarily revise skills using execution trajectories collected from current tasks, leaving the evol...

📖 Read original article


431. Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity ​

Author: Jiaqi Yao, Julia Kowal
Published: 8/18/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2608.16612v1 Announce Type: cross Abstract: An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practi...

📖 Read original article


432. Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning ​

Author: Peng Du, Kiran Kamble, Rakshith Vasudev, Zhizhuo Yang, Rohith Nadimpally, Arjun Krishna, Waseem Alshikh, Daniel M. Bikel
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16620v1 Announce Type: cross Abstract: Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synthetic tool-u...

📖 Read original article


433. HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes ​

Author: Yujia Li, Yiqun Zhang, Zihan Cheng, Yijie Huang, Tenglong Ye, Zihan Wang, Xiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16622v1 Announce Type: cross Abstract: Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked target or its supporting evidence. We therefore extend harmful meme detec...

📖 Read original article


434. When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness ​

Author: Mahdi Dhaini, Adam Dejl, Juraj Vladika, Volkan "Ozer, Barbara Plank, Gjergji Kasneci
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16627v1 Announce Type: cross Abstract: Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it remains unclear how different types of NLEs compare in their effects on ...

📖 Read original article


435. Toward Better Assessment of LLMs' Performance in Clinical Error Detection ​

Author: Yifan Zhang, Rahmatollah Beheshti
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.16643v1 Announce Type: cross Abstract: Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are t...

📖 Read original article


436. Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents ​

Author: Zhijian Li, Chao Ren, Peijin Wang, Xian Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.16651v1 Announce Type: cross Abstract: Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations. However, conventional planners often rely on predefined maps and fixed environmental assumptions, limiting their adaptability in dynami...

📖 Read original article


437. X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization ​

Author: Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2608.16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localizatio...

📖 Read original article


438. Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning ​

Author: Daniel Nowak Assis, Jean Paul Barddal, Fabr'icio Enembreck
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16659v1 Announce Type: cross Abstract: Ensembles of decision trees are well-established methods for data stream classification. In ensemble learning, Hoeffding Trees are widely adopted as base learners, performing periodic split attempts according to the Hoeffding bound. Recent studies, h...

📖 Read original article


439. Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL ​

Author: Yi Ai
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.16663v1 Announce Type: cross Abstract: Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In enterprise schemas, SQL can execute successfully while using the wrong relationship role or aggregation grain. We...

📖 Read original article


440. Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank ​

Author: Shanwen Wang, Xin Sun, Danfeng Hong, Junyu Dong, Patrick Le Callet
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16681v1 Announce Type: cross Abstract: Although semi-supervised semantic segmentation ($\text{S}^4$) utilizes abundant unlabeled data to reduce manual labeling burdens, independent training of labeled and unlabeled data causes the former to dominate, which severely degrades pseudo-label q...

📖 Read original article


441. UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures ​

Author: Homa Esfahanizadeh, Matin Mortaheb, Jinfeng Du, Harish Viswanathan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, cs.MM, math.IT

arXiv:2608.16696v1 Announce Type: cross Abstract: Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidth, latency, and energy budgets. Because the task driving downstream decisions evolves over time, a task-specifi...

📖 Read original article


442. Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors ​

Author: Hang Zhang, Kaifeng Zhang, Yixiao Ma, Weijie Xu, Ye Zhu, Kai Ming Ting
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16700v1 Announce Type: cross Abstract: Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise their legal right to have their data $D_f$ removed from a machine learning model. This process is typically ac...

📖 Read original article


443. Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors ​

Author: David Eric Austin, Kaheer Suleman, Jackie Chi Kit Cheung
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16707v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitatio...

📖 Read original article


444. MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter ​

Author: Vignesh Nagarajan, Sriram Venkatapathy
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.16709v1 Announce Type: cross Abstract: A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number into readable prose can quietly add claims the model never made. MIRROR is a research prototype built to separa...

📖 Read original article


445. Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI ​

Author: Chiara Tappermann, Steffen Renisch, Lars Ole Schwen, Hans Meine, Horst K. Hahn, Eike Petersen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16725v1 Announce Type: cross Abstract: Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdevel...

📖 Read original article


446. GoalEvolve: From Handcrafted Algorithm Priors to Goal-Driven Evolution of Physical Design Algorithms ​

Author: Haixu Liu, Lei Zhou, Yuhao Ren, Yumao Wu, Zhiang Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2608.16733v1 Announce Type: cross Abstract: Physical design algorithms operate within tightly coupled, multi-stage optimization flows, where stage-local gains may vanish or induce downstream degradation. Existing program-evolution frameworks often rely on stage-local objectives or undifferenti...

📖 Read original article


447. TDD-Agent: Test-Driven Reasoning for Code Generation ​

Author: Hongyue Yu, Kefan Li, Jiakun Li, Hongzheng Chai, Yuan Yuan, Rui He, Junyi Wei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.16742v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which lim...

📖 Read original article


448. Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments ​

Author: Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16747v1 Announce Type: cross Abstract: Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactua...

📖 Read original article


449. TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation ​

Author: Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16765v1 Announce Type: cross Abstract: Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to...

📖 Read original article


450. Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis ​

Author: Reza Fayyazi, Michael Zuzak, Shanchieh Jay Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.16775v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cyber...

📖 Read original article


451. Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching ​

Author: Ye Lu, Shen Wang, Zhaoyang Zhang, Yihan Yan, Li Liu, Runze Liu, Fanghui Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR, cs.MM

arXiv:2608.16791v1 Announce Type: cross Abstract: Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic...

📖 Read original article


452. Neurosymbolic Embodied Agents ​

Author: Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL

arXiv:2608.16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-ho...

📖 Read original article


453. Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot ​

Author: Hui Mao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CE, cs.AI

arXiv:2608.16795v1 Announce Type: cross Abstract: Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates que...

📖 Read original article


454. UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation ​

Author: Rongcheng Lin, Yan Sun, Jamey Zhang, Guanglei Xiong, Ivan Ji, Xianjie Chen, Shujian Bu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.16797v1 Announce Type: cross Abstract: Industrial recommenders rely on two model families that have evolved largely independently: feature-interaction models over multi-field user/item features, and sequential models over user-behavior histories. Production systems couple them only loosel...

📖 Read original article


455. ClawGym II: Exploring Black-Box RL on Agent Harness ​

Author: Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to l...

📖 Read original article


456. Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models ​

Author: Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Sixue Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.16805v1 Announce Type: cross Abstract: Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metr...

📖 Read original article


457. When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents ​

Author: Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.16806v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code generation, driving their gradual evolution from text generation models into the core of agents capable of percei...

📖 Read original article


458. CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated? ​

Author: Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleator...

📖 Read original article


459. Model Hypnosis: Strong control of AI via additive subliminal effects ​

Author: Enric Boix-Adsera, Benedict Tessler
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16834v1 Announce Type: cross Abstract: We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occ...

📖 Read original article


460. HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL ​

Author: Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.16837v1 Announce Type: cross Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality an...

📖 Read original article


461. Proteus: Incremental Memory Activation for Long-Context Sequence Modeling ​

Author: Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.16844v1 Announce Type: cross Abstract: The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throug...

📖 Read original article


462. Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text ​

Author: Benjamin Belay
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.16868v1 Announce Type: cross Abstract: A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal stat...

📖 Read original article


463. AutoSR: Automatic Symbolic Regression by Searching Research States ​

Author: Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SC, cs.AI, cs.LG, cs.NA, math.NA

arXiv:2608.16876v1 Announce Type: cross Abstract: We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numeri...

📖 Read original article


464. Improving the matrix multiplication exponent with modern optimization and AlphaEvolve ​

Author: Emilien Dupont, Marvin Eisenberger, Borislav Kozlovskii, Abbas Mehrabian, Francisco J. R. Ruiz, Abigail See, Renfei Zhou, Josh Alman, Virginia Vassilevska Williams, Matej Balog
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DS, cs.AI, cs.CC, cs.LG

arXiv:2608.16884v1 Announce Type: cross Abstract: The current best bounds on the matrix multiplication exponent $\omega$ are obtained through a refinement of the laser method called combination loss analysis (Duan et al., 2022; Williams et al., 2024; Alman et al., 2025). In this note, we address the...

📖 Read original article


465. Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory ​

Author: Bingxin Xu, Yuzhang Shang, Emilio Ferrara
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.16889v1 Announce Type: cross Abstract: Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: errors compound beyond the policy's ability to correc...

📖 Read original article


466. mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA ​

Author: Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Yuxuan Zhao, Zehua Xie, Jin Ma, Ying Shan, Weiming Hu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2411.15041v2 Announce Type: replace Abstract: Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOSEEK and Encyclopedic-VQA, due to their limited and frozen knowledge scope, often leading to ambiguous and ina...

📖 Read original article


467. Evidence of conceptual mastery in the application of rules by Large Language Models ​

Author: Jos'e Luiz Nunes, Guilherme FCF Almeida, Brian Flanagan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.HC

arXiv:2503.00992v2 Announce Type: replace Abstract: In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure to match the diversity of thought generated by LLMs to that observed in a human sample. We then conducted two ...

📖 Read original article


468. SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling ​

Author: Shixuan Sun, Siyuan Liang, Jianjie Huang, Jingzhi Li, Xiaochun Cao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2508.09105v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) significantly improve the knowledge coverage and contextual understanding of Large Language Models (LLMs) by introducing external knowledge sources. Howev...

📖 Read original article


469. Calibrated Generative AI as Meta-Reviewer: A Systemic Functional Linguistics Discourse Analysis of Reviews of Peer Reviews ​

Author: Gabriela C. Zapata, Bill Cope, Mary Kalantzis, Duane Searsmith
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2509.15035v2 Announce Type: replace Abstract: This study investigates the use of generative AI to support formative assessment through machine generated reviews of peer reviews in graduate online courses in a public university in the United States. Drawing on Systemic Functional Linguistics an...

📖 Read original article


470. The Fragility of Strategic Thinking in Large Language Models ​

Author: Enric Junque de Fortuny, Veronica Roberta Cappelli
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.GT

arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation. However, can we trust LLMs to think strategically in complex situations...

📖 Read original article


471. Budget-Aware Tool Use Enables Effective Agent Scaling ​

Author: Tengxiao Liu, Zifeng Wang, Jin Miao, I-Hung Hsu, Jun Yan, Jiefeng Chen, Rujun Han, Fangyuan Xu, Yanfei Chen, Ke Jiang, Samira Daruki, Yi Liang, William Yang Wang, Tomas Pfister, Chen-Yu Lee
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.17006v2 Announce Type: replace Abstract: Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thinking in tokens but also acting via tool calls that directly constrain environmental interaction. However, we...

📖 Read original article


472. MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration ​

Author: Yakun Zhu, Yutong Huang, Shengqian Qin, Zhongzhen Huang, Shaoting Zhang, Xiaofan Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.23049v2 Announce Type: replace Abstract: Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-stage process, requiring proactive EHR data acquisition, scenario-dependent calculator selection, and multi-s...

📖 Read original article


473. Agentic Test-Time Scaling for WebAgents ​

Author: Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2602.12276v2 Announce Type: replace Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on agentic, multi-step tasks remains less well-understood: small per-step errors can compound over long horizon...

📖 Read original article


474. The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents ​

Author: Shrey Shah, Levent Ozgur
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2603.00801v2 Announce Type: replace Abstract: Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial ranking ...

📖 Read original article


475. ML-AutoResearch: Training Machine Learning Research Agents with Automatically Generated Environments ​

Author: Ziyang Cai, Amir Saeidi, Harkirat Behl
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.17216v2 Announce Type: replace Abstract: With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal. However, training agents to autonomously execute the engineering-heavy labor of machine learning (ML) research requires massive, process-level ...

📖 Read original article


476. FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification ​

Author: Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu, Libin Zheng, Wei Liu, Shaowu Pan, Shimin Di, Min-Ling Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.04074v4 Announce Type: replace Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We present FactReview, an audit pipeline that extracts review-relevant claims, grounds ...

📖 Read original article


477. Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability ​

Author: Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and find that cross-dom...

📖 Read original article


478. An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing ​

Author: Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou, Malcolm Yoke Hean Low
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2605.13221v2 Announce Type: replace Abstract: In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational t...

📖 Read original article


479. SAPO: Step-Aligned Policy Optimization for Reasoning-Based Generative Recommendation ​

Author: Zaiyi Zheng, Liang Wu, Guanghui Min, Yaochen Zhu, Liangjie Hong, Chen Chen, Jundong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.17648v2 Announce Type: replace Abstract: Generative recommendation treats next-item prediction as autoregressive item-identifier generation. Specifically, items are encoded as semantic identifiers (SIDs), which are short coarse-to-fine token sequences whose early tokens capture broad sema...

📖 Read original article


480. BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization ​

Author: Zhengyang Ni, Feng Yan, Yu Guo, Fei Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.GR

arXiv:2605.26182v2 Announce Type: replace Abstract: Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability. Existing brick generation methods either rely on heuristic...

📖 Read original article


481. Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models ​

Author: Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho, Tuhin Kundu, Chris Dongjoo Kim, Zelun Luo, Jieyu Zhang, Linda Shapiro, Ranjay Krishna
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.03988v3 Announce Type: replace Abstract: Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoi...

📖 Read original article


482. A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems ​

Author: Pollob Chandra Ray, Sabah Binte Noor, Fazlul Hasan Siddiqui
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.14582v2 Announce Type: replace Abstract: Efficient route optimization play a vital role in ensuring both safety and punctuality in railway operations. It is very crucial particularly in heterogeneous multi-gauge railway networks with varying train speed, stopping pattern, infrastructure c...

📖 Read original article


483. A Machine-Learned Comorbidity Index ​

Author: Suleman Baloch, Kishlay Jha, Alberto M. Segre, Philip M. Polgreen, Bijaya Adhikari
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.17450v2 Announce Type: replace Abstract: Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outco...

📖 Read original article


484. Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries ​

Author: Ylli Prifti, Pasquale De Meo, Alessandro Provetti
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.PL, cs.SE

arXiv:2606.20615v3 Announce Type: replace Abstract: AI agents now act as first-class members of the software development lifecycle, but the instruments teams use to direct them enforce nothing: process encoded in prompts is flexible but unenforceable, while workflow formalisms are enforceable but do...

📖 Read original article


485. PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows ​

Author: Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.06008v3 Announce Type: replace Abstract: While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and lon...

📖 Read original article


486. Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns ​

Author: Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.11621v2 Announce Type: replace Abstract: Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) wh...

📖 Read original article


487. Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization ​

Author: David Aguado, Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.00270v2 Announce Type: replace Abstract: Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solving routing problems such as the Traveling Salesman Problem (TSP). However, the generalization capabilities of...

📖 Read original article


488. Where did the ambiguity go? Examining how multimodal models interpret polysemous words ​

Author: Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2608.00410v2 Announce Type: replace Abstract: Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning,...

📖 Read original article


489. DiffImaginE: Imagine to Verify Entity Types with Diffusion ​

Author: Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.03025v4 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visua...

📖 Read original article


490. RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation ​

Author: Shuhao Yan, Changhao He, Peng Hu, Xi Peng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05714v2 Announce Type: replace Abstract: Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, external...

📖 Read original article


491. Runtime Observability for Heterogeneous Attention Memory ​

Author: Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05863v2 Announce Type: replace Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract th...

📖 Read original article


492. GSBF: Gaussian Splatting for Environment-Aware Beamforming ​

Author: Yijie Bian, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT

arXiv:2608.05896v2 Announce Type: replace Abstract: Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur s...

📖 Read original article


493. Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents ​

Author: Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka, Jie M. Zhang, Michael Luck, Elizabeth Black, Daniele Quercia
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.08601v2 Announce Type: replace Abstract: To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job-specific risks introduced by agents. To address this gap, we make t...

📖 Read original article


494. Context Is Not Authority: Structured Runtime Governance for Financial Market Agents ​

Author: Rui Tang, Qiangqiang Liu, Yichi Zhang, Youwei Yang, Xi Chen, Chen Dong
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, stat.ML

arXiv:2608.09025v2 Announce Type: replace Abstract: Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present SAGE-Fin, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, ...

📖 Read original article


495. Automating and Scaling Behavioral Scientific Research on AI Agents ​

Author: Soo Yong Lee, Jongha Lee, Jaewan Chun, Hyunjin Hwang, Fanchen Bu, Ziv Ben-Zion, Taekwan Kim, Denny Borsboom, Jaemin Yoo, Kijung Shin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.10030v2 Announce Type: replace Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system t...

📖 Read original article


496. SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure ​

Author: Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.11079v2 Announce Type: replace Abstract: Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather tha...

📖 Read original article


497. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research ​

Author: Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.11216v2 Announce Type: replace Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomo...

📖 Read original article


498. When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs ​

Author: Utkarsh Bahuguna
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.11403v2 Announce Type: replace Abstract: Self-consistency via majority vote reduces per-problem accuracy on most GPQA Diamond problems for small instruction-tuned models: 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B. The obvious remedy is a verifier-free confidence gate. This...

📖 Read original article


499. Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation ​

Author: Liming Liu, Mingze Wang, Tuo Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.12385v2 Announce Type: replace Abstract: As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential...

📖 Read original article


500. Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension ​

Author: Alison R. Panisson, Maria Eduarda W. M. Vianna, Italo Firmino da Silva, Heitor Henrique da Silva, Rafaela Fernandes Savaris, Bernardo Pandolfi Costa, Martin Augusto Gagliotti Vigil, Jim Lau, Agenor Hentz, Andr'ea Sabedra Bordin, Alexandre Leopoldo Gon\c{c}alves, Roberto Rodrigues-Filho
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.13447v2 Announce Type: replace Abstract: Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artific...

📖 Read original article


501. MobileMem: Learning from a Year of Mobile Experiences ​

Author: Xinle Deng, Yida Xue, Xiangyuan Ru, Yijun Chen, Buqiang Xu, Mingjun Mao, Xinjie Liu, Haoming Xu, Shuofei Qiao, Mengru Wang, Chen Jiang, Yuchen Eleanor Jiang, Lizhong Wang, Jason Wang, Li Zeng, Haofen Wang, Guilin Qi, Huajun Chen, Ningyu Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, cs.MM

arXiv:2608.13606v2 Announce Type: replace Abstract: The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-...

📖 Read original article


502. From Monte Carlo to neural networks approximations of boundary value problems ​

Author: Lucian Beznea, Iulian Cimpean, Oana Lupascu-Stamate, Ionel Popescu, Arghir Zarnescu
Published: 8/18/2026, 4:00:00 AM
Categories: math.PR, cs.AI, cs.LG, cs.NA, math.AP, math.NA

arXiv:2209.01432v4 Announce Type: replace-cross Abstract: In this paper we study probabilistic and neural network approximations for solutions to Poisson equation subject to Holder data in general bounded domains of $\mathbb{R}^d$. We aim at two fundamental goals. The first, and the most important, ...

📖 Read original article


503. ShadowNet for Data-Centric Quantum System Learning ​

Author: Yuxuan Du, Yibo Yang, Tongliang Liu, Zhouchen Lin, Bernard Ghanem, Dacheng Tao
Published: 8/18/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2308.11290v2 Announce Type: replace-cross Abstract: Understanding the dynamics of large quantum systems is hindered by the curse of dimensionality. Statistical learning offers new possibilities in this regime through neural network protocols and classical shadows, while both methods have limit...

📖 Read original article


504. A Bi-directional Multi-solution Scalable Grover Search Algorithm ​

Author: Debanjan Konar, Zain Hafeez, Vaneet Aggarwal
Published: 8/18/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2404.15616v2 Announce Type: replace-cross Abstract: Grover's search algorithms, including various Partial Grover Searches (PGS), suffer from scaling issues when multiple solutions are sought, as the number of iterations scales with the number of solutions or marked states, making implementatio...

📖 Read original article


505. DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Variations ​

Author: Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao, Qingming Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2405.07780v3 Announce Type: replace-cross Abstract: This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced. We argue that the variation in these distributions can be broken down hierarchica...

📖 Read original article


506. TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability ​

Author: Fengji Ma, Hei Victor Cheng, Chenxing Li, Li Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2405.17678v2 Announce Type: replace-cross Abstract: Achieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, especially under large adversarial perturbations. Through empirical analyses, we identify three critical yet...

📖 Read original article


507. MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4 ​

Author: Vahid Azizi, Fatemeh Koochaki
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2406.00971v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs, which process image and text modalities simultaneously, have demonstrated the ability to learn and unde...

📖 Read original article


508. Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study ​

Author: Jerson Francia, Derek Hansen, Benjamin Schooley, Matthew Taylor, Shydra Valynn Murray, Rebekah Cornelius, Greg Snow
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2406.13049v3 Announce Type: replace-cross Abstract: Personalized phishing is difficult to defend against because messages can be tailored to a target's work, interests, and social context. Large language models may make such tailoring faster and easier, but it remains unclear whether messages ...

📖 Read original article


509. Quantum Large Language Models via Tensor Network Disentanglers ​

Author: Borja Aizpurua, Fernando Loren, Saeed S. Jahromi, Sukhbinder Singh, Roman Orus
Published: 8/18/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2410.17397v2 Announce Type: replace-cross Abstract: We introduce a framework for seamlessly integrating quantum computing into pretrained large language models (LLMs). The key idea is to construct a hybrid quantum-classical representation that exactly reproduces the original model, providing a...

📖 Read original article


510. MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization ​

Author: Zelin Zang, Yuhao Wang, Jinlin Wu, Hong Liu, Yue Shen, Zhen Lei, Stan Z. Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.19504v3 Announce Type: replace-cross Abstract: Dimensionality reduction (DR) plays a crucial role in various fields, including data engineering and visualization, by simplifying complex datasets while retaining essential information. However, achieving both high DR accuracy and strong exp...

📖 Read original article


511. Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching ​

Author: Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2412.18911v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs. As an effective approach for DiT acceleration, feature caching methods are designed to cache the fea...

📖 Read original article


512. Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks ​

Author: Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2501.06286v2 Announce Type: replace-cross Abstract: Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models remain vulnerable to irrelevant context. We investigate this evidence bottleneck in the English HotpotQA...

📖 Read original article


513. Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities ​

Author: Qirun Dai, Dylan Zhang, Jiaqi W. Ma, Hao Peng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2501.12147v2 Announce Type: replace-cross Abstract: Selecting appropriate training data is crucial for instruction fine-tuning of large language models (LLMs), which aims to (1) elicit strong capabilities, and (2) achieve balanced performance across different tasks. Influence-based methods sho...

📖 Read original article


514. ConfRetro: a 3D-aware template-free method for enhancing retrosynthesis via molecular conformer information ​

Author: Jiaxi Zhuang, Yu Zhang, Ying Qian, Aimin Zhou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2501.12434v3 Announce Type: replace-cross Abstract: Motivation: Retrosynthesis plays a crucial role in organic synthesis and drug discovery, focusing on identifying a set of reactants capable of synthesizing a target product molecule. Although the existing approaches have shown promising resul...

📖 Read original article


515. Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects ​

Author: Levente Z'olyomi, Tianze Wang, Sofiane Ennadir, Oleg Smirnov, Lele Cao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2502.04899v3 Announce Type: replace-cross Abstract: The proliferation of digital interactions across diverse domains, such as healthcare, e-commerce, gaming, and finance, has resulted in the generation of vast volumes of event stream (ES) data. ES data comprises continuous sequences of timesta...

📖 Read original article


516. DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning ​

Author: Hongye Qiu, Yue Xu, Yi Wang, Meikang Qiu, Wenjie Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2502.11603v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, particularly gender bias, raising fairness concerns. Existing prompt-based debiasing strategies share a key ...

📖 Read original article


517. Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation ​

Author: Giorgio Franceschelli, Mirco Musolesi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG

arXiv:2502.13207v4 Announce Type: replace-cross Abstract: Despite the increasing use of large language models for creative tasks, their outputs often lack diversity. Common solutions, such as sampling at higher temperatures, can compromise the quality of the results. Dealing with this trade-off is s...

📖 Read original article


518. Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching ​

Author: Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
Published: 8/18/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, stat.ME

arXiv:2502.14424v3 Announce Type: replace-cross Abstract: Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate representation learning as Distribution Matching (DM), learning an augmentation-invariant encoder whose induced...

📖 Read original article


519. Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning ​

Author: Liang Lu, Yuan Jiang, Christoph Treude, Shuzheng Gao, Jingyu Xiao, Xiaohong Su, Michael R. Lyu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2503.09020v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been widely adopted in commercial code completion engines, significantly enhancing coding efficiency and productivity. However, even functionally correct LLM-generated code may exhibit non-functional quality ...

📖 Read original article


520. Leveraging Machine Unlearning for Cost-Efficient Preference Alignment ​

Author: Xiaohua Feng, Yuyuan Li, Huwei Ji, Jiaming Zhang, Li Zhang, Tianyu Du, Chaochao Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2504.06659v2 Announce Type: replace-cross Abstract: Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable challenges. These approaches require high-quality datasets of positive preference ...

📖 Read original article


Author: Matthew Dahl, Eric Mart'inez
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2505.02763v2 Announce Type: replace-cross Abstract: One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without calling for much discretion. Yet it remains an open question how well AI models actually perform on suc...

📖 Read original article


522. WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales ​

Author: Drew Prinster, Xing Han, Anqi Liu, Suchi Saria
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2505.04608v5 Announce Type: replace-cross Abstract: Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address an...

📖 Read original article


523. Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair ​

Author: David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2505.07372v3 Announce Type: replace-cross Abstract: This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs). Current APR systems are constrained by the limited availability of high-quality tra...

📖 Read original article


524. DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models ​

Author: Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2505.14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable their...

📖 Read original article


525. PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models ​

Author: Chang Liu, Bohao Zhao, Jingtao Ding, Huandong Wang, Yong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.23863v3 Announce Type: replace-cross Abstract: Understanding chaotic dynamics is a fundamental problem across scientific disciplines, including climate science, neuroscience, and fluid dynamics, yet direct experimentation and intervention in such systems are often infeasible. Chaotic syst...

📖 Read original article


526. VirnyFlow: Optimizing ML Pipelines for Accuracy, Fairness, and Stability at Scale ​

Author: Denys Herasymuk, Anastasiia Mozghova, Nazar Protsiv, Vladyslav Sydorak, Julia Stoyanovich
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2506.01584v2 Announce Type: replace-cross Abstract: Developing machine learning (ML) systems for real-world deployment requires navigating context-dependent trade-offs among accuracy, fairness, stability, and other objectives. Existing AutoML frameworks optimize pipelines efficiently, but they...

📖 Read original article


527. LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking ​

Author: Vahid Azizi, Fatemeh Koochaki
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2506.07449v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have driven their adoption in recommender systems through Retrieval-Augmented Generation (RAG) frameworks. However, existing RAG approaches predominantly rely on flat, similarity-based retrieval...

📖 Read original article


528. Contraction-Aware Reinforcement Learning for Nonlinear Control with Statistical Robustness ​

Author: Minjae Cho, Hiroyasu Tsukamoto, Huy T. Tran
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.15700v2 Announce Type: replace-cross Abstract: Control contraction metrics (CCMs)-defined by Riemannian metrics under which a closed-loop system is incrementally exponentially stable-offer a constructive framework for synthesizing contracting policies in nonlinear path-tracking problems. ...

📖 Read original article


529. From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology ​

Author: Zhicheng Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.HC

arXiv:2506.16697v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk o...

📖 Read original article


530. A validity-guided workflow for robust large language model research in psychology ​

Author: Zhicheng Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY

arXiv:2507.04491v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measurement unreliability: ...

📖 Read original article


531. Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges ​

Author: Yidong Jiang, Jiangtong Li, Daiwei Cheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2507.09562v2 Announce Type: replace-cross Abstract: The Segment Anything Model (SAM) has transformed image segmentation by introducing a prompt-based paradigm that enables strong zero-shot generalization. In this framework, prompts serve as a semantic interface between human intent and machine...

📖 Read original article


532. Reprojection-Guided 3D Gaussian Splatting Diffusion for Weakly Supervised Single-Image Normal Estimation ​

Author: Yanxing Liang, Yinghui Wang, Wei Li, Tao Yan, Jiaxing Shen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2508.05950v4 Announce Type: replace-cross Abstract: We propose CLONE, a Continuous Latent Optimization framework for Normal Estimation via 3D Gaussian splatting. The core idea is to construct an image-geometry-image consistency strategy that unifies explicit geometric representation with diffe...

📖 Read original article


533. Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Representation Alignment ​

Author: Yanru Sun, Emadeldeen Eldele, Zongxia Xie, Yucheng Wang, Wenzhe Niu, Qinghua Hu, Chee Keong Kwoh, Min Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.07195v2 Announce Type: replace-cross Abstract: Recent advances have demonstrated that Large Language Models (LLMs) can be effectively adapted for time series forecasting, revealing strong potential beyond natural language tasks. However, their performance remains constrained by two fundam...

📖 Read original article


534. ProteoKnight: Convolution-based Phage Virion Protein Classification and Uncertainty Analysis ​

Author: Samiha Afaf Neha, Md. Ishrak Khan, Abir Ahammed Bhuiyan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2508.07345v2 Announce Type: replace-cross Abstract: \textbf{Introduction:} Accurate prediction of Phage Virion Proteins (PVP) is essential for genomic studies due to their crucial role as structural elements in bacteriophages. Computational tools, particularly machine learning, have emerged fo...

📖 Read original article


535. CulTrace: Tracing Internal Cultural Reasoning in Large Language Models ​

Author: Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awareness in LLMs by analys...

📖 Read original article


536. PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning ​

Author: Yunxiao Wang, Meng Liu, Sicheng Zhao, Lizi Liao, Liqiang Nie
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.09521v3 Announce Type: replace-cross Abstract: Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adopt an appropriate strategy, and respond in a natural, human-like manner. Despite advances in large langu...

📖 Read original article


537. Efficient Code Embeddings from Code Generation Models ​

Author: Daria Kryvosheieva, Saba Sturua, Michael G"unther, Han Xiao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2508.21290v2 Announce Type: replace-cross Abstract: jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. It makes i...

📖 Read original article


538. Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning ​

Author: Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Xuanshan Zhou, Jiayu Yao, Jiafeng Guo, Xueqi Cheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.06461v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex visual environments. While existing enhancement approaches require additional training, rely on external...

📖 Read original article


539. Privacy-Preserving Decentralized Federated Learning via Explainable Adaptive Differential Privacy ​

Author: Fardin Jalil Piran, Zhiling Chen, Yang Zhang, Qianyu Zhou, Jiong Tang, Farhad Imani
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2509.10691v3 Announce Type: replace-cross Abstract: Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sensitive information through inversion, reconstruction, and membership inference attacks. Differential pr...

📖 Read original article


540. Geometrically Constrained and Token-Based Probabilistic Spatial Transformers ​

Author: Johann Schmidt, Sebastian Stober
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.11218v2 Announce Type: replace-cross Abstract: Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such tran...

📖 Read original article


541. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models ​

Author: Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.22415v5 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token hidden state into t...

📖 Read original article


542. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing ​

Author: Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and s...

📖 Read original article


543. DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts ​

Author: Linjin He, Xinda Qi, Dong Chen, Zhaojian Li, Xiaobo Tan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2510.00358v2 Announce Type: replace-cross Abstract: Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control remains challenging due to highly nonlinear dynamics. Existing model-based and bio-inspired controllers rely on simplified assumptions ...

📖 Read original article


544. Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data ​

Author: Usman Akram, Yiyue Chen, Haris Vikalo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2510.04927v2 Announce Type: replace-cross Abstract: Automatic modulation classification (AMC) is a core enabler of cognitive wireless systems, providing spectrum awareness and supporting adaptive communication at the network edge. However, training AMC models on centrally aggregated data incur...

📖 Read original article


545. A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs ​

Author: Chengwei Wu, Xingrui Zhuo, Mingyang Gao, Xinghe Cheng, Zhichao Yan, Jiapu Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.06039v2 Announce Type: replace-cross Abstract: Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language text with verifiable Knowledge Graph (KG) facts. Yet existing Chinese benchmarks primarily assess gene...

📖 Read original article


546. Sleeping Kelly ​

Author: Ben Abramowitz
Published: 8/18/2026, 4:00:00 AM
Categories: q-fin.GN, cs.AI

arXiv:2510.15911v3 Announce Type: replace-cross Abstract: The Sleeping Beauty problem is a problem of imperfect recall that has received considerable attention. One approach to solving the Sleeping Beauty problem is to allow Sleeping Beauty to make decisions based on her beliefs, and then characteri...

📖 Read original article


547. Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing ​

Author: Zan Li, Rui Fan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE

arXiv:2510.17088v3 Announce Type: replace-cross Abstract: Financial anomalies arise from heterogeneous mechanisms - price shocks, liquidity freezes, contagion cascades, and momentum reversals - yet existing detectors produce uniform anomaly scores without revealing which mechanism is failing or wher...

📖 Read original article


548. Retrofit: Continual Learning with Controlled Forgetting for Binary Security Detection and Analysis ​

Author: Yiling He, Junchi Lei, Hongyu She, Shuo Shao, Xinran Zheng, Yiping Liu, Zhan Qin, Lorenzo Cavallaro
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.11439v3 Announce Type: replace-cross Abstract: Binary security has increasingly relied on deep learning to reason about malware behavior and program semantics. However, the performance often degrades as threat landscapes evolve and code representations shift. While continual learning (CL)...

📖 Read original article


549. High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid ​

Author: Even Marius Nordhagen, H{\aa}vard Homleid Haugen, Magnus Sikora Ingstad, Aram Farhad Shafiq Salihi, Thomas Nils Nipen, Ivar Ambj{\o}rn Seierstad, Inger-Lise Frogner, Mariana Clare, Simon Lang, Matthew Chantry, Peter Dueben, J{\o}rn Kristiansen
Published: 8/18/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI

arXiv:2511.23043v2 Announce Type: replace-cross Abstract: We present a probabilistic data-driven weather model providing ensembles of high spatial resolution realizations of 87 variables at arbitrary ensemble size and forecast length. The model uses a global stretched grid, dedicating 2.5 km resolut...

📖 Read original article


550. jina-vlm: Small Multilingual Vision Language Model ​

Author: Andreas Koukounas, Georgios Mastrapas, Florian H"onicke, Sedigheh Eslami, Guillaume Roncari, Han Xiao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2512.04032v4 Announce Type: replace-cross Abstract: We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. The model couples a SigLIP2 vision encoder with a Qwen3 language decoder and mak...

📖 Read original article


551. Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies ​

Author: Mingming Zhang, Na Li, Zhuang Feiqing, Hongyang Zheng, Jiangbing Zhou, Wang Wuyin, Sheng-jie Sun, XiaoWei Chen, Junxiong Zhu, Lixin Zou, Chenliang Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2601.02754v3 Announce Type: replace-cross Abstract: With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertiser environments. The current approaches focus on reinforcement learning (RL) and generative models. Thes...

📖 Read original article


552. The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI ​

Author: Jacob Erickson
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2601.03222v2 Announce Type: replace-cross Abstract: As conversational AI systems become a larger part of the media landscape, they raise questions about whose interests they serve and the risks they may pose to users. These systems do more than provide information: they increasingly offer advi...

📖 Read original article


553. QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging ​

Author: Zhaofeng Zhong, Wei Yuan, Tong Chen, Liang Qu, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.03506v2 Announce Type: replace-cross Abstract: Recent large reasoning models (LRMs) have achieved strong performance on complex reasoning tasks by generating a long chain-of-thought (Long-CoT). However, such lengthy reasoning is often unnecessary for simple queries, leading to additional ...

📖 Read original article


554. Backpropagation-Free Test-Time Adaptation for Lightweight EEG-Based Brain-Computer Interfaces ​

Author: Siyang Li, Jiayi Ouyang, Zhenyao Cui, Ziwei Wang, Tianwang Jia, Feng Wan, Dongrui Wu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2601.07556v3 Announce Type: replace-cross Abstract: Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) face significant deployment challenges due to inter-subject variability, signal non-stationarity, and computational constraints. While test-time adaptation (TTA) mitigates dist...

📖 Read original article


555. AWED-PIPER: Agents, Web Applications & Expert Detectors for Personally Identifiable Information Protection & Fine-grained Named Entity Recognition across 36 languages for 6.6 Billion Speakers ​

Author: Prachuryya Kaushik, Ashish Anand
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2601.10161v3 Announce Type: replace-cross Abstract: Named Entity Recognition (NER) and Personally Identifiable Information (PII) anonymization are critical tasks in Natural Language Processing (NLP) for information extraction and privacy preservation. We introduce AWED-PIPER, an open-source fr...

📖 Read original article


556. Sequential LLM Release Facilitates Manipulation in Regulated Markets ​

Author: Eilam Shapira, Moshe Tennenholtz, Roi Reichart
Published: 8/18/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CL, cs.MA

arXiv:2601.11496v3 Announce Type: replace-cross Abstract: AI agents increasingly mediate bargaining, negotiation and persuasion for people and firms. Such markets extend software-mediated commerce, but add a governance problem: independent model releases change delegates available to participants. G...

📖 Read original article


557. Aletheia: What Makes RLVR For Code Verifiers Tick? ​

Author: Vatsal Venkatkrishna, Indraneil Paul, Iryna Gurevych
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2601.12186v4 Announce Type: replace-cross Abstract: Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generation has lagged behind that of execution feedback due to the pr...

📖 Read original article


558. Robust Privacy: Inference-Stage Privacy through Certified Robustness ​

Author: Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2601.17360v3 Announce Type: replace-cross Abstract: An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representatives of the model's training data. The inference interface thus acts as a side channel for privacy leakag...

📖 Read original article


559. Credit Fairness: Online Fairness In Shared Resource Pools ​

Author: Seyed Majid Zahedi, Rupert Freeman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.OS

arXiv:2601.17944v2 Announce Type: replace-cross Abstract: We study repeated allocation of shared resources among agents with time-varying demands and capped linear utilities. In this setting, independently maximizing the minimum endowment-normalized utility in each round satisfies sharing incentives...

📖 Read original article


560. Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads ​

Author: Chendong Song, Meixuan Wang, Hang Zhou, Hong Liang, Yuan Lyu, Zixi Chen, Yuwei Fan, Zijie Zhou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.21351v4 Announce Type: replace-cross Abstract: Attentio-FFN disaggregation (AFD) is an emerging architecture for LLM decoding that separates state-heavy, KV-cache-dominated Attention computation from stateless, compute-intensive FFN computation, connected by per-step communication. While ...

📖 Read original article


561. SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking ​

Author: Muhammad Taha Mukhtar, Syed Musa Ali Kazmi, Khola Naseem, Muhammad Ali Chattha, Andreas Dengel, Sheraz Ahmed, Muhammad Naseer Bajwa, Muhammad Imran Malik
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.04525v3 Announce Type: replace-cross Abstract: Very-high-resolution remote-sensing imagery provides a scalable basis for delineating informal settlements, but sparse annotations, severe imbalance between informal-settlement and background pixels, and cross-city heterogeneity in urban morp...

📖 Read original article


562. DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter ​

Author: Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng, Yuan Meng, Haojun Jiang, Shaoxuan Xie, Guocai Yao, Alois Knoll, Zhenshan Bing, Xinlong Wang, Zhenguo Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2602.05513v4 Announce Type: replace-cross Abstract: Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of effectively combining these modalities, we propose DECO, a decoupled multimodal diffusion transformer th...

📖 Read original article


563. Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization ​

Author: Matteo Pannacci, Andrea Fanti, Elena Umili, Roberto Capobianco
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.09761v2 Announce Type: replace-cross Abstract: In this work we address the problem of training a Reinforcement Learning agent to follow multiple temporally-extended instructions expressed in Linear Temporal Logic in sub-symbolic environments. Previous multi-task work has mostly relied on ...

📖 Read original article


564. Zero-Shot Instruction Following in RL via Structured LTL Representations ​

Author: Mathias Jackermeier, Mattia Giuri, Jacques Cloete, Alessandro Abate
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.14344v2 Announce Type: replace-cross Abstract: We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training. In this setting, linear temporal logic (LTL) has recently been adopted as a powerful framework fo...

📖 Read original article


565. ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization ​

Author: Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG, math.OC

arXiv:2602.15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibil...

📖 Read original article


566. LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights ​

Author: Kasun Dewage, Marianna Pensky, Suranadi De Silva, Shankadeep Mondal
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17510v2 Announce Type: replace-cross Abstract: We introduce LoRA-CRAFT (\textbf{C}ross-layer \textbf{R}ank \textbf{A}daptation via \textbf{F}rozen \textbf{T}ucker), abbreviated CRAFT throughout, an extremely parameter-efficient fine-tuning (PEFT) method that applies Tucker tensor decompos...

📖 Read original article


567. OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models ​

Author: Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.DB

arXiv:2602.18094v2 Announce Type: replace-cross Abstract: Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID). However, in real-world scenarios...

📖 Read original article


568. Exact Attention Sensitivity and the Geometry of Transformer Stability ​

Author: Seyed Morteza Emadi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation. Our main result is the exact identity $|J_\tau(u)|{\infty\to1}=\theta(p)/\tau$ for the Jacobian $J\tau(u)$ of the tempered softm...

📖 Read original article


569. Reasoning-Based Personalized Generation for Users with Sparse Data ​

Author: Bo Ni, Branislav Kveton, Samyadeep Basu, Subhojyoti Mukherjee, Leyao Wang, Franck Dernoncourt, Sungchul Kim, Seunghyun Yoon, Zichao Wang, Ruiyi Zhang, Puneet Mathur, Jihyung Kil, Jiuxiang Gu, Nedim Lipka, Yu Wang, Ryan A. Rossi, Tyler Derr
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2602.21219v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) personalization holds great promise for tailoring responses by leveraging personal context and history. However, real-world users usually possess sparse interaction histories with limited personal context, such as c...

📖 Read original article


570. SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance ​

Author: Minghan Yang, Lan Yang, Ke Li, Honggang Zhang, Kaiyue Pang, Yizhe Song
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.21819v3 Announce Type: replace-cross Abstract: Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending th...

📖 Read original article


571. Automating the Detection of Requirement Dependencies Using Large Language Models ​

Author: Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel C. Briand, Ramesh S., Arun Adiththan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2602.22456v2 Announce Type: replace-cross Abstract: Requirements are inherently interconnected through various types of dependencies. Identifying these dependencies is essential, as they underpin critical decisions and influence a range of activities throughout software development. However, t...

📖 Read original article


572. Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs ​

Author: Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud, Simon Woodhead
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.02830v2 Announce Type: replace-cross Abstract: Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions. One of the key approaches to do this has been through the use of knowledge tracing (KT) mod...

📖 Read original article


573. Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast ​

Author: Mehmet Yigit Avci (and for the Alzheimer's Disease Neuroimaging Initiative), Akshit Achara (and for the Alzheimer's Disease Neuroimaging Initiative), Andrew King (and for the Alzheimer's Disease Neuroimaging Initiative), Jorge Cardoso (and for the Alzheimer's Disease Neuroimaging Initiative)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.04113v3 Announce Type: replace-cross Abstract: Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisition characteristics have been shown to contribute substantially to this predictability. Whether the sam...

📖 Read original article


574. Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs ​

Author: Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual ...

📖 Read original article


575. Informative Perturbation Selection for Uncertainty-Aware Post-hoc Explanations ​

Author: Sumedha Chugh, Ranjitha Prasad, Nazreen Shah
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2603.14894v3 Announce Type: replace-cross Abstract: Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model explanations. Post-hoc model-agnostic explanation methods addresses this challenge by learning a surroga...

📖 Read original article


576. Data-knowledge dual-driven intelligent framework for full-chain, experiment-efficient synthesis of 2D dendrites ​

Author: Wenqiang Huang, Xuhang Gu, Susu Fang, Shen'ao Xue, Huanhuan Xing, Junjie Jiang, Junying Zhang, Shen Zhou, Zheng Luo, Jin Zhang, Fangping Ouyang, Shanshan Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2603.16959v2 Announce Type: replace-cross Abstract: Exemplified by the chemical vapor deposition growth of two-dimensional dendrites, which has potential applications in catalysis and presents a parameter-intensive, data-scarce and reaction process-complex model problem, we devise a machine in...

📖 Read original article


577. FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion ​

Author: Hugo Caselles-Dupr'e, Mathis Koroglu, Guillaume Jeanneret, Arnaud Dapogny, Matthieu Cord
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.17555v2 Announce Type: replace-cross Abstract: Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's native resolution often loses fine-grained structure, whereas high...

📖 Read original article


578. Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought ​

Author: Zichen Xie, Wenxi Wang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2603.18334v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly assist secure software development, their ability to meet the rigorous demands of Rust program verification remains unclear. Existing evaluations treat Rust verification as a black box, assessing m...

📖 Read original article


579. SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs ​

Author: Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu
Published: 8/18/2026, 4:00:00 AM
Categories: physics.comp-ph, cs.AI, cs.DC, cs.LG

arXiv:2603.20253v4 Announce Type: replace-cross Abstract: Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To ad...

📖 Read original article


580. Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence ​

Author: Peter Fasogbon, Ugurcan Budak, Patrice Rondao Alface, Hamed Rezazadegan Tavakoli
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2603.21933v2 Announce Type: replace-cross Abstract: The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream processing. However, most of the existing pruning strategies depend on camera parameters, rendered images, ...

📖 Read original article


581. VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models ​

Author: Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Zhongzheng Ren, Daniel S Weld, Ranjay Krishna
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.24575v2 Announce Type: replace-cross Abstract: Scalable Vector Graphics (SVG) are essential for technical illustration and digital design, offering resolution independence and semantic editability. In practice, original vector files are frequently lost, leaving only rasterized versions (e...

📖 Read original article


582. I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering ​

Author: Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu, Gillian K. Hadfield
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.03904v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing over expressing uncertainty. We study epistemic abstention for factual questions with verifiable answer...

📖 Read original article


583. Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models ​

Author: Davood Soleymanzadeh, Xiao Liang, Minghui Zheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2604.07084v2 Announce Type: replace-cross Abstract: Open-loop end-to-end neural motion planners have recently been proposed to improve motion planning for robotic manipulators. These methods enable planning directly from sensor observations without relying on a privileged collision checker dur...

📖 Read original article


584. RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies ​

Author: Jenai Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2604.09860v4 Announce Type: replace-cross Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often ...

📖 Read original article


585. Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification ​

Author: Jiho Noh, Mukhesh Raghava Katragadda, Raymond Carl, Soon Lee
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.21137v3 Announce Type: replace-cross Abstract: Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse ...

📖 Read original article


586. Structural Generalization on SLOG without Hand-Written Rules ​

Author: Zichao Wei
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.26157v4 Announce Type: replace-cross Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations. Existing approaches either rely on hand-written algebraic rules (AM-Parser) or fail to generalize structural...

📖 Read original article


587. Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR ​

Author: Kazuki Egashira, Mark Vero, Jasper Dekoninck, Florian E. Dorner, Robin Staab, Martin Vechev
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.02909v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designed for tasks with verifiable ground-truth answers, real-world v...

📖 Read original article


588. Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models ​

Author: Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.08974v2 Announce Type: replace-cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently trac...

📖 Read original article


589. Evolving Ensemble of Agents ​

Author: Zongmin Yu, Liu Yang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG

arXiv:2605.09018v4 Announce Type: replace-cross Abstract: We introduce the Evolving Ensemble of Agents (EvE), a decentralized framework that organizes existing, highly capable coding agents into a live, co-evolving system for algorithmic discovery. Rather than reinventing the wheel within the ``LLMs...

📖 Read original article


590. AgentMV: A State-Guided Multi-Agent Framework for Budget-Aware Music Video Generation ​

Author: Huimin Wang, Chang Xia, Leilei Ouyang, Yongqi Kang, Yu Fu, Yuqi Ouyang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MA

arXiv:2605.10723v2 Announce Type: replace-cross Abstract: Generating a complete music video from a song requires more than synthesizing visually plausible clips for individual lyric prompts. A practical system must maintain long-range visual consistency, coordinate recurring motifs, synchronize edit...

📖 Read original article


591. Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting ​

Author: Amritansh Maurya, Navjot Singh, Mohammed Javed, Omar Moured
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CV, cs.LG

arXiv:2605.20254v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, because Table Question-Answering (TQA) requires precise cell retrieval and multi-step structure...

📖 Read original article


592. SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Activation Sparsity ​

Author: Ting Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.21333v2 Announce Type: replace-cross Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combination that has produced a persistent quality gap relative to dense Transformers. We present SymbolicLigh...

📖 Read original article


593. EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control ​

Author: Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2605.21862v2 Announce Type: replace-cross Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact, occlusion, and object motion, and the geometry that later deci...

📖 Read original article


594. Periodic Topological Deep Learning for Polymer Design and Discovery ​

Author: Yasharth Yadav, Tze Kwang Gerald Er, Atsushi Goto, Kelin Xia
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.26833v2 Announce Type: replace-cross Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine learning approaches represent polymers as molecular graphs of a single repeati...

📖 Read original article


595. The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution ​

Author: Deepak Panigrahy, Aakash Tyagi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR, cs.DC, cs.PF

arXiv:2605.27599v3 Announce Type: replace-cross Abstract: Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-base...

📖 Read original article


596. Annealed Softmax Greedy in Many-Armed Bayesian Bandits ​

Author: William Overman, Mohsen Bayati
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.31034v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward. These updat...

📖 Read original article


597. SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation ​

Author: Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.00380v2 Announce Type: replace-cross Abstract: Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which i...

📖 Read original article


598. Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics ​

Author: Mohammad Amanlou, Yasaman Amou-Jafari, Fereshte Bagheri, Fatemeh Boloukazari, Mehrad Liviyan, Elahe Khodaverdi Nadrabadi, Shahab Sherafat, Behnam Bahrak
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2606.01375v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly entering students' learning practices, but their educational value may depend on whether they are used to support reasoning or to complete tasks without engaging in the underlying reasoning. This ...

📖 Read original article


599. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments ​

Author: Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2606.03770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost efficiency, low latency, and optimal resource u...

📖 Read original article


600. MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models ​

Author: Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari, Gholamreza Haffari, Trang Vu, Lizhen Qu, Dinh Phung
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, eess.AS

arXiv:2606.05177v2 Announce Type: replace-cross Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text. We introduce MCBench, a benchmark with 1196 scenarios spanning four safety categorie...

📖 Read original article


601. The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models ​

Author: Patrick Keough
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2606.05183v2 Announce Type: replace-cross Abstract: Pass/fail safety evaluation reports whether a model refused. It does not report how far a model went to please the user, and we show these are close to different measurements. We audited sycophancy across three Gemini generations, scoring N=8...

📖 Read original article


602. Compositional Boundaries for Density Fusion ​

Author: Ratan Bahadur Thapa, Ali Darijani, J"urgen Beyerer, Steffen Staab
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT, stat.ME

arXiv:2606.05871v2 Announce Type: replace-cross Abstract: Distributed uncertainty-management systems often combine local probabilistic models along aggregation trees chosen by communication, privacy, or scheduling constraints. The final density should depend on the weighted sources, not on the parti...

📖 Read original article


603. LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents ​

Author: Aofan Yu, Chenyu Zhou, Tianyi Xu, Zihan Guo, Rong Shan, Zhihui Fu, Jun Wang, Weiwen Liu, Yong Yu, Weinan Zhang, Jianghao Lin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.06087v2 Announce Type: replace-cross Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and exposes skill content as plaintext. We present LatentSkill, a f...

📖 Read original article


604. Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries ​

Author: Linfeng Cao, Ming Shi, Ness B. Shroff
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.08410v2 Announce Type: replace-cross Abstract: Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and unknown preferences, existing methods infer preferences only ...

📖 Read original article


605. Culturally-Aware AI for Cross-Boundary Community Learning: Undergraduate Innovation at the Intersection of Computation and Design ​

Author: Jiaojiao Zhao, Weisheng Zhang, Jiawen Cai, Haibin Gao, Luyao Zhang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.GR, cs.HC, cs.MM

arXiv:2606.09041v2 Announce Type: replace-cross Abstract: Research on artificial intelligence in education (AIED) is rapidly expanding, yet technical progress often lacks human-centered grounding and adequate attention to cultural context. Community-Based Learning, a pedagogy rooted in social work, ...

📖 Read original article


606. Speculative Rollback Correction for Quality-Diverse Web Agent Imitation ​

Author: Longkun Hao, Hongyu Lin, Hao Li, Zhuowen Liu, Zhichao Yang, Haojie Hao, Dongshuo Huang, Haitao Yang, Hongyu Ge, Ming jie Xie, Yanjun Wu, Zi Hao Yin, Yan Bai, Yihang Lou
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.12485v2 Announce Type: replace-cross Abstract: Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, determining the optimal timing for expert intervention presents a critical challenge in this context. Del...

📖 Read original article


607. SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models ​

Author: Feng Wu, Harsh Deep, Eric Lehman, Sanyam Kapoor, Guoshuai Zhao, Rahul G. Krishnan, Gari Clifford, Li-wei H Lehman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.19888v2 Announce Type: replace-cross Abstract: Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data. While recent self-supervised...

📖 Read original article


608. Empowering Polymeric Materials Discovery by Artificial Intelligence ​

Author: Chenyao Ma, Linda Zhang, Yuheng Chen, Wei Du, Shangwen Fang, Zihao Jiang, Chuanyu Liu, Xinyu Ma, Rui Su, Gang Wang, Muyao Yu, Dong Zhong, Jie Zhu, Weibo Gong, Huan Gu, Limin Li, Chen Shen, Rui Wu, Zhenghao Wu, Kan Xu, Min Zhou, Donglin He, Xiayun Huang, Shan Jiang, Pengfei Ou, Jiayu Peng, Yuwei Zhang, Jie Zhao, Di Zhang, Piao Ma, Zhenghao Li, Hao Li
Published: 8/18/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI

arXiv:2606.20753v2 Announce Type: replace-cross Abstract: Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing. Yet their rational design remains exceptionally challenging because material performance emerges from comple...

📖 Read original article


609. Red-Teaming the Agentic Red-Team ​

Author: Dario Pasquini, Michal Bazyli, Taras Fedynyshyn, Artem Sorokin
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2606.24496v2 Announce Type: replace-cross Abstract: The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has b...

📖 Read original article


610. LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression ​

Author: Zhuowen Liu, Longkun Hao, Shiyu Feng, Xiaowen Chang, Ruiqun Li, Changqun Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.03057v2 Announce Type: replace-cross Abstract: The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. As a hardware-agnostic and highly compatible approach, low-rank compression has been widely adopted to r...

📖 Read original article


611. Statistical Adversaries: Natural Backdoor-like Adversarial Features in Clean Vision Datasets ​

Author: Paul K. Mandal, Pavan Reddy, Tristan Malatynski
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR, cs.LG

arXiv:2607.05516v2 Announce Type: replace-cross Abstract: Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave as backdoor-like triggers without being maliciously inserted. We call ...

📖 Read original article


612. Efficient Safety Alignment of Language Models via Latent Personality Traits ​

Author: Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR

arXiv:2607.07918v2 Announce Type: replace-cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility ...

📖 Read original article


613. LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning ​

Author: Ning Liu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.10139v3 Announce Type: replace-cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward...

📖 Read original article


614. Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants ​

Author: Dipesh Tharu Mahato
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.13039v2 Announce Type: replace-cross Abstract: A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access p...

📖 Read original article


615. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values ​

Author: Jan Betley, Johannes Treutlein, Jan Dubi'nski, Harry Mayne, Karol Ga{\l}\k{a}zka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2607.14345v4 Announce Type: replace-cross Abstract: People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed...

📖 Read original article


616. Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature ​

Author: Maximilian K"ahler, Katja Konermann, Lisa Kluge, Markus Schumacher
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.14882v2 Announce Type: replace-cross Abstract: With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Cl...

📖 Read original article


617. Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft ​

Author: Zeynep Engin, Tim Gordon, Viviana Bastidas, Tom Crick, Jon Crowcroft, Jean-Martin Denis, David J. Hand, Lauren Maffeo, Jakob M"okander, Irene Ng, Anastasija Nikiforova, Giulio Quaggiotto, David Uriel Socol de la Osa, Rhonda Syler, Philip Treleaven, Stefaan Verhulst
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.ET, cs.SI, cs.SY, eess.SY

arXiv:2607.18483v2 Announce Type: replace-cross Abstract: The digital substrate - data, algorithms, infrastructure, platforms, applications - is being governed without adequate conceptual foundations. The ability and legitimacy required to govern this substrate, and to govern with it, are simultaneo...

📖 Read original article


618. SCPP: A Unified Python Library for Soft Clustering ​

Author: Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje, Ali Sajedifar, Manny Chalak, Ava Zerafatangiz, Sadegh Eskandari
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.19620v2 Announce Type: replace-cross Abstract: In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training, prediction, memb...

📖 Read original article


619. G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection ​

Author: Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung, Moongu Jeon
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.19942v3 Announce Type: replace-cross Abstract: This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited view...

📖 Read original article


620. Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies ​

Author: Isak Hwang, Yoon Pyo Lee, Syed Bahauddin Alam
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22067v2 Announce Type: replace-cross Abstract: Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforces. We evaluate an open-weight 31-billion-parameter multimodal model (Gemma 4 31B-IT) on the U.S. Nuclea...

📖 Read original article


621. When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation ​

Author: Renjie Liang, Zijian Xu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22771v2 Announce Type: replace-cross Abstract: Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive ...

📖 Read original article


622. Moral Hazard in Multi-Agent Language Models ​

Author: Dane Malenfant
Published: 8/18/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.23982v5 Announce Type: replace-cross Abstract: Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr"om's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theory-grounded controll...

📖 Read original article


623. EEG Emotion Recognition From AI-Generated Biodigital Architecture Images ​

Author: Hongye Yang, Eva Guttmann-Flury
Published: 8/18/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.HC

arXiv:2607.24808v2 Announce Type: replace-cross Abstract: Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-experiment involving 336 participants identified 60 images, selected from an initial pool of 600, that elici...

📖 Read original article


624. ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection ​

Author: Paul Julius K"uhn, Saptarshi Neil Sinha, Tiago Kleist, Richard Hoffmann, Arjan kuijper, Michael Weinmann
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27065v3 Announce Type: replace-cross Abstract: While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pip...

📖 Read original article


Author: Mohammad Asif, Azizuddin Khan, Mohd Azam, Anurag Rajkumar Bombarde
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2607.28687v2 Announce Type: replace-cross Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically synthesizes recent te...

📖 Read original article


626. WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization ​

Author: Fanzhe Wei (Metask Lab), Li Liu (Metask Lab), Ziyang Wang (Metask Lab), Chenyu Wang (Metask Lab)
Published: 8/18/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.28699v3 Announce Type: replace-cross Abstract: KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now. We give it a provably sound runtime meter -- a "DTrace for KV quantizat...

📖 Read original article


627. UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations ​

Author: Ziyue Kang, Nan Nan, Chenhao Lin, Xiaohong Guan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG

arXiv:2608.00576v2 Announce Type: replace-cross Abstract: High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded representations with fixed tracks or slots. Converting richly orchestrated scores into compact forms is the...

📖 Read original article


628. Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't ​

Author: Ravi Satya Durga Prasad Yenugula
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.02829v3 Announce Type: replace-cross Abstract: Model families are typically trained size by size, each from scratch. Can a pretrained large model instead be converted into a smaller sibling? We characterize the 1.4B->410M conversion in Pythia end to end. Representations align strongly acr...

📖 Read original article


629. The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk ​

Author: Francis Heylighen
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.03361v2 Announce Type: replace-cross Abstract: AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of...

📖 Read original article


630. SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation ​

Author: Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Zhengyu Chen, Yikun Ban, Shuang Qiu, Zhongxiang Dai
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.04419v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal wh...

📖 Read original article


631. Agentic AI: User Empowerment or Foreclosure? ​

Author: David Gamba, Daniel M. Romero, Grant Schoenebeck
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.06510v2 Announce Type: replace-cross Abstract: Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and one that depends on more than the technology. We conduct a c...

📖 Read original article


632. Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations ​

Author: Ji\v{r}'i Klepl, Maty'a\v{s} Brabec, Martin Kruli\v{s}
Published: 8/18/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.08085v2 Announce Type: replace-cross Abstract: Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource usage. These optimizations require a deep understanding of low-level hardware details and the intricacies...

📖 Read original article


633. Population-Scalable Multi-Agent World Modeling ​

Author: Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li, Yimin Sheng, Tianxi Tan, Zhenkai Zhang, Jianyi Zhu, Yong-Lu Li
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08600v2 Announce Type: replace-cross Abstract: World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed ...

📖 Read original article


634. Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories ​

Author: Shuyan Huang, Kai Du, Andrew Lan
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.10319v2 Announce Type: replace-cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As developers collaborate with coding agents over time, their preferences emerge through repeated in...

📖 Read original article


635. Persistent Recursive Worlds Enable Autonomous Software Evolution ​

Author: Beichen Huang, Zhenyu Liang, Bowen Zheng, Ran Cheng
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.MA, cs.NE

arXiv:2608.10450v3 Announce Type: replace-cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, managers or shared context. We introduce EvoX G...

📖 Read original article


636. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? ​

Author: Xiaohongshu Dots Studio, Evolvent AI
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.10875v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks ...

📖 Read original article


637. Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models ​

Author: Yoshihiko Kayama
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, nlin.CG

arXiv:2608.11657v2 Announce Type: replace-cross Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous, closed-loop dynamical system. By establishing a non-linear homeostatic feedba...

📖 Read original article


638. Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework ​

Author: Avinash Agarwal, Vridhi Jain
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.11891v2 Announce Type: replace-cross Abstract: Purpose: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. This paper assesses India's foundation-model ecosystem and examines whether apparent ca...

📖 Read original article


639. Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation ​

Author: Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
Published: 8/18/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.11980v2 Announce Type: replace-cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for ite...

📖 Read original article


640. No One to Blame: A Framework of Constitutive AI Unaccountability ​

Author: Long Hoang Nguyen, Eva Sp"athe, Sebastian Lins, Ali Sunyaev
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.12104v2 Announce Type: replace-cross Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparenc...

📖 Read original article


641. One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL ​

Author: Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.12253v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collaps...

📖 Read original article


642. Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review ​

Author: Joel Abenhaim
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.12440v2 Announce Type: replace-cross Abstract: This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validat...

📖 Read original article


643. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory ​

Author: Le Zhang, Ke Sun
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.HC

arXiv:2608.12627v2 Announce Type: replace-cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for ag...

📖 Read original article


644. Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs ​

Author: Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.13368v2 Announce Type: replace-cross Abstract: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialize...

📖 Read original article


645. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency ​

Author: Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos
Published: 8/18/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.14065v2 Announce Type: replace-cross Abstract: Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on o...

📖 Read original article


646. From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics ​

Author: Meng Li, Chuqi Chen, Zhengqing Gao, Xi Zhou, Xiao Sun, Yang Xiang, Huaxi Huang
Published: 8/18/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GR

arXiv:2608.14120v2 Announce Type: replace-cross Abstract: Lagrangian modeling is vital to fluid dynamics, as it characterizes particle transport and complements the Eulerian representation. However, Lagrangian trajectories are less commonly available than Eulerian fields, while most neural operators...

📖 Read original article


647. Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination ​

Author: Shuo Liang, Yixing Ma, Pengfei Zhou, Zhenglin Wan, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
Published: 8/18/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.14391v2 Announce Type: replace-cross Abstract: Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector a...

📖 Read original article