Skip to content

arXiv cs.AI - 2026-08-25 ​

585 items collected.


1. KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference ​

Author: Srihari Unnikrishnan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2608.21362v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, lim...

📖 Read original article


2. AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance ​

Author: Ali Toygar Abak
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.21363v1 Announce Type: new Abstract: A protocol is presented for recording the governance decisions of automated AI runtimes. When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check...

📖 Read original article


3. Reviewing Model Collapse and Countermeasures ​

Author: Xihao Xie, Beichen Hu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.21366v1 Announce Type: new Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors. The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-genera...

📖 Read original article


4. AI Learning and Conceptual Transfer in the Game of Hidden Rules ​

Author: Christo Mathew, Wentian Wang, Jacob Feldman, Lazaros K. Gallos, Paul B. Kantor, Vladimir Menkov, Hao Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.21372v1 Announce Type: new Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, gen...

📖 Read original article


5. LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform ​

Author: Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21374v1 Announce Type: new Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce ...

📖 Read original article


6. SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG ​

Author: Yong-eun Cho
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21375v1 Announce Type: new Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores. Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity...

📖 Read original article


7. RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students ​

Author: Ria Sidhu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.21379v1 Announce Type: new Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has alrea...

📖 Read original article


8. There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items ​

Author: V. S. Raghu Parupudi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on th...

📖 Read original article


9. Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference ​

Author: Sandeep Bokkasam, Pankaj D
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21393v1 Announce Type: new Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around l...

📖 Read original article


10. Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering ​

Author: Toneema Zubair
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21408v1 Announce Type: new Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-s...

📖 Read original article


11. The Abstention Protocol: RCA for Clos Fabrics ​

Author: Madhava Gaikwad, Deepak Pandey
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.NI

arXiv:2608.21412v1 Announce Type: new Abstract: Root cause analysis (RCA) in large datacenter networks is challenging because telemetry is noisy, partial, and asynchronous. Score-based approaches degrade under these conditions, often yielding unstable or incorrect attributions. We present \textsc{Co...

📖 Read original article


12. Retrieval-grounded robot program generation and simulation-based correction via Model Context Protocol ​

Author: Zhichao Zhou, Siyuan Chen, Omkar Salunkhe, Ebru Turanoglu Bekar, Johan Stahre, Anders Skoogh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.21417v1 Announce Type: new Abstract: Flexible manufacturing requires industrial robots to be reprogrammed rapidly as product variants change. This paper presents a language-model-based workflow that generates, validates, and iteratively corrects ABB RAPID robot programs from natural langu...

📖 Read original article


13. Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenance, Temporal Reasoning, and Decision Traceability ​

Author: Grama Chethan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.21418v1 Announce Type: new Abstract: Manufacturing knowledge graphs that integrate data from heterogeneous industrial systems face a trust deficit: consumers cannot determine whether queried data is valid, whether it was valid when a decision was made, where it originated, or how it was a...

📖 Read original article


Author: David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar, Arnav Podichetty, Rachael Samberg, Ipek Nil Sancak, Yuhan Shao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.CY

arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues for learning about film history and the evolution of narrative techniques. But the creation of stable benchmarks bu...

📖 Read original article


15. Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities ​

Author: Timothy Merritt, Alejandro Jarabo-Pe~nas, Juan Bravo-Arrabal, Maria-Theresa Bahodi, Anders Lyhne Christensen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.HC, cs.RO

arXiv:2608.21444v1 Announce Type: new Abstract: Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of...

📖 Read original article


16. Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review ​

Author: Louis Peter, Nils Gumpfer, Jana Fischer, Christin Seifert, Jennifer Hannig
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21449v1 Announce Type: new Abstract: Time series arise in a wide range of application domains and are analyzed using machine learning in decision-critical settings. Time series classification (TSC) is one of the most widely studied and relevant tasks. In this context, ensuring the transpa...

📖 Read original article


17. Enhanced Artificial Neural Networks Using QHAdamW in Air Quality Forecasting ​

Author: Mary Joy Daniel Vinas
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.LG, cs.NE

arXiv:2608.21463v1 Announce Type: new Abstract: The study employed an Artificial Neural Network in combination with the optimized Adaptive Moment Estimation (Adam) algorithm, currently the only AQI forecasting model available in the Philippines. The modified QHAdamW - Quasi-Hyperbolic Momentum (QHAd...

📖 Read original article


18. Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning ​

Author: Qifan Shi, Zhaolu Kang, Chenghua Zhu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21501v1 Announce Type: new Abstract: Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: evidence about success, a transport operator that converts this evidence into token-level advantages, and an update geometry that turns advan...

📖 Read original article


19. Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation models ​

Author: Zhiyong Zhou, Song Gao, Qianheng Zhang, Feng Zhang, Zhenhong Du
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21567v1 Announce Type: new Abstract: Human mobility serves as an essential proxy for understanding social, economic, and environmental dynamics in urban systems. Geospatial transferability, which measures a model's capability in a new location or unseen region, is a critical dimension for...

📖 Read original article


20. A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification ​

Author: Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, Paulo Henrique Eleuterio Falsetti, Jo~ao Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.21570v1 Announce Type: new Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hold between 1 and 9 billion parameters, are oriented toward the graphics processing unit, and answer in...

📖 Read original article


21. Robust Lightweight Deep Learning Models for Oral Cancer Screening ​

Author: Siddhant Bharadwaj, Aakash Shedsale, Tejashree Subramanya, Mohd. Azfar, Praveen Birur, Debnath Pal, Shankararama Sharma, Anupama Shetty, Rajesh Sundaresan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21583v1 Announce Type: new Abstract: Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care screening via smartphones offers a scalable solution, developing robust AI for resource-constrained set...

📖 Read original article


22. Data-Driven Dynamic Algorithm Dispatch with Large Language Models ​

Author: Rushil Shah, Emmanuel Lujan, Rabab Alomairy, Alan Edelman
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.NA, math.NA

arXiv:2608.21584v1 Announce Type: new Abstract: We introduce a large language model (LLM)-driven approach for generating dynamic algorithmic dispatch heuristics in high-performance linear algebra. By combining prompt engineering with LLaMA 3 and a curated performance database, the model learns to sy...

📖 Read original article


23. K-Bench: measuring model performance on real scientific agent requests ​

Author: Aubrey Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.21601v1 Announce Type: new Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. Th...

📖 Read original article


24. Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANs ​

Author: Nijesh Upreti, Vaishak Belle
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21605v1 Announce Type: new Abstract: Logic Tensor Network-Enhanced Generative Adversarial Networks (LTN-GANs) inject background knowledge by grounding each logical axiom as a predicate and training the generator to raise its satisfaction, a fuzzy truth value in $[0,1]$. Previous LTN-GAN w...

📖 Read original article


25. Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals ​

Author: Junaid Farooq
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21610v1 Announce Type: new Abstract: Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by similarity, discarding the hierarchical structure of the source. We introduce Semantic Compression Trees (SCT), a hi...

📖 Read original article


26. SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning ​

Author: Yujie Zhang, Bin Gao, Tulika Mitra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2608.21614v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity through sparse expert...

📖 Read original article


27. Measuring Activation Control in Large Language Models ​

Author: Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.21664v1 Announce Type: new Abstract: Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control the...

📖 Read original article


28. From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation ​

Author: Yuan An, Emily Wang, Benjamin Wang, Ruhma Hashmi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.21668v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decis...

📖 Read original article


29. Context as an Environment: Programmatic Context Management for Long-Horizon Agents ​

Author: Yin Lin, Elaine Ang, Erkang Zhu, Bolin Ding, Jingren Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21690v1 Announce Type: new Abstract: LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier interactions or extract selected information into fixed memory representations, committing to what to ...

📖 Read original article


30. From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism ​

Author: Jing Liu, Yongxing Qi, Muchen Jiang, Chengnan Hu, Qingqing Peng, Haoming Wang, Yuqing Wang, Yang Yu, Xu Zhang, Ting Wu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.21702v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity, optionally followed by reranking--often returns documents that share keywords with the query without...

📖 Read original article


31. ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling ​

Author: Deyi Li, Qi Xu, Lingyao Li, Tiansheng Wang, Muxuan Liang, Mei Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.21712v1 Announce Type: new Abstract: Transformer-based models are widely used for clinical prediction from electronic health records (EHRs), yet their architectures still require substantial manual tuning, and the optimal configuration may vary across tasks and hospitals. Neural architect...

📖 Read original article


32. Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention ​

Author: Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21721v1 Announce Type: new Abstract: Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying...

📖 Read original article


33. ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents ​

Author: Yu Qian, Hong Miao, Boyang Guo, Tingyi Jiang, Shan Zhao, Tianxing Le, Lintian Li, Meng Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21755v1 Announce Type: new Abstract: Long-horizon agents need memory that identifies relevant experience, resolves revisions, and exposes checkable provenance. We present ECHO (Embodied Context and History Orchestration), an auditable memory architecture and service prototype inspired by ...

📖 Read original article


34. What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation ​

Author: Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh, Seon Ho Kim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.CV

arXiv:2608.21761v1 Announce Type: new Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challengin...

📖 Read original article


35. Physics-Knowledge-Guided Hybrid Neural Learning for Arctic Sea Ice Concentration Evolution and Short-Range Prediction ​

Author: Maqun Zhang, Feng Gao, Wankun Chen, Hui Yu, Yanhai Gan, Junyu Dong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21767v1 Announce Type: new Abstract: Accurate modeling of sea ice concentration (SIC) evolution is essential for polar climate assessment and short?range sea ice prediction. Numerical and data-driven approaches constitute major foundations for SIC modeling, but the former often require co...

📖 Read original article


36. HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries ​

Author: Shangxuan Tian, Yanhui Chen, Carlos Queiroz
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.IR, cs.LG

arXiv:2608.21792v1 Announce Type: new Abstract: Document classification in regulated industries is constrained by data residency, limited cold-start labels, scarce review capacity, and costly model-governance procedures. We present HIRA, a training-free, on-premises retrieval-augmented cascade for d...

📖 Read original article


37. Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning ​

Author: Qiqian Fu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.21811v1 Announce Type: new Abstract: Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2-VL-2B answers 3.6% of rollouts correctly, 85-97% of GRPO rollout groups are entirely wrong and contribute zero g...

📖 Read original article


Author: Jiahao Xie, Guangmo Tong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.21825v1 Announce Type: new Abstract: Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph information from visual observations. Existing methods typically rely on fixed KNN-based heuristics for candidate edge selection and fail to capt...

📖 Read original article


39. Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents ​

Author: Chengyang Gu, Le Zhang, Jingbo Zhou, Yize Chen, Yu Shi, Siqi Bao, Zheng-Fan Wu, Hua Wu, Hui Xiong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21830v1 Announce Type: new Abstract: Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. H...

📖 Read original article


40. GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? ​

Author: Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.21833v1 Announce Type: new Abstract: Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability...

📖 Read original article


41. LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization ​

Author: Hui Zeng, Pengfei Yang, Yanxin Chen, Fusong Ju, Xinran Wei
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21836v1 Announce Type: new Abstract: Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We identify a benchmark-...

📖 Read original article


42. AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations ​

Author: Rachel Poonsiriwong (Pub), Chayapatr (Pub), Archiwaranguprok, Constanze Albrecht, Monchai Lertsutthiwong, Pattie Maes, Pat Pataranutaporn
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.21841v1 Announce Type: new Abstract: Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-patter...

📖 Read original article


43. MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance ​

Author: Haoyu Wang, Guangyuan Dong, He Liang, Zijing Zhang, Jiachen Luo, Chuang Liu, Chao Xue, Hao Tang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.21867v1 Announce Type: new Abstract: LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds o...

📖 Read original article


44. HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews ​

Author: Ao Chen, Xiaojiang Peng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21868v1 Announce Type: new Abstract: Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PHQ-8 profile. This process is hierarchical: relevant evidence is often sparse and context-dependent within local q...

📖 Read original article


45. From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning ​

Author: Chenghao Zhang, Yikai Mao, Shanqi Liu, Haoyu Gao, SaiSai Hu, Dan Roth
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21897v1 Announce Type: new Abstract: Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. We stu...

📖 Read original article


46. Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning ​

Author: Chenghao Zhang, Canran Xiao, SaiSai Hu, Dan Roth
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21898v1 Announce Type: new Abstract: Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent states, or infeasible tasks. We address the gap between scalable environm...

📖 Read original article


47. Consistency Is Not Coherence: Orientation Search for Certified Alignments Between 4D Defence Upper Ontologies ​

Author: Fabio Rovai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.LO

arXiv:2608.21914v1 Announce Type: new Abstract: We align three upper ontologies that sit under UK and NATO defence data infrastructure: the Information Exchange Standard (IES), the Higher Quality Data Model (HQDM) that underpins the National Digital Twin, and Basic Formal Ontology (BFO). No public a...

📖 Read original article


48. ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation ​

Author: Weichu Liu, Yuxuan Hu, Yirong Sun, Ningning Mao, Ziyun Zhang, Jian Chen, Mingyang Xu, Qishan Zhong, Chengming Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21925v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage-aware reasoning and seaml...

📖 Read original article


49. GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI ​

Author: Zhesheng Zhang, Jiahao Lu, Wei Liu, Cong Pan, Jianhua Yang, Yixiang Chen, Hongyuan Yu, Mengqi Zhang, Kailin Lyu, Zhumin Chen, Keji He
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.RO

arXiv:2608.21928v1 Announce Type: new Abstract: In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis o...

📖 Read original article


50. Multimodal Prompt Learning with Irregular EHRs for Robust Monitoring of Critical Care Patients ​

Author: Yixin Yang, Yueyang Sun, Weichen Liu, Xianbing Zhao, Sicen Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21941v1 Announce Type: new Abstract: Accurate assessment of patients in intensive care units (ICUs) is essential for timely clinical intervention and improved patient outcomes. Multimodal electronic health records (EHRs), including structured physiological time series and longitudinal cli...

📖 Read original article


51. TessIndex: Capability Verified Identity System for the Agent Economy ​

Author: Mehul Goenka, Tejas Pathak, Siddharth Asthana
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.MA, cs.NI

arXiv:2608.21942v1 Announce Type: new Abstract: Software systems have traditionally been organized around applications where human users act as principal decision-makers. Recent developments in agentic capabilities alter this paradigm: software agents now autonomously translate high-level goals into...

📖 Read original article


52. SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality ​

Author: Hyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim, Dongjin Lee, Jaehoon Lee, Sungroh Yoon, Dahuin Jung
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21952v1 Announce Type: new Abstract: Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling and providing a viable alternative to Transformers. Building upon this, Mamba-2 introduces the Structured State ...

📖 Read original article


53. Repo2Skill-Evo: Repository Skills Go Stale in Silence ​

Author: Chenyuan Duan, Ge Shi, Zineng Mao, Ge Zhang, Hao Liang, Yinzhu Piao, Yuchen Wu, Zhixin Yao, Kaiyu Huang, Wenhao Huang, Linzhuang Sun, Shen Yan, Wentao Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.21964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific procedural knowledge: which APIs to call, which scripts to run, and which conventions the current release expects. ...

📖 Read original article


54. Closed-loop AI achieves certifiable engineering design ​

Author: Tianyi Yu, Chengxing Tao, Haoxuan Shen, Huiyang Li, Rugang Chen, Long Teng, Lilin Wang, Yan Li, Qingbin Chen, Chaogang Xu, Lizhong Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21976v1 Announce Type: new Abstract: Agentic AI has automated parts of scientific discovery, including paper generation, expert-level coding, therapeutic proposal, and autonomous experimentation. Complex physical engineering design remains a gap, because candidates must satisfy simultaneo...

📖 Read original article


55. Beyond Similarity: Heterogeneous Graph Learning for Multi-Objective Food Substitution in Charitable Food Agencies ​

Author: Naimur Rahman Chowdhury, Limon Bin Hossain
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21979v1 Announce Type: new Abstract: Charitable food agencies play an important role in alleviating food insecurity by distributing donated food to people in need. However, they rely on ad hoc in-kind donations and often face shortages of specific foods, so they offer substitutes. A good ...

📖 Read original article


56. Redteaming Leading Arabic LLMs with ASAS ​

Author: Fidaa Abed, Haidar Khan, M Saiful Bari, Babar Khan, Abdalghani Abujabal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.21985v1 Announce Type: new Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We...

📖 Read original article


57. DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction ​

Author: Joe Yu, Shibin Thomas Stanley Paul, Sven Mayer
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22014v1 Announce Type: new Abstract: Automated prompt and skill optimization typically produces a single static instruction that is reused across inference instances until the next optimization cycle. However, this approach cannot adapt when the required context, constraints, and evidence...

📖 Read original article


58. SPAR-Hate: An Auditor-Guided Multi-Agent Framework for Bilingual Hate Speech Parsing ​

Author: Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22018v1 Announce Type: new Abstract: Hate speech detection has recently shifted from coarse-grained classification to structured parsing, where systems must jointly identify hateful targets, arguments, and target-level labels. However, existing studies primarily emphasize benchmark evalua...

📖 Read original article


59. One-Step Evolution for Long-Time Extrapolation: An Error-Bound-Informed and Prior-Guided Neural Residual Framework for Autonomous PDEs ​

Author: Maqun Zhang, Feng Gao, Wankun Chen, Hui Yu, Yanhai Gan, Junyu Dong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22026v1 Announce Type: new Abstract: Accurate simulation of the long-time evolution of systems governed by partial differential equations (PDEs) is central to scientific computing. Among existing deep learning?based approaches for solving PDEs, neural operators typically rely on extensive...

📖 Read original article


60. More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning ​

Author: Orion Powers, Daniella Seum, Khaled Slhoub
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22048v1 Announce Type: new Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphasize accuracy while fewer quantify local runtime and energy, characterize failure modes, or apply paired statistic...

📖 Read original article


61. GenCoord: Skill-Path Commitments under Private Information ​

Author: Peng He, Junning Zhu, Haohan Yuan, Jianpeng Liang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22055v1 Announce Type: new Abstract: Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither local view determines who should act, what should be handed off, or how the joint task should continue. We introd...

📖 Read original article


62. MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds ​

Author: Minjae Seo, Wonwoo Choi, Geonwoo Han, Taekyoung Kwon, Yongsu Kim, Sang Seo, Jaewon Noh, Hankyul Baek, Seongyun Seo, Myoungsung You
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.22061v1 Announce Type: new Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that ...

📖 Read original article


63. Search Broadly, Seek Evidence on Both Sides, Decide Narrowly: Evidence-Admissible GraphRAG for Longitudinal Clinical Event Verification ​

Author: Xingtao Lin, Yubo Feng, Weixin Liu, Hangqi Ren, Junchao Zhou, Caiwan Sun, You Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22062v1 Announce Type: new Abstract: Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clinical events. This task is challenging because evidence is distributed across structured records, notes, laboratory...

📖 Read original article


64. From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers ​

Author: Bartolomeo Bogliolo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2608.22063v1 Announce Type: new Abstract: Agents built on Large Language Models (LLMs) increasingly reach enterprise data through the Model Context Protocol (MCP), and many MCP database servers maximize flexibility by exposing a single generic SQL execution tool. This paper proposes the Domain...

📖 Read original article


65. Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays ​

Author: Edwin Ouko, Emmanuel Lujan, Alan Edelman, Robert Metcalfe
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CE

arXiv:2608.22068v1 Announce Type: new Abstract: Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations, provide opportunities to enhance energy production capacity and increase fault tolerance. The development and adoption of these emerging g...

📖 Read original article


66. Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation ​

Author: Laxmigayathri Challa, Yuhan Zhou, Ana Cleveland, Haihua Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22085v1 Announce Type: new Abstract: Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream ...

📖 Read original article


67. Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks ​

Author: Amit Roth, Ivan Bercovich, Yonathan Efroni
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22103v1 Announce Type: new Abstract: As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically rel...

📖 Read original article


68. Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings ​

Author: Aarzoo Dhiman, Farzana Haque, Kartikae Grover, Lydia Brian Smith, William Stephen Jones
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22108v1 Announce Type: new Abstract: Breast Cancer Multidisciplinary Team (MDT) meetings manage increasingly complex cases under considerable time pressure, and documentation requirements can reduce clinical efficiency and decision quality. Existing AI based MDT workflows rely on cloud-ba...

📖 Read original article


69. Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning ​

Author: Zhaoda Du, Qiaojie Zheng, Xiaoli Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22128v1 Announce Type: new Abstract: Printability assessment in additive manufacturing is typically conducted at the geometry level before printing to determine whether a computer-aided design (CAD) model or stereolithography (STL) file can be successfully fabricated. Task suitability, in...

📖 Read original article


70. MegaMem: A Retrieval Solution for Ultra-Large Context Windows ​

Author: Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22137v1 Announce Type: new Abstract: Modern language models and agents increasingly require persistent memory for complete codebases, long interaction histories, and heterogeneous enterprise records. The key challenge is to keep hundreds of millions of tokens searchable while passing only...

📖 Read original article


71. Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations ​

Author: Samira Golsefid
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.22138v1 Announce Type: new Abstract: Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed. We present a graded, multi-family, failure-aware framework for stress-testing reasoning models. It perturbs eac...

📖 Read original article


72. MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data ​

Author: Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22141v1 Announce Type: new Abstract: Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- faithful evidence, and cross-query adaptation must be supported together. Enterprise mem- ory extends retrieval be...

📖 Read original article


73. Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems ​

Author: Henry Fordjour Ansah (Louisiana State University of New Orleans), Shreya Banerjee (Louisiana State University of New Orleans), Pranish Ghimire (Louisiana State University of New Orleans)
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22143v1 Announce Type: new Abstract: Qualitative mechanical problem-solving (QMPS) refers to solving qualitative problems from the mechanical domain. Qualitative problems can be solved with minimal discipline-specific information, without any robust quantitative calculation, generally by ...

📖 Read original article


74. AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems ​

Author: Zhixu Du, Yiran Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22160v1 Announce Type: new Abstract: Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at production rates beyond any human line, and their adoption is accelerating. But when their joint decis...

📖 Read original article


75. Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification ​

Author: Qian Ma, Anna Squicciarini, Sarah Rajtmajer
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.22161v1 Announce Type: new Abstract: Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blin...

📖 Read original article


76. MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning ​

Author: Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew, Silvio Savarese, Junnan Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22167v1 Announce Type: new Abstract: Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL frameworks stop at the policy update. For every new domain, the user is left with two hard systems problems: s...

📖 Read original article


77. Role-Specialized Mixture-of-Agents with Open-Weight LLMs for Clinical Prediction ​

Author: Jun Hou, Yi Fang, Xuan Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22176v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly applied to clinical prediction tasks such as in-hospital mortality and readmission from electronic health records (EHRs). Privacy and compliance constraints motivate systems that can be deployed locally, wh...

📖 Read original article


78. Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents ​

Author: Kang Chen, Junjie Nian, Yixin Cao, Yugang Jiang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.22191v1 Announce Type: new Abstract: Software-engineering agents solve repository-level tasks through long, stochastic tool-use trajectories, and repeated attempts often find fixes missed by one run. Test-time scaling is difficult because patches lack canonical answer forms, while sibling...

📖 Read original article


79. Query-Driven Multimodal Information Extraction from Long Documents ​

Author: Yikai Gao, Ding Xia, Xi Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.MM

arXiv:2608.22214v1 Announce Type: new Abstract: In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing eviden...

📖 Read original article


80. Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models ​

Author: Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.MM

arXiv:2608.22232v1 Announce Type: new Abstract: Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and inve...

📖 Read original article


81. Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents ​

Author: Zedong Liu, Jiaan Wu, Xinyang Ma, Le Xu, Kai Wang, Yuanchao Hu, Dingwen Tao, Guangming Tan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22237v1 Announce Type: new Abstract: Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading interfaces often expose entire objects even when only sparse evidence is needed. This over-reading increases token and latency costs and can dilute task...

📖 Read original article


82. Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency ​

Author: Zhihong Cao, Chen Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.22266v1 Announce Type: new Abstract: In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized assistants. A critical aspect of this evolution is the ability to tailor strategic interactions to a user's unique n...

📖 Read original article


83. HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory ​

Author: Yuanhua Lin, Yile Li, Zhiyuan Zhao, Jing Shang, Jian Sun
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22310v1 Announce Type: new Abstract: Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing methods often rely on LLMs to compress or rewrite dialogue histories and use the transformed memories as retrieval evidence. Despite the progress in or...

📖 Read original article


84. Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture ​

Author: Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fern'andez
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.22347v1 Announce Type: new Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal but complete system - a recurrent reasoner with adaptive halting, a homeostatic control field, and a v...

📖 Read original article


85. Addressing the Selection Problem in Explainable AI ​

Author: Claire Vlases, Katelyn Morrison
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.22356v1 Announce Type: new Abstract: Explainable AI (XAI) research has produced a plethora of explanation techniques, yet user studies repeatedly show that available explanations are not effective in practice. We argue that, given the siloed nature of conventional XAI, users are strugglin...

📖 Read original article


86. Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA ​

Author: Jingbo Wang, Sendong Zhao, Haochun Wang, Bing Qin, Ting Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22363v1 Announce Type: new Abstract: Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on English, limiting its relevance to linguistically diverse patients and clinicians. Recent multilingual medical VQA be...

📖 Read original article


87. WAM-OPD: On-Policy Distillation for World Action Models ​

Author: Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.22364v1 Announce Type: new Abstract: World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilities during distillation and later encounter states that are poorly represented by offline data. We study whether on...

📖 Read original article


88. LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis ​

Author: Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckstr"om
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC

arXiv:2608.22417v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended...

📖 Read original article


89. Where World Models Break: Natural-Input Failure Discovery ​

Author: Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang, Xuyang Chen, Hongzong Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22421v1 Announce Type: new Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic prediction failures of world models could dangerously propagate through the control pipeline, as subseq...

📖 Read original article


90. Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding ​

Author: Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22429v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inference latency and fails to effectively resolve the spatial-structural g...

📖 Read original article


91. When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing ​

Author: Taehyeon An, Jaehyeong Park, Donghyuk Shin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.22438v1 Announce Type: new Abstract: Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic perso...

📖 Read original article


92. Small Reasoning Models are Instruction Followers in Function Calling ​

Author: Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.22472v1 Announce Type: new Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calling accuracy through fine-tuning, reinforcement learning (RL), and multi-agent frameworks, particularl...

📖 Read original article


93. When Does AI for PDEs Yield Scientific Evidence? ​

Author: Wenshuo Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22504v1 Announce Type: new Abstract: Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measure...

📖 Read original article


94. ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts ​

Author: YuanHang Xiao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22510v1 Announce Type: new Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acq...

📖 Read original article


95. HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems ​

Author: Christos Sardianos, Iliana Pla, Vasilis Efthymiou, Iraklis Varlamis, Thomas Lagkas, Panagiotis Sarigiannidis, Georgios Th. Papadopoulos
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22512v1 Announce Type: new Abstract: Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orchestrated intrusion campaigns have been reported. Yet, when such a system causes harm, no method can robustly esta...

📖 Read original article


96. CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories ​

Author: Zheyuan Deng, Binghang Lu, Hanqi Feng, Shirley Huang, Dianzhuo Wang, Yuanda Xu, Zhiwei Zhang, Yige Sun, Changhong Mou, Runyu Zhang, Yuexing Hao, Barnabas Poczos, Xiaomin Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22533v1 Announce Type: new Abstract: Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone to procedural failures: misreading application state,...

📖 Read original article


97. STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control ​

Author: Mengxi Luo, Changjia Chen, An Cao, Zirong Huang, Wanyi Dai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22538v1 Announce Type: new Abstract: Policy-governed agents must interpret case evidence while following an authorized procedure. We present \textsc{Stage}, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministic ...

📖 Read original article


98. Scaling Curriculum Learning For Autonomous Driving ​

Author: Cevahir Koprulu, David Paz, Feng Tao, Yuliang Guo, Xinyu Huang, Ufuk Topcu, Liu Ren
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22549v1 Announce Type: new Abstract: Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousands of traffic scenarios and billions of interactions within a matter of days. Although such high-throughput feeds ...

📖 Read original article


99. ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation ​

Author: Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR

arXiv:2608.22559v1 Announce Type: new Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require black-box LLM judges, and typically assume criteria aggregate independ...

📖 Read original article


100. CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents ​

Author: Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang, Haojian Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22577v1 Announce Type: new Abstract: Long-horizon GUI agents can retain a complete interaction trace cheaply as textual action records, but expose only a few past events to the policy in high-fidelity pixels. We formulate this as conditional fidelity restoration: each event persists in su...

📖 Read original article


101. Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning ​

Author: Sparsh Tiwari, Gesina Schwalbe, Bettina Finzel
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22584v1 Announce Type: new Abstract: Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefined symbols from possibly later defined relational reasoning thereon. However, anchoring high-level p...

📖 Read original article


102. Coalition-Aware Skill Reliability for Self-Evolving Agents ​

Author: Qiyan Zhao, Xiaofeng Zhang, Bo Liu, Minda Chen, Wei Xiong, Jingyang Chen, Guanting Ye, Wenhao Yu, Xiaosong Yuan, Shijie Han, Da-Han Wang, Jianmin Ji, Fei Huang, Xu-Yao Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22610v1 Announce Type: new Abstract: Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central mechanism for enabling large language model (LLM)-based self-evolving agents to learn from past experience. Yet ex...

📖 Read original article


103. DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue ​

Author: Qi Zhang, Heajun An, Prakriti Dumaru, Sang Won Lee, Lifu Huang, Pamela J. Wisniewski, Jin-Hee Cho
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22615v1 Announce Type: new Abstract: Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidanc...

📖 Read original article


104. CAI-DLLM: Convergence Aware Inference for Diffusion Language Models ​

Author: Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22646v1 Announce Type: new Abstract: Diffusion language models can generate many tokens in parallel, but they still require repeated denoising steps during inference. This makes generation costly, especially when the model continues to recompute tokens that are already stable. To address ...

📖 Read original article


105. A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems ​

Author: Xiaoyu Zhang, Qiuye Sun, Jiachen Xu, Zhongming Yao, Yushuai Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22672v1 Announce Type: new Abstract: Energy system operation contains a loop of work that automation has never taken over: posing the optimization problem the current cycle should solve, disposing of infeasibility, sequencing a solution into interlocked switching orders, assembling eviden...

📖 Read original article


106. Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses ​

Author: Jiachen Xu, Torben Bach Pedersen, Zhongming Yao, Xiaoyu Zhang, Yushuai Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22676v1 Announce Type: new Abstract: Robustness to a bad tool return means answering it in the way that return calls for, which depends on how the tool went wrong. A tool that has failed and a tool that returns a well-formed falsehood are different problems with different remedies. We ask...

📖 Read original article


107. Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf ​

Author: Davood Wadi, Yu Ma
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC

arXiv:2608.22697v1 Announce Type: new Abstract: Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results...

📖 Read original article


108. CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery ​

Author: Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong, Haochen Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22708v1 Announce Type: new Abstract: Tool use in LLM systems faces a structural trade-off. Progressive disclosure keeps the prompt small by showing only the tools relevant to the current task, while prompt caching rewards a request prefix that stays fixed across calls; every change to the...

📖 Read original article


109. SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale ​

Author: Jiaqi Liu, Maolin Ran, Xiaoyang Lu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22725v1 Announce Type: new Abstract: Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent frameworks generate each shot in isolation, so context ...

📖 Read original article


110. LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans ​

Author: Parisa Ghanad Torshizi, Stacy Marsella
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO

arXiv:2608.22731v1 Announce Type: new Abstract: Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the cont...

📖 Read original article


111. The Compaction Cliff in Long-Running AI Agent Memory ​

Author: Saber Zerhoudi, Jelena Mitrovic, Michael Granitzer
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.22752v1 Announce Type: new Abstract: A safety rule and an episodic log compete for the same tokens in an AI agent's context. When the budget overflows, both are summarized at the same rate; only the rule needs exact wording to remain enforceable. On 20 production agent configurations, Cla...

📖 Read original article


112. Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models ​

Author: Chenhui Liu, Jianpeng Zhou, Jiahai Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22762v1 Announce Type: new Abstract: Knowledge graph question answering (KGQA) is a key task for evaluating KG-augmented Large Language Models (LLMs), and complex KGQA that requires multi-hop reasoning is especially challenging. Solving a complex query involves two coupled phases: candida...

📖 Read original article


113. The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory ​

Author: Qi Feng, Chris Ding, Jicong Fan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22767v1 Announce Type: new Abstract: Long-term language-model agents accumulate memories across interactions, but their retrievers typically do not accumulate retrieval experience. Semantic retrieval is efficient, but embedding similarity does not always reflect whether a memory contains ...

📖 Read original article


114. TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts ​

Author: Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, Mingchen Wan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22788v1 Announce Type: new Abstract: Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for ...

📖 Read original article


115. Performance of a domain-specific large language model in answering patient questions in psychiatry ​

Author: Alexander J. Hish, Arjun Nagendran, Scott N. Compton
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22797v1 Announce Type: new Abstract: Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We develope...

📖 Read original article


116. Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL ​

Author: Kate Gwimm, Carson Eisenach
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22830v1 Announce Type: new Abstract: Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands of tables, and no model can ingest a full catalog at once. We argue that the most effective place to intervene i...

📖 Read original article


117. Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku ​

Author: Xiansheng Luo, Chaowei Zhang, Zewei Zhang, Yi Zhu, Jipeng Qiang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22832v1 Announce Type: new Abstract: The social interactions among crowds via \textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both viewpoint conflicts and consensus, providing fine-grained discriminative social signals that can benefit fake news de...

📖 Read original article


118. FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks ​

Author: Hang Wang, Jin Zhang, Guoliang Xu, Pengyue Lu, Yao Li, Zijiao Zhang, Tianyu Huang, Weiqi Xiong, Yulong Wang, Chuqiao Lu, Wenkang Huang, Kai Yang, Yadong Li, Hui Li, Xingzhong Xu, Xiao Xu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22842v1 Announce Type: new Abstract: Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial documents, with FinixDoc-VL, a 4B-...

📖 Read original article


119. GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis ​

Author: Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao, Wei Liu, Jian Luan, Chenliang Li, Lixin Zou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22847v1 Announce Type: new Abstract: Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific env...

📖 Read original article


120. Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron ​

Author: Sahong Park, Suhwan Park, Hoyoung Lee, Gakyung Kwon, Wonbin Ahn, Jaewon Choi, Alejandro Lopez-Lira, Yoon Kim, Chanyeol Choi, Hyeongwoo Kong, Yongjae Lee
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, q-fin.GN

arXiv:2608.22852v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific investment preferences. We study whether a model's overall investment stance can be calibrated to a spec...

📖 Read original article


121. Proxy reliance in large language model decisions is uncalibrated to predictive evidence ​

Author: Zengqing Wu, Chuan Xiao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY

arXiv:2608.22887v1 Announce Type: new Abstract: Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated...

📖 Read original article


122. CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents ​

Author: Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22899v1 Announce Type: new Abstract: Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. H...

📖 Read original article


123. Beyond Observed Auxiliary Relations: Environment-Conditioned Modeling for Multi-Behavior Recommendation ​

Author: Seunghan Lee, Hyunsik Yoo, Jian Kang, Susik Yoon, SeongKu Kang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.22920v1 Announce Type: new Abstract: Multi-behavior recommendation (MBR) leverages auxiliary behavioral signals, such as clicks and add-to-cart, to enhance target behavior prediction like purchases. While recent graph neural network-based approaches have achieved strong performance by sys...

📖 Read original article


124. Concepts for Securing Agentic AI Coding and the Terok Environment ​

Author: Ji\v{r}'i Vysko\v{c}il, Franz P"oschel, Andreas Kn"upfer
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.22930v1 Announce Type: new Abstract: Agentic AI is a fascinating new tool for software development. It is a huge step forward compared to "conventional" AI assisted coding, which in turn was a considerable breakthrough earlier. AI support through LLMs is a young and very fast-moving field...

📖 Read original article


125. What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels ​

Author: Jiawei He, Mengyu Shi, Jie jia, Xikai Yang, Dong Sun
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22960v1 Announce Type: new Abstract: Coding agents are increasingly evaluated not only by whether they solve a task, but also by how they execute it. However, existing process-level evaluations often treat action prediction, task uncertainty, and step attribution as if they were the same ...

📖 Read original article


126. Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents ​

Author: Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.22963v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated text. Over long trajectories, this text can dominate...

📖 Read original article


127. ParallelWorld: Test-Time Scaling for Embodied Reasoning ​

Author: Min Chen, Shengjun Zhang, Yuxin Li, Zhang Zhang, Xin Fei, Chong Xia, Yueqi Duan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.22971v1 Announce Type: new Abstract: Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning, and interaction within physical environments. Recent studies have shifted the paradigm of embodied reasoning fr...

📖 Read original article


128. Toward Effective and Reliable LLM Agents via Dynamic Ontology ​

Author: Xiaohui Zhang, Zequn Sun, Chengyuan Yang, Yuanning Cui, Lingbing Guo, Wei Hu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22974v1 Announce Type: new Abstract: Large language model (LLM) agents rely heavily on knowledge encoded in model parameters or presented as unstructured context. In domain-specific tasks, this leaves important semantic connections implicit. This often results in incomplete evidence use a...

📖 Read original article


129. Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B ​

Author: Defu Lin, Wenhui Chen, Ziyao Lin, Jianlin Chen, Peiji Long, Chi Man Vong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22975v1 Announce Type: new Abstract: Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constraint through four resource walls: a perceptual Shannon wall for bounded state, a horizon wall for query-independent...

📖 Read original article


130. SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems ​

Author: Xiang Wang, Shigang Quan, Tingzhen Chang, Kang Yang, Sitong Chen, Yabo Fan, Xingxing Wang, Zhaodian He
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22979v1 Announce Type: new Abstract: Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quantization is compact but introduces boundary distortion, whereas dense soft quantization couples representation qu...

📖 Read original article


131. PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts ​

Author: Weiwei Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE

arXiv:2608.23001v1 Announce Type: new Abstract: Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing unrelated metrics and citations to change even when the resulting PDF still builds. PatchWrite instead constrains how candidate edits become committed ...

📖 Read original article


132. PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies ​

Author: Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23028v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks ...

📖 Read original article


133. Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction ​

Author: Peter Kuhn, Chris Pang, Sonakshi Chauhan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23030v1 Announce Type: new Abstract: We study how an AI system can identify and model other agents in its environment from observation alone, which is a capability necessary for cooperative behaviour in the real world. This problem is less constrained than inverse reinforcement learning a...

📖 Read original article


134. MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks ​

Author: Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao, Long Chen, Weigao Sun, Feida Zhu, Yiran Zhong, Steven Hoi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23035v1 Announce Type: new Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot...

📖 Read original article


135. AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces ​

Author: Sungho Park, Wonjoong Kim, Rongyuan Tan, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, cs.SE

arXiv:2608.23041v1 Announce Type: new Abstract: LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual a...

📖 Read original article


136. From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation ​

Author: Xiangxin Zhang, Zhanwei Zhang, Zhihang Fu, Binbin Lin, Wenxiao Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23045v1 Announce Type: new Abstract: Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judging t...

📖 Read original article


137. LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications ​

Author: Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a...

📖 Read original article


138. Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies ​

Author: Xiaotong Tan, Chunli Qiu, Xin Liu, Qing Huang, Guangli Zhou, Bo Gao, Xiaoyan Song, Shuyan Wang, Xiuqin Wang, Wufeng Xue, Ruobing Huang, Dong Ni, Guowei Tao, Jun Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23061v1 Announce Type: new Abstract: Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliability, hallucination, and limited interpretability. We compared the performance of LLMs and reasoning strategies for ...

📖 Read original article


139. From Generation to Simulation: How Far Are World Models from Being True Simulators? ​

Author: Tong Wang, Huan Deng, Mucheng Yang, Yang He, Xiaohui Kuang, Gang Zhao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.23070v1 Announce Type: new Abstract: With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the...

📖 Read original article


140. AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models ​

Author: Saurav Singla, Aarav Singla, Advik Gupta, Parnika Gupta
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23078v1 Announce Type: new Abstract: Large language models increasingly operate over large collections of tools, functions, APIs, and specialized agents. As the candidate action space grows, a function-calling model must process more schemas, consume more prompt tokens, and distinguish am...

📖 Read original article


141. POOL: Propagated Uncertainty Over Lookalikes ​

Author: Rounak Sharma, Ananya B. Sai, Soumyabrata Pal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23086v1 Announce Type: new Abstract: Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to prioritize human review, route uncertain cases to stronger models, or choose abstention thresholds on development...

📖 Read original article


142. Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement ​

Author: Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li, Xin Cong, Yuzhuo Bai, Kangyang Luo, Maosong Sun
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23098v1 Announce Type: new Abstract: Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry systems follow a one-shot generation paradigm, which reduces users to prompt providers and weakens th...

📖 Read original article


143. AI emotional support is better only when chosen, but shifts preferences even when it is not ​

Author: Yaoxi Shi, Cathy Mengying Fang, Guy LabanPattie Maes, Amit Goldenberg
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.23196v1 Announce Type: new Abstract: People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's cho...

📖 Read original article


144. Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy ​

Author: Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, Alexandros Potamianos
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23205v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates valuable opportunities to study model behavior not only at the surface level but also at the granularity of individu...

📖 Read original article


145. What is mathematics now, and what should it be? ​

Author: Jeremy Avigad
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, math.HO

arXiv:2608.23218v1 Announce Type: new Abstract: Advances in neural theorem provers have been impressive, but the successes obscure a broader vision of what AI can do for mathematics and how mathematicians can engage with AI. This essay advances a more expansive and optimistic point of view.

📖 Read original article


146. Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data ​

Author: Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that contain reasoning-rich content but lack explicit chain-of-thought annotations. The method trains a model to generat...

📖 Read original article


147. Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records ​

Author: Zeyd Boukhers, Lingxiao Kong, Xenophon Zabulis, Georgios Toubekis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DL

arXiv:2608.23263v1 Announce Type: new Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which every reference is resolvable. The Europeana Data Mode...

📖 Read original article


148. Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance ​

Author: Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibi...

📖 Read original article


149. Apodex 1.1: Scaling Agentic Intelligence for Complex Work ​

Author: Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen, X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Z. Chen, Z. Cheng, Z. Feng, Z. Liang, Z. Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.23283v1 Announce Type: new Abstract: General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. ...

📖 Read original article


150. EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models ​

Author: Xuetong Li, Gaofeng Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is safe for the right multimodal reason. Safelooking behavior may reflect ...

📖 Read original article


151. Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning ​

Author: Zixuan Wang, Yanrui Miao, Zhengxi Lu, Teng Pan, Yiwen Qiu, Hongxing Li, Peng Qiu, Ruiqing Zhang, Yongliang Shen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23318v1 Announce Type: new Abstract: Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guid...

📖 Read original article


152. Walking on the DARKSIDE ​

Author: Aldo Gangemi, Emanuele Bottazzi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2608.23370v1 Announce Type: new Abstract: Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an input rests on a fabricated authority, a misapplied mechanism, or a surreptitious analogy, an unsteered LLM tend...

📖 Read original article


153. Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting ​

Author: Seyed Mohammad Hossein Hashemi, Mohsen Hooshmand, Parvin Razzaghi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, stat.AP

arXiv:2608.23373v1 Announce Type: new Abstract: Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly available surveillance datasets typically pair numeric epidemiological signals with textual information that is noisy, loosely structured, only indirectly...

📖 Read original article


154. MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction ​

Author: Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23397v1 Announce Type: new Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis alone does not show that an agent respected evidence and care-process constraints. We introduce MediSk...

📖 Read original article


155. SkillAlchemy: Open-World Agent Skill Creation ​

Author: Hengjun Wang, Shuyue Wei, Boyi Liu, Jun Yang, Yongxin Tong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23417v1 Announce Type: new Abstract: Agent skills are reusable procedural artifacts that extend language agents with specialized workflows, tool conventions, and domain behaviors at inference time. However, creating reliable skills still depends largely on human authorship, model priors, ...

📖 Read original article


156. Characterizing Necessary Losers to Explain Tournaments Losers ​

Author: Contet Cl'ement, Umberto Grandi, J'er^ome Mengin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23446v1 Announce Type: new Abstract: We study the problem of formally explaining why a candidate was not selected by a given tournament rule, by identifying sub-tournaments in which the candidate loses independently of how the rest of the tournament is completed. We define destructive min...

📖 Read original article


157. StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models ​

Author: Jinghan Tan, Yuanzheng Wang, Lu Chen, Zijun Chen, Yuqian Wang, Maosong Sun
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23475v1 Announce Type: new Abstract: As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstr...

📖 Read original article


158. Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation ​

Author: Naman Garg, Sarika Jain, George Fazekas
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23484v1 Announce Type: new Abstract: We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation through a multi-modal and personalized conversational recommender system. Our submitted system employs a three-st...

📖 Read original article


159. SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning ​

Author: Jialong Liu, Yuling Shi, Ning Yang, Xiaodong Gu, Zuchao Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23493v1 Announce Type: new Abstract: Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-training Large Language Models (LLMs) remains underexplored. We propose Self-R...

📖 Read original article


160. Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty ​

Author: Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23497v1 Announce Type: new Abstract: Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safe...

📖 Read original article


161. EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards ​

Author: Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23525v1 Announce Type: new Abstract: Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, an...

📖 Read original article


162. Correcting a learned physical invariant improves world-model rollouts ​

Author: Richard Bao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23526v1 Announce Type: new Abstract: World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pendulum video learns a scalar that its own latent transition treats as approximately conserved. A label-free searc...

📖 Read original article


163. How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles ​

Author: Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23543v1 Announce Type: new Abstract: While AI assistance can improve human task performance in the short term, it may also undermine the development of skills in the longer term. We examine this tension in a controlled logic-puzzle experiment involving on-demand AI assistance, where parti...

📖 Read original article


164. Prime Agent: A Self-Improving RLM Harness ​

Author: Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian M"uller, Elie Bakouch, Daniel Auras, Mika Senghaas, Fares Obeid, Konstantin Dunas, Johannes Hagemann, Sami Jaghouar
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE

arXiv:2608.23552v1 Announce Type: new Abstract: Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A pe...

📖 Read original article


165. ReWorld: An Interactive World Model with Long-Horizon Memory ​

Author: Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu, Yihua Du, Wei Wang, Tianyi Gui, Lianghua Huang, Yingcong Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23565v1 Announce Type: new Abstract: An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: control wants a short horizon, memory wants an unbounded one. ReWorld separates the two during training and...

📖 Read original article


166. Correcting Variable Importance Scored by Random Forests ​

Author: Guancheng Zhou, Haiping Xu, Jason Liu, Donghui Yan
Published: 8/25/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG

arXiv:2606.10770v1 Announce Type: cross Abstract: Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a variety of tasks such as assisting model interpretation, model selection and diagnosis, and cost-bounded learning e...

📖 Read original article


167. Small Language Model enabled Autonomous agent for Language-Conditioned Cognitive Radar ​

Author: Minhaj Uddin Ahmad, Zakia Zaman, Shunqiao Sun, Mizanur Rahman
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.SY, eess.SY

arXiv:2608.11596v1 Announce Type: cross Abstract: Modern radar systems require adapting their processing strategies in response to changing interference, clutter, and data availability. This paper introduces a framework for a small language model (SLM)-driven autonomous agent designed for language-c...

📖 Read original article


168. Triangular Fuzzy Rescaling Distance ​

Author: Eddy Soria, Aida Valls, Ana Beatriz Hern'andez-Lara
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.GM

arXiv:2608.19234v1 Announce Type: cross Abstract: Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets, particularly Triangular Fuzzy Numbers (TFNs). A crucial aspect of many fuzzy methods is the quantification of ...

📖 Read original article


169. Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation ​

Author: Yi-Chun Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MM

arXiv:2608.21364v1 Announce Type: cross Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly. Incremental interpretation, therefore, depends not only on what is rep...

📖 Read original article


Author: Nimol Thuon
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.21365v1 Announce Type: cross Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage. This paper presents KSE-Web, ...

📖 Read original article


171. PepLLM: ESM-Guided Llama for Structured Protein-Peptide Binding Interface Analysis ​

Author: Hao Qian, Shikui Tu, Lei Xu
Published: 8/25/2026, 4:00:00 AM
Categories: q-bio.BM, cs.AI, cs.LG

arXiv:2608.21367v1 Announce Type: cross Abstract: Protein-peptide interactions are central to cellular regulation and peptide-based drug discovery, yet existing computational methods mainly focus on interaction classification, binding-site prediction, or peptide binder generation. These formulations...

📖 Read original article


172. Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning ​

Author: Stephanie Okoye
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21369v1 Announce Type: cross Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical as...

📖 Read original article


173. On the Role of Citations in Preference Data ​

Author: Yu Hou, Hal Daum'e III, Rachel Rudinger, William Walden
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is u...

📖 Read original article


174. Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models ​

Author: Thantham Jittham
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA

arXiv:2608.21377v1 Announce Type: cross Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjecting LLMs t...

📖 Read original article


175. RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception ​

Author: Oguzhan Baser, Mirac Sozen, Kaan Kale, Sandeep Chinchali, Sriram Vishwanath
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.IT, math.IT

arXiv:2608.21380v1 Announce Type: cross Abstract: With the increased adoption of robotic agents operating in human environments by scanning and sharing 3D representations (e.g., for fleet learning, cloud-based planning, or collaborative mapping), collected point clouds reveal not just the objects in...

📖 Read original article


176. Determinants of Starting Salaries for Filipino Graduates: An Explainable Machine Learning Approach ​

Author: Alexander Gabriel A. Aranes, John Michael C. Magpantay, Reginald Neil C. Recario, Jamlech Iram N. Gojo Cruz, Rodolfo C. Camaclang III
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.21383v1 Announce Type: cross Abstract: Filipino graduates face a persistent disconnect between educational preparation and labor market outcomes, where starting salary is a key signal of entry-level valuation. Current Philippine research is dominated by descriptive tracer studies that doc...

📖 Read original article


177. Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems ​

Author: Ivan Dobrovolskyi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21384v1 Announce Type: cross Abstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity. We quantify this overhead across nine production tokenizers and...

📖 Read original article


178. A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram ​

Author: Michail Zafeiropoulos, Despoina Antonakaki, Sotiris Ioannidis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.SI

arXiv:2608.21385v1 Announce Type: cross Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have n...

📖 Read original article


179. Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning? ​

Author: John C. Howell
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21386v1 Announce Type: cross Abstract: Given a task described by a few examples, how should a model be specialized to it? Four mechanisms are available -- zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork -- yet the operati...

📖 Read original article


180. Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens ​

Author: Alexander Loth, Martin Kappes, Marc-Oliver Pahl
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CR

arXiv:2608.21389v1 Announce Type: cross Abstract: Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human...

📖 Read original article


181. A Survey Instrument to Assess Students' AI and Generative AI Knowledge ​

Author: Aditya Johri, Cory Brozina, Akriti Bagale
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.21391v1 Announce Type: cross Abstract: In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI), including generative artificial intelligence (GenAI), has proliferated, so has the need to educate s...

📖 Read original article


182. ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD ​

Author: Blossom Treesa Bastian, Keerthi S. Shetty, Manish Kolachalam, Rani Malhotra, Ashish Dutta
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.21395v1 Announce Type: cross Abstract: NoMaD [31] is a learned vision-navigation policy that unifies goal-conditioned navigation and exploration in a single goal-masked diffusion policy. In an unseen environment, however - where neither a goal image nor a topological map is available - it...

📖 Read original article


183. Runtime Action Interference for AI Control of AlphaStar in StarCraft II ​

Author: Jaymari Chua, Chen Wang, Liming Zhu, Lina Yao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY, cs.HC, cs.MA

arXiv:2608.21398v1 Announce Type: cross Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions. We contribute \emph{runtime action interference} (RAI), an A...

📖 Read original article


184. Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts ​

Author: Farida Mohsen, Thowayba Elkaffash, Mohammad Reza Chalak Qazani, Mohamed Mabrok, Nader Meskin, Ali Safa
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.21407v1 Announce Type: cross Abstract: Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overh...

📖 Read original article


Author: Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2608.21409v1 Announce Type: cross Abstract: In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent, defined by jurisdiction, temporal validity, and the hierarchy of authoritative sources. The recent ...

📖 Read original article


186. Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding ​

Author: Yisong Xiao, Aishan Liu, Yongxin Huang, Zonghao Ying, Shiji Zhao, Tianlin Li, Yong Han, Jian Yang, Xianglong Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21415v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social g...

📖 Read original article


187. Operational digital twin clinics enable task-based evaluation of embodied AI ​

Author: Xinyuan Wu, Jingrao Zhang, Mengdi Xu, Henry K. Chu, Mingguang He, Danli Shi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.21416v1 Announce Type: cross Abstract: Embodied artificial intelligence (AI) must be tested in the clinical environments where it will operate, but building realistic, robot-testable settings is costly and difficult to scale. Here we show that routine clinic images can be transformed into...

📖 Read original article


188. Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing ​

Author: Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR, cs.MA

arXiv:2608.21423v1 Announce Type: cross Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize ...

📖 Read original article


189. Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation ​

Author: Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu, Chenxi Huang, Yingwei Song, Liyuan Lillian Ma, Yang Ran, Youhua Li, Yongxin Ni
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21425v1 Announce Type: cross Abstract: Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the...

📖 Read original article


190. Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics ​

Author: Ran Chen, Jiaxing Ren, Zhikun Zhang, Yunhao Hou, Junbao Zhuo, Bochao Zou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.21440v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments b...

📖 Read original article


191. Constructing Predictive Surgical Path for AI-based Capsulorhexis Skill Transfer ​

Author: Mohammad Javad Ahmadi, Hamid D. Taghirad
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.21441v1 Announce Type: cross Abstract: Automated training of surgeons is one of the most crucial factors that significantly minimize surgical training risks and expenses. With recent advances in artificial intelligence (AI) knowledge and available data from various surgeries, AI's involve...

📖 Read original article


192. FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows ​

Author: Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz, Devanshu Bansal, Shivani Jain, Nicholas Saban, Chirag Maheshwari, Anand Kannappan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21460v1 Announce Type: cross Abstract: Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue to underperform on subjective and creative design tasks. A major contributor to this performance gap is the lack...

📖 Read original article


193. CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance ​

Author: Erik Thureck, Leo S. R"dian
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.21462v1 Announce Type: cross Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, th...

📖 Read original article


194. Complexity Induction: Compositional Generalization via Structured Label Distortion ​

Author: Aleksandr Abramov
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.21464v1 Announce Type: cross Abstract: We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geometric sha...

📖 Read original article


195. Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs ​

Author: Musa Tur Farazi, K G Subarno Bithi
Published: 8/25/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.21482v1 Announce Type: cross Abstract: Background: Multimodal fracture classifiers may benefit from patient and anatomical metadata, but they can also become brittle when contextual information is missing or mismatched. Methods: We studied 1493 radiographs from the Bangladeshi OrthoFrac-X...

📖 Read original article


196. TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models ​

Author: Chang Sun, Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21487v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diverse tasks. However, during continual adaptation, both catastrophic forgetting and zero-shot degradation occur, s...

📖 Read original article


197. KAN-Robust-Bench: A Benchmark for Evaluating the Robustness of Kolmogorov-Arnold Networks ​

Author: Mohammad Meymani, Roozbeh Razavi-Far
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21488v1 Announce Type: cross Abstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exposed to adversarial threats. While adversarial attacks fall into various categories, the most promin...

📖 Read original article


198. Selection of Heart Sound Segments for Synchronous Classification of Multi-channel Heart Sounds ​

Author: Marcelo Nogueira, Jorge H. Oliveira, Carlos F. Ferreira, Miguel T. Coimbra, Al'ipio M. Jorge
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21499v1 Announce Type: cross Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main auscultation spots. Despite this, automatic heart sound analysis algorithms mostly classify patients using a...

📖 Read original article


199. SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation ​

Author: Yibo Peng, Long Lian, David Wagner, Sizhe Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.21500v1 Announce Type: cross Abstract: Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform ." To prevent arbitrary...

📖 Read original article


Author: Weixuan Ding, Shang Liu, Hanyu Pei, Zeyan Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21543v1 Announce Type: cross Abstract: Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which ...

📖 Read original article


201. Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents ​

Author: Baicheng Chen, Zheyuan Liu, Jingyu Zhang, Kaize Ding, Ningshan Ma, Yue Huang, Meng Jiang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning: previous u...

📖 Read original article


202. Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation ​

Author: Lorenz Brehme, Adam Jatowt
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21558v1 Announce Type: cross Abstract: Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that can assess RAG performance on proprietary data. Existing datasets, such as HotpotQA, challenge current RAG syste...

📖 Read original article


203. Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning ​

Author: Yuyang Luo, Kai Shu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21577v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fairness is a critical safety requirement. In practice, these models are continually updated through continual learning (CL) to adapt to evolving tasks an...

📖 Read original article


204. Power-Performance Characterization of TinyML Systems ​

Author: Yujie Zhang, Dhananjaya Wijerathne, Zhaoying Li, Tulika Mitra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21646v1 Announce Type: cross Abstract: TinyML systems are enabling machine learning (ML) inference at the edge. However, there is little quantitative analysis of such systems. This paper presents a systematic performance and power characterization of diverse TinyML applications on microco...

📖 Read original article


205. Why This, Not That? Mining User Profiles for Pair-wise Counterfactuals ​

Author: Meysam Varasteh, Veronika Bogina, Noam Koenigstein, Robin Burke
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.21662v1 Announce Type: cross Abstract: The topic of explanation in recommender systems has seen steady research attention since the earliest days of the field. With some exceptions, this work has focused on the explanation of single items in a recommendation list and, especially recently,...

📖 Read original article


206. SynEHR: Joint Modeling Inter-visit Temporal Evolution and Intra-visit Clinical Structure for Longitudinal EHR Synthesis ​

Author: Ximiao Li, Lin Jiang, Rongchao Xu, Dahai Yu, Zhe He, Guang Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21673v1 Announce Type: cross Abstract: Longitudinal electronic health records (EHRs) document patients' sequences of clinical visits over time, preserving the temporal evolution of disease progression and care delivery. However, real longitudinal EHRs are difficult to access because they ...

📖 Read original article


207. Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing ​

Author: Anuj Kumar, Heiko Zimmermann, Josiah Bjorgaard, Jacan Chaplais, Nikolaos Bouklas, Matteo Salvador, Alexander Lavin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, physics.app-ph, physics.comp-ph

arXiv:2608.21677v1 Announce Type: cross Abstract: Recent mesh-based simulation advances have, in no small part, relied on neural surrogates of two distinct families: global models that route information through a small set of latent tokens, and local models that perform message passing across mesh e...

📖 Read original article


208. Scalable quantum simulation of continuous-time generative models via tensor networks ​

Author: Nathan X. Kodama, L. Andrew Wray, Sam Cochran, Chad Rigetti, Shravan Veerapaneni, Michael J. Keiser
Published: 8/25/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2608.21700v1 Announce Type: cross Abstract: Continuous-time flow and diffusion models are widely used across many application domains, from large-scale deployment in computer vision and protein folding to emerging adoption for modeling language, time series, and quantum states. After training,...

📖 Read original article


209. Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data ​

Author: Renfei Zhang, Niloofar Mireshghallah
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21727v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifia...

📖 Read original article


210. Architecture as Capability Equalizer for Coding Agents ​

Author: Arquimedes Canedo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2608.21747v1 Announce Type: cross Abstract: LLM-based coding agents generate complete software systems from high-level descriptions, yet little is known about how the format of architecture specifications affects the quality of generated code or whether this effect depends on model capability....

📖 Read original article


211. LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices ​

Author: Riadul Islam, Joey Mule, Dhandeep Challagundla, Shahmir Rizvi, Sean Carson, Rachit Saini
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV

arXiv:2608.21764v1 Announce Type: cross Abstract: Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual signals that reduce redundant data processing and support sustainable edge computing. However, the asynchronous ...

📖 Read original article


212. Evaluation Awareness in Language Models: Representation, Verbalization, and Control ​

Author: Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and c...

📖 Read original article


213. Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web ​

Author: Qijia Chen, Giulio Jacucci
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2608.21794v1 Announce Type: cross Abstract: GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently con...

📖 Read original article


214. SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering ​

Author: Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21796v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visual content. Typically, current methods fuse multimodal features to retrieve external information, subs...

📖 Read original article


215. ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models ​

Author: Maraz Mia, Shovan Roy, Mir Mehedi A. Pritom, Maanak Gupta
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.21803v1 Announce Type: cross Abstract: As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become essential for regulatory compliance and trust. However, the current auditing paradigm relies on an impl...

📖 Read original article


216. More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers ​

Author: Shuai Chen, Tong Bao, Jitong Peng, Chengzhi Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2608.21806v1 Announce Type: cross Abstract: Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, us...

📖 Read original article


217. PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models ​

Author: Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo, Kyungsu Kim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2608.21819v1 Announce Type: cross Abstract: Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement...

📖 Read original article


218. Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores ​

Author: Yu Pan, Hongfeng Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.21829v1 Announce Type: cross Abstract: Retrieval-augmented generation treats the document store as a frozen input, and the systems that instead let an agent curate one never measure what curation does to the store. We invert the framing: the knowledge base is the model. A training agent a...

📖 Read original article


219. LLMs are Few-Shot Decision-Makers: Generalized Context-Aware Microgrid Frequency Control through Prompt Decision Transformer ​

Author: Xu Yang, Chenhui Lin, Haotian Liu, Kaihang Deng, Yunhe Li, Wenchuan Wu
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2608.21858v1 Announce Type: cross Abstract: The rapid evolution of energy structures has positioned microgrids as pivotal components of next-generation power systems, offering enhanced resilience and renewable energy integration. However, the inherent low inertia, complex dynamics, and poor mo...

📖 Read original article


220. ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning ​

Author: Weihang Pan, Zhengxu Yu, Yuxiang Zhang, Wenzhi Li, Zhongming Jin, Binbin Lin, Xiaofei He, Jieping Ye
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21860v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by introducing explicit intermediate reasoning. However, advanced Large Reasoning Models (LRMs) often exhibit over...

📖 Read original article


221. HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning ​

Author: Yucan Guo, Xiaohan Wang, Miao Su, Saiping Guan, Zhongni Hou, Jiajun Chai, Wei Lin, Guojun Yin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21863v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, exis...

📖 Read original article


222. BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications ​

Author: Md Asaduzzaman Jabin, Zihao Wu, Tianming Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.21864v1 Announce Type: cross Abstract: The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often endure lesion noises, modality misalignment, hallucination, and missed contextual grounding in comple...

📖 Read original article


223. GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration ​

Author: Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi, Aranbi Dutta, Aman Chadha, Vinija Jain, Amitava Das
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21869v1 Announce Type: cross Abstract: Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising trajectory toward policy-violating content such as explicit nudity or gr...

📖 Read original article


224. Pruned Traffic Trees: Native Semantic Compression with a Protocol-Structured Model Family for Encrypted Traffic Classification ​

Author: Yuantu Luo, Jun Tao, Xiangyu Xu, Linxiao Yu, Kangying Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.21874v1 Announce Type: cross Abstract: Deep learning has achieved strong performance in encrypted traffic classification (ETC), yet its computational cost limits deployment on resource-constrained network devices such as routers and middleboxes. Existing compression methods mainly operate...

📖 Read original article


225. A Scalable Vector Graphics Latent Space ​

Author: Leonardo Zini, Elia Frigieri, Lorenzo Baraldi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21893v1 Announce Type: cross Abstract: Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a continuous, dense, and invertible latent space for vector representations, the kind of foundational building block th...

📖 Read original article


226. Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks ​

Author: Aaditya Pratap, Harsh Kasyap, Somanath Tripathy
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.21895v1 Announce Type: cross Abstract: Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-served models. Therefore, the safety of LLMs depends on the defense mechanisms used, and their effectiven...

📖 Read original article


227. Bi-EZP: LLM-Guided Bilevel Program Evolution for Ensemble Zero-Cost Proxy Discovery ​

Author: Yutao Lai, Kezhao Lai, Hai-Lin Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.21927v1 Announce Type: cross Abstract: Zero-cost proxies enable neural architecture search (NAS) to rank candidate networks from statistics computed at initialization, avoiding repeated training. However, different proxies capture different properties and often produce inconsistent rankin...

📖 Read original article


228. EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning ​

Author: Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng, Shuo Ren, Jiajun Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.21946v1 Announce Type: cross Abstract: Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy up...

📖 Read original article


229. Bulbul: A Dataset for Dialectal Arabic Speech Recognition ​

Author: Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas, Nada Almarwani, Samah Aloufi, Saad Ezzini, Maged S. Al-Shaibani, Doaa Dalaq, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Mohamed Mehdi Trigui, Dania Refai, Layan Refai, Mohamed Akrout, Mustafa Jarrar, Wasfi G. Al-Khatib, Alaa Dalaq, Darin El-Nakla, Samir Abdaljalil, Abdulrahman Al-Fakih, Nour El Imane Zeghib, Moussa Redah, Salmane Chafik, Mohamed El-Attar, Rima Grati, Sarah Kohail, Malak Alkhorasani, Khadijah Al Safwan, Ismail M. Mudhaffar, Ali Altam, Ahmed Al-Shaikh, Adnan Saeed, Hamzah Luqman
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21950v1 Announce Type: cross Abstract: Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading...

📖 Read original article


230. NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus ​

Author: Elisavet Lydia Alvanaki, Je Yang, Biruk Seyoum, Luca P. Carloni
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2608.21962v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate register-transfer-level (RTL) designs from natural-language specifications. However, assessing functional correctness at early stages remains a fundamental challenge. Existing oracle-free...

📖 Read original article


231. Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments ​

Author: Khang Luong, Nam Nguyen, Hoang Ta, Hung The Tran, Tuan Dam
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.21995v1 Announce Type: cross Abstract: We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration process is dominated by stochastic variance. VarDE is built on a fundamental principle: sampling effor...

📖 Read original article


232. Barycentric Fused Gromov-Wasserstein Balancing for Causal Inference under Multiple Treatments ​

Author: Yuki Murakami, Takumi Hattori, Kohsuke Kubota
Published: 8/25/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG, stat.ML

arXiv:2608.22024v1 Announce Type: cross Abstract: Estimating heterogeneous single and interaction treatment effects from observational data under multiple simultaneous treatments is crucial for decision-making. To mitigate estimation variance, previous studies balance representation distributions be...

📖 Read original article


233. Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets ​

Author: Aashish Panta, Hugo Lee, Giorgio Scorzelli, Kyongsik Yun, Valerio Pascucci
Published: 8/25/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.22045v1 Announce Type: cross Abstract: Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access, and explore. Although many datasets are publicly available, using them often requires familiarity with reposit...

📖 Read original article


234. CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders ​

Author: Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen
Published: 8/25/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.22059v1 Announce Type: cross Abstract: Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited. As the model space expands from general-purpose to broad-medical and specialty-specific encoders, se...

📖 Read original article


235. Discovering Dual-Origin Slow Wind from Solar Orbiter with Self-Supervised Contrastive Learning ​

Author: Henry Han, Jorge Yero Salazar
Published: 8/25/2026, 4:00:00 AM
Categories: astro-ph.SR, cs.AI, cs.LG

arXiv:2608.22065v1 Announce Type: cross Abstract: Whether the slow solar wind originates from one coronal source or two distinct channels remains a central open question in heliophysics. Resolving this requires unsupervised separation of two populations that arrive at nearly the same bulk speed and ...

📖 Read original article


236. ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology ​

Author: Duncan Stothers, Ren-Chin Wu, William Lotter
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.22066v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inference requires applying a large image encoder to every foreground tile despite the subsequent attent...

📖 Read original article


237. Inferring Action from Future Latent State for Robotic Manipulation ​

Author: Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Jie Cheng, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG

arXiv:2608.22067v1 Announce Type: cross Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling....

📖 Read original article


238. Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction ​

Author: Ahmet Tu\u{g}rul Bayrak, Fatma Nur Korkmaz, Bekir Berker T"urker, Mustafa Serta\c{c} T"urkel, Alper Kaplan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22071v1 Announce Type: cross Abstract: Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for turn-ending prediction...

📖 Read original article


239. Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms ​

Author: Khivishta Boodhoo, Josh Plumbly, Nicholas Watson
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22076v1 Announce Type: cross Abstract: Rising energy demand, fossil fuel depletion and climate change highlight the need for more efficient energy production and consumption. Offshore oil and gas platforms face challenges related to inefficient energy use, system failures, accessibility a...

📖 Read original article


240. On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study ​

Author: Daniel Rodriguez-Cardenas, David Nader Palacio, Anna Schmedding, Yiyang Lu, Aadil Mallick, Bill Hudson, Chris Gourley, Michael Roytman, Chris Shenefiel, Evgenia Smirni, Denys Poshyvanyk
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.22089v1 Announce Type: cross Abstract: Modern software systems require earlier and more scalable vulnerability severity assessment to reduce exposure to high-impact security flaws. Security analysts typically assign CVSS scores, but this manual triage does not scale with the growth of dis...

📖 Read original article


241. Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators ​

Author: Yujiao Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22090v1 Announce Type: cross Abstract: Large language models can produce fluent reasoning traces whose local semantic errors propagate to an incorrect conclusion, while unconstrained self-correction may preserve, amplify, or introduce errors. Existing diffusion language models provide ite...

📖 Read original article


242. Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos ​

Author: Xiaoyang Liu, Kai Han
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22102v1 Announce Type: cross Abstract: We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of deformable objects represented by 3D Gaussians. Given a static multi-view scan for geometric initialization, our met...

📖 Read original article


243. TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans - A Case Study in Craniosynostosis 3D Photography ​

Author: Sanjay Bhandari, Nawazish Khan, Alzbeta Novotna, Tiffany Jeong, Loretta Bowman, Michael Hernandez, Tobi Somorin, Viraj Govani, Jesse Goldstein, Shireen Elhabian
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22131v1 Announce Type: cross Abstract: Craniosynostosis severity analysis increasingly relies on statistical shape models (SSMs) to quantify cranial morphology, but most existing workflows depend on computed tomography or heavily curated three-dimensional (3D) photographs. Raw clinical 3D...

📖 Read original article


244. SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning ​

Author: Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CE

arXiv:2608.22132v1 Announce Type: cross Abstract: Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rew...

📖 Read original article


245. Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion ​

Author: Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22140v1 Announce Type: cross Abstract: Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly understood. We evaluate four open-weight instruction-tuned models and frontier models across four reasoning benchma...

📖 Read original article


246. Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints ​

Author: Gwen Yidou-Weng, Edward Sun, Tianyi Ma, Metin Alp Dogan, Benjie Wang, Allen Peng, Guy Van den Broeck, Yuchen Cui
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.22149v1 Announce Type: cross Abstract: LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and existing remedies trade formal guarantees against plan quality: soft methods (affordance scoring, grounded decoding)...

📖 Read original article


247. Why Does Robustness Reduce Superposition? ​

Author: Adam Elimadi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22155v1 Announce Type: cross Abstract: The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial e...

📖 Read original article


248. Joint Causal Structure and Cluster Discovery Using Variational Inference ​

Author: Avni Rajpal, Anubhav Kumar, Rishabh Karnad, Mohammad Emtiyaz Khan, P. K. Srijith
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.22212v1 Announce Type: cross Abstract: Causal discovery aims to understand the relationships between individual random variables. In many applications, such as brain imaging and climate modeling, it is more meaningful to consider interactions among groups of variables. Existing methods as...

📖 Read original article


249. Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN ​

Author: Eliuvish Han Cui
Published: 8/25/2026, 4:00:00 AM
Categories: stat.AP, cs.AI, cs.LG, stat.ML

arXiv:2608.22223v1 Announce Type: cross Abstract: Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broadly with cheaper information, then spend scarce confirmatory amyloid measurements where they support the decisi...

📖 Read original article


250. FreKoo++: Learning Continuous Spectral Dynamics for Temporal Domain Generalization ​

Author: En Yu, Xiaoyu Yang, Wei Duan, Guangquan Zhang, Jie Lu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22224v1 Announce Type: cross Abstract: Temporal Domain Generalization (TDG) aims to learn from historical domains and generalize to unseen future distributions under concept drift. Nevertheless, prevailing TDG methods struggle with complex real-world streaming scenarios involving both mul...

📖 Read original article


251. Improving Few-Step Language Flows with Untied Self-Conditioning ​

Author: Bocheng Li, Linli Xu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22244v1 Announce Type: cross Abstract: Flow-matching language models refine all token positions in parallel and can trade sampling steps for latency, yet generation quality still degrades sharply with few sampling steps. We trace a source of this degradation to a train--inference mismatch...

📖 Read original article


252. Training-Free VLM Personalization via Calibrated Residual Decoding ​

Author: Jiaao Yu, Yujian Ma, Xianming Hu, Pengran Wang, Ang Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22263v1 Announce Type: cross Abstract: Vision-language models can be personalized in a training-free manner by directly providing user profiles, preferences, or visual references at inference time, without updating model parameters. However, direct personalized prompting does not guarante...

📖 Read original article


253. GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets ​

Author: Saif Ahmed, Ashadulla Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Mohammad Abdul Qayum, Mohsin Sajjad, Mohammad Ashrafuzzaman Khan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.22272v1 Announce Type: cross Abstract: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a pretrai...

📖 Read original article


254. Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition ​

Author: Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.22273v1 Announce Type: cross Abstract: Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme r...

📖 Read original article


255. Length-Adaptive Decoding for Masked Diffusion Machine Translation ​

Author: Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22274v1 Announce Type: cross Abstract: Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies ...

📖 Read original article


256. OVIBench: Benchmarking Online Video Question Answering under Interruption ​

Author: Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22279v1 Announce Type: cross Abstract: Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may inte...

📖 Read original article


257. Learning from the Test: Self-Referential Differential Testing for Deep RL Agents ​

Author: Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2608.22284v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) has achieved significant success in complex decision-making problems. As DRL systems are increasingly deployed in real-world applications, ensuring their quality and reliability is paramount. Current works primarily ...

📖 Read original article


258. LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model ​

Author: Ergan Shang, Weijing Tang, Yinqiu He
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22295v1 Announce Type: cross Abstract: Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and ...

📖 Read original article


259. The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction ​

Author: Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju, Hanxiang Ren, Ruizhe Liu, Yu Zhang, Qian Luo, Feng Chen, Pei Zhou, Yi Ma, Yanchao Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.22301v1 Announce Type: cross Abstract: Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and languag...

📖 Read original article


260. Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms ​

Author: Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22335v1 Announce Type: cross Abstract: Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed promp...

📖 Read original article


261. TransHands: Repurposing Human Pose Encoders as Hand Pose Encoders ​

Author: Milo Piccioli, Gianluca Amprimo, Claudia Ferraris, Gabriella Olmo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22341v1 Announce Type: cross Abstract: Lifting 3D hand poses from 2D monocular representations remains challenging due to the limited availability of large-scale, diverse 3D-annotated hand datasets, in contrast to the abundance of human body motion data. We address this limitation by tran...

📖 Read original article


262. Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment ​

Author: Jahangir Alam SM, Md Khalid Syfullah, Saad Ahmed, Munira Akter Mou, A K Z Rasel Rahman, A. K. M. Masudur Rahman, Mohammed Sowket Ali
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22346v1 Announce Type: cross Abstract: This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academi...

📖 Read original article


263. SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models ​

Author: Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao, Ruichu Cai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22354v1 Announce Type: cross Abstract: Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory but potentially becoming unstable under extreme-context extrapolation. By tracking RWKV-7 over sequences of up to 100M tokens, we empirically identify a distinc...

📖 Read original article


264. Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video ​

Author: Masoud Jalayer, Changyi Li, Yu Xiao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.SD

arXiv:2608.22359v1 Announce Type: cross Abstract: Automatically analyzing hours-long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logistics, construction, and manufacturing. Yet current pipelines that process short, fixed-size windows with a visi...

📖 Read original article


265. Self-Supervised Graph Representation Learning for In-The-Wild Wearable and Smartphone based Emotion Recognition ​

Author: Ioannis N. Ziogas, Leontios J. Hadjileontiadis, Ahsan H. Khandoker, Aamna Al Shehhi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2608.22387v1 Announce Type: cross Abstract: Wearable and smartphone-based emotion recognition (WER) remains a challenging setting in affective computing, due to the notorious difficulty and bias associated with in-the-wild label collection. The high inter-and intra-subject emotional variabilit...

📖 Read original article


266. ProBel: Propaganda Detection with Techniques, Spans, and Explanations ​

Author: Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori, Giovanni Da San Martino, Firoj Alam
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22388v1 Announce Type: cross Abstract: Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span identification. However, it remains unclear how supervision at these levels interacts when learned jointly acr...

📖 Read original article


267. KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget ​

Author: Ahmed Abolfadl
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SE

arXiv:2608.22389v1 Announce Type: cross Abstract: Regulation (EU) 2024/886 obliges European payment service providers to settle euro credit transfers in under ten seconds, around the clock. This removes both the overnight batch window in which anti-money-laundering (AML) analytics traditionally ran ...

📖 Read original article


268. Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings ​

Author: Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.22420v1 Announce Type: cross Abstract: Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of potential applications. However, existing methods face considerable challenges in cross-subject decoding, primari...

📖 Read original article


269. Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator ​

Author: Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22432v1 Announce Type: cross Abstract: Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Tur...

📖 Read original article


270. EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting ​

Author: Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.22449v1 Announce Type: cross Abstract: Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based methods typically map observations directly to future motions, overlooking the underlying manipulation process that ...

📖 Read original article


271. Functional compatibility as a determinant of persistent neural learning ​

Author: Hossein Javidnia
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22462v1 Announce Type: cross Abstract: Artificial neural networks can acquire new capabilities but often damage existing ones when they continue to learn. This stability-plasticity problem has motivated replay, regularization and constrained-update methods, yet it remains unclear whether ...

📖 Read original article


272. BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning ​

Author: Md Toufikuzzaman, Ahmad Mousavi, Dongwon Lee
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.22557v1 Announce Type: cross Abstract: Existing LLM unlearning methods struggle with robustness: unbounded forget losses degrade model coherence, fixed-weight balancing cannot adapt as retain difficulty shifts mid-training, and methods that work on one benchmark falter under scaling or re...

📖 Read original article


273. Hybrid Panels: Toward Human-AI Collaboration in Survey Research ​

Author: Julia Romberg, Tobias Gummer, Gabriella Lapesa, Tanja Kunz, Claudia Wagner
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC

arXiv:2608.22582v1 Announce Type: cross Abstract: Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and dat...

📖 Read original article


274. Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support ​

Author: Kushagra Yadav, Nalin Prabhath, Amit Lamba, Goeun Han, Yining Mao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22583v1 Announce Type: cross Abstract: Clinical records contain rich evidence about patient state, but converting that evidence into reliable, structured knowledge graphs remains difficult because extraction errors, ontology mismatch, missing relations, and temporal ambiguity can propagat...

📖 Read original article


275. Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video ​

Author: Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.22586v1 Announce Type: cross Abstract: External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sens...

📖 Read original article


276. AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge ​

Author: Vivek Chavan, J"org Kr"uger
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.HC

arXiv:2608.22617v1 Announce Type: cross Abstract: Assembly and disassembly processes rely on expert knowledge that is difficult to document, reuse, and transfer. This paper presents a data-centric approach for extracting structured task knowledge from expert demonstrations using egocentric and exoce...

📖 Read original article


277. Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains ​

Author: Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW. Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, Parisa Rashidi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22622v1 Announce Type: cross Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) coul...

📖 Read original article


278. GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering ​

Author: Meenu Ravi, Shailik Sarkar, Lulwah AlKulaib, Yordanos Tessema, Chang-Tien Lu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22634v1 Announce Type: cross Abstract: Current work on improving reliability in large language model (LLM)- generated answers has primarily leveraged Retrieval-Augmented Generation (RAG), knowledge-graph augmentation, and reinforcement learning. While these methods are adept at enhancing ...

📖 Read original article


279. Do Not Copy/Paste: Soft Barriers for Copying in AI-Assisted Programming ​

Author: Iyiola E. Olatunji, Alberick Euraste Djire, Jacques Klein, Tegawend'e F. Bissyand'e
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.22638v1 Announce Type: cross Abstract: Copying a function from a chat window into an editor takes less than a second. For many uses of AI coding tools, that speed is the point; in settings such as programming education, code review, and security-sensitive development, it can also be the p...

📖 Read original article


280. Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules ​

Author: Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22642v1 Announce Type: cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. To address these challenges, we present \textb...

📖 Read original article


281. Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code ​

Author: Alberick Euraste Djire, Iyiola E. Olatunji, Melissa Tessa, Earl T. Barr, Jacques Klein, Tegawend'e F. Bissyand'e
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.22652v1 Announce Type: cross Abstract: LLMs are increasingly used for code generation, yet they frequently hallucinate non-existent software packages, creating exploitable entry points into the software supply chain. We make four contributions to this problem. First, we show that prior ev...

📖 Read original article


282. Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs ​

Author: Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA

arXiv:2608.22657v1 Announce Type: cross Abstract: Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate...

📖 Read original article


283. Hyperbolic Hierarchical Clustering for Visual Representation Learning ​

Author: Jianan Wei, Guikun Chen, Zhiyuan Weng, Chunchao Guo, Yujia Wang, Wenguan Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22665v1 Announce Type: cross Abstract: We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental component of modern vision backbones like vision Transformers, facilitating...

📖 Read original article


284. RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation ​

Author: Sen Wang, Yiming Sun, Jiaxuan He, Pengfei Zhu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.22678v1 Announce Type: cross Abstract: UAV vision-language navigation (UAV-VLN) is commonly evaluated as goal reaching, but inspection-oriented deployment requires the agent to stop within a valid inspection region and avoid falsely confirming visually or semantically similar distractors....

📖 Read original article


285. Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing ​

Author: Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.22695v1 Announce Type: cross Abstract: Agent ecosystems now include thousands of MATS components (Models, Agents, Tools, and Skills), yet their discovery still relies on in-context routing. These systems read a registry (names, hints, or descriptions, as context budget permits), pick a ca...

📖 Read original article


286. TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge ​

Author: Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar, Souvik Kundu, Zhishan Guo, Abdullah Al Arafat, Adnan Siraj Rakin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.22716v1 Announce Type: cross Abstract: Despite their remarkable success, machine learning models, particularly in vision applications, are alarmingly vulnerable to a range of security threats. One key factor in the attack landscape is the distinction between white-box and black-box threat...

📖 Read original article


287. DiaRelay: Relaying Dialogue Context with a Constant-Size Memory for Emotion Recognition in Conversation ​

Author: Zihao Zhou, Bin Yang, Jinghui Qin, Kebing Jin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22745v1 Announce Type: cross Abstract: Emotion Recognition in Conversation (ERC) requires models to identify subtle emotional cues that are often distributed across distant dialogue turns. Existing methods typically incorporate dialogue history through a fixed context window. However, sho...

📖 Read original article


288. Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation ​

Author: Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22757v1 Announce Type: cross Abstract: Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they s...

📖 Read original article


289. XTC: Head-Aware Sampling by Excluding Top Choices ​

Author: Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22758v1 Announce Type: cross Abstract: Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or truncating its low-probability tail. These strategies overlook a common regime of open-ended generation in which several con...

📖 Read original article


290. Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time ​

Author: Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.22761v1 Announce Type: cross Abstract: Large Language Models generate text autoregressively, but open-ended generation is prone to verbatim looping, in which models repeat spans already present in context. Standard defenses such as repetition, presence, and frequency penalties and n-gram ...

📖 Read original article


291. TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents ​

Author: Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22793v1 Announce Type: cross Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fu...

📖 Read original article


292. Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation ​

Author: Jianxiang Liu, Gaojing Zhang, Chuan Wen, Qipeng Liu, Yuxuan Zhao, Ning Guo, Wenzhao Lian
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.22800v1 Announce Type: cross Abstract: Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but t...

📖 Read original article


293. SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support ​

Author: Ahnaf Atef Choudhury, Ramkrishna Saha
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22802v1 Announce Type: cross Abstract: Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in...

📖 Read original article


294. Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization ​

Author: Sunhee Hwang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22820v1 Announce Type: cross Abstract: Deep learning models often produce performance disparities across demographic groups, due to the training data imbalance with respect to sensitive attributes such as gender or age. To address this problem, existing work has explored fair representati...

📖 Read original article


295. Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments ​

Author: Ryuki Hyodo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.22833v1 Announce Type: cross Abstract: Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordinary computers. We p...

📖 Read original article


296. Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning ​

Author: Shuting Xie, Nathaniel Lesperance, Graham W. Taylor
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22839v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for scientific decision support, yet reliable confidence estimation remains difficult in black-box settings. We study uncertainty estimation for hierarchical taxonomic reasoning generated by a black-...

📖 Read original article


297. The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models ​

Author: Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22876v1 Announce Type: cross Abstract: We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight audit, two forward passes, no training or gradients, that localizes exactly where causality breaks. Attention-mask inspection is inc...

📖 Read original article


298. AraDetox: A Multi-Dialect Arabic Detoxification Dataset ​

Author: Mo El-Haj
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22894v1 Announce Type: cross Abstract: Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 ...

📖 Read original article


299. Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text ​

Author: Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22908v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent advances, existing SLMs still exhibit weaker instruction-following behavior and limited generalization across d...

📖 Read original article


300. Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling ​

Author: Akifumi Wachi, Takumi Tanabe, Youhei Akimoto
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR

arXiv:2608.22915v1 Announce Type: cross Abstract: Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that this composition creates a two-stage failure: an imperfect safety prox...

📖 Read original article


301. Deep Learning-Based Multi-User Communication Design for Dense IoT Networks: Interference-Aware Finite-Blocklength Communication and Preliminary MIMO Extensions ​

Author: Arkadeep Sinha, Shubham Paul, R. Manivasakan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT

arXiv:2608.22923v1 Announce Type: cross Abstract: Dense IoT networks require reliable communication despite limited spectrum and substantial multi-user interference while maintaining manageable receiver complexity. This work introduces a deep-learning-based end-to-end multi-user communication design...

📖 Read original article


302. What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation ​

Author: Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Hongyao Zuo (Tianjin University), Ziwen Gong (Hainan University), Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Tong Yang (Peking University), Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Xiaohui Li (Huawei Technologies), Haoli Bai (Huawei Technologies)
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.22948v1 Announce Type: cross Abstract: Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery o...

📖 Read original article


303. WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans ​

Author: Jun Zhang, Qiao Zhao, Cheng Cui, Jianying Qu, Zhongkai Sun, Jianwen Yang, Changda Zhou, ZhuoXin Liu, Shubin Han
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.22959v1 Announce Type: cross Abstract: While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formu...

📖 Read original article


304. Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval ​

Author: Kishore Konda
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.22980v1 Announce Type: cross Abstract: Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat an embedding as an indivisible point in a high-dimensional space. In this work, we propose the Hypergraph Embedd...

📖 Read original article


305. A Physical Response-and-Memory Model for Muon Optimization ​

Author: Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cond-mat.stat-mech, cs.AI, physics.comp-ph

arXiv:2608.22994v1 Announce Type: cross Abstract: Training large language models is costly. How low a loss the same compute can ultimately reach depends on how each step's gradient is converted into a weight update; the rule that performs this conversion is the optimizer. From SGD and AdamW to the r...

📖 Read original article


306. Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG ​

Author: Zhe Jin, Zhimin Lin, Bin Zheng, Junhua Fang, Huihua Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23011v1 Announce Type: cross Abstract: Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing their retrieval index. We argue ...

📖 Read original article


307. SplitLite: Low-Rank Residual Compression for Split Learning ​

Author: Tao Li, Yulin Tang, Qi Guo, Xianhao Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23018v1 Announce Type: cross Abstract: Federated fine-tuning of on-device large language models (LLMs) faces a significant computing burden. To overcome this limitation, split learning (SL) has emerged as a promising solution, which offloads the primary training workload to a powerful ser...

📖 Read original article


308. Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality ​

Author: Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong, Zhaokun Wang, Wenyi Li, Shiyao Guo, Jinyu Guo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23020v1 Announce Type: cross Abstract: Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise forgetting without compromising general utility remains challenging. Existing sequence- and token-level methods p...

📖 Read original article


309. Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs ​

Author: Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23026v1 Announce Type: cross Abstract: Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language...

📖 Read original article


310. FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning ​

Author: Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23031v1 Announce Type: cross Abstract: Federated Learning (FL) enables distributed clients to collaboratively train models without sharing raw data, making it promising for leveraging massive devices in communication networks. In distillation-based FL, each client applies its local model ...

📖 Read original article


311. The Multilingual FrameNet Corpus ​

Author: Beatrice Fiuman`o, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23037v1 Announce Type: cross Abstract: This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing existing language-specific corpora across nine additional languages: Brazilian Portuguese,...

📖 Read original article


312. Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking ​

Author: Abdul Ghafoor, Muhammad Arslan Manzoor, Yufang Hou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23047v1 Announce Type: cross Abstract: Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Inco...

📖 Read original article


313. Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia ​

Author: Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.IR, cs.MM

arXiv:2608.23065v1 Announce Type: cross Abstract: Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recogniz...

📖 Read original article


314. Shaping the Evolutionary Dynamics of Robot Morphology via Adaptive Control Learning ​

Author: Junru Song, Yang Yang, Yaqing Xu, Ying Wen, Wei Peng, Guozhen Li, Wei'en Zhou, Wen Yao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.23100v1 Announce Type: cross Abstract: Robot co-design via bi-level optimization couples within-lifetime controller learning for fitness evaluation with cross-generational morphological evolution. Prior work has established that well-adapted morphology facilitates faster control learning,...

📖 Read original article


315. PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors ​

Author: Nathan Duboisset, Zhaolan Huang, Felix Bie{\ss}mann, Roudy Dagher, Antoine Lavandier, Emmanuel Baccelli
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23101v1 Announce Type: cross Abstract: Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However...

📖 Read original article


316. Molecular LLM Agents: From Architectural Design to Scientific Autonomy ​

Author: Jiatong Li, Wengyu Zhang, Weida Wang, Yuxuan Ren, Wei Liu, Chenyang Mao, Yuqiang Li, Yatao Bian, Changmeng Zheng, Xiaoyong Wei, Qing Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23104v1 Announce Type: cross Abstract: Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon chemical objects acros...

📖 Read original article


317. DeMixPert: Decomposed Response Modeling with Gaussian Mixtures for OOD Single-Cell Perturbation Prediction ​

Author: Jiawen Liu, Xuechenxiao Cao, Yutong Li, Bing Liu, Jiaming Liang, Tinghe Zhang, Xiaoqi Sheng, Hongmin Cai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23114v1 Announce Type: cross Abstract: Predicting transcriptome-wide responses to unseen genetic perturbations remains a major computational challenge because accurate prediction requires recovering both perturbation-specific transcriptional shifts and heterogeneous cellular responses. Ex...

📖 Read original article


318. Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study ​

Author: Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23120v1 Announce Type: cross Abstract: Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the E...

📖 Read original article


319. LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space ​

Author: Jinghui Zhang, Lang Gao, Ao Li, Mingzhe Li, Ruihong Zeng, Zirui Song, Kentaro Inui, Xiuying Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23124v1 Announce Type: cross Abstract: Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, existing approaches to author modeling and personal...

📖 Read original article


320. Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation ​

Author: Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.23138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-...

📖 Read original article


321. Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization ​

Author: Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23149v1 Announce Type: cross Abstract: The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-Lingual Ranking Preference Optimization (CRPO), a n...

📖 Read original article


322. Counterfactual Transition Graphs: Evaluating Cross-Class Transition Quality ​

Author: Syed Muhammad Hamza Zaidi, Szymon Bobek, Grzegorz J. Nalepa, Myra Spiliopoulou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23164v1 Announce Type: cross Abstract: Counterfactual (CF) explanations for time-series classifiers are usually evaluated one example at a time: what minimal edit flips this single window's prediction? We argue that the more informative question for diagnostic interpretability is structur...

📖 Read original article


323. NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration ​

Author: Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui
Published: 8/25/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.23179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are poorly understood. An essential prerequisite is to assess such agents in a realistic but risk-free envi...

📖 Read original article


324. BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations ​

Author: Taqi Hamoda, Hayat Rajani, Nuno Gracias
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23215v1 Announce Type: cross Abstract: Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations that inextricably mix intrinsic seabed reflectivity with transient viewing geometries. Existing self-supervised le...

📖 Read original article


325. AI Surrogate Modeling for Real-Time Tokamak Equilibrium Prediction: Benchmarking Neural Architectures and Validation on EXL-50U ​

Author: Guoyang Shi, Zitong Zhang, Siqi Ding, Jianguo Chen, Yapeng Zhang, Jiayi Zhi, Hanyue Zhao, Tianyuan Liu
Published: 8/25/2026, 4:00:00 AM
Categories: physics.plasm-ph, cs.AI

arXiv:2608.23217v1 Announce Type: cross Abstract: Fast and reliable plasma equilibrium prediction is essential for real-time tokamak operation and control, but conventional Grad-Shafranov (GS) solvers are often too costly for real-time deployment. We develop an AI surrogate framework and benchmark f...

📖 Read original article


326. Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation ​

Author: Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen, Xiaojun Zhu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.23224v1 Announce Type: cross Abstract: Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces m...

📖 Read original article


327. Retrieval-Augmented Classification of Environmental Mitigations in Hydropower Licensing Documents ​

Author: Hong-Jun Yoon, Tom Ruggles, Joanna Lee, Debjani Singh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.23241v1 Announce Type: cross Abstract: Identifying and classifying environmental mitigation obligations in Federal Energy Regulatory Commission hydropower licensing documents is a labor-intensive task requiring deep domain expertise. We formulate this as a multi-label classification probl...

📖 Read original article


328. Credal Large Language Models for Semantic Commitment under Uncertainty ​

Author: Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, stat.ML

arXiv:2608.23244v1 Announce Type: cross Abstract: Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a single predictive distribution, conflating epistemic ignorance with genuin...

📖 Read original article


329. Multi-Winner Voting with Argumentative Ballots ​

Author: Ryuta Arisaka, Hirotaka Ono
Published: 8/25/2026, 4:00:00 AM
Categories: cs.GT, cs.AI

arXiv:2608.23247v1 Announce Type: cross Abstract: We introduce multi-winner voting with argumentative ballots (MVArg) and investigate theoretical properties. As our conceptual contribution, we generalise approval ballots to argumentative ballots, thereby allowing voters to express defeasible prefere...

📖 Read original article


330. Future Querying: Can LLMs Serve as Implicit Medical World Models? ​

Author: Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx, Philippe Modard, Elke Giets, Ludovic Denoyer
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23248v1 Announce Type: cross Abstract: Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured text. To address this, we introduce future querying, a paradigm that probes whether large language m...

📖 Read original article


331. E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models ​

Author: Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang, Shang Gao, Hualong Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23253v1 Announce Type: cross Abstract: Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Existing pruning methods largely rely on attention scores and directly aggregate outputs across attentio...

📖 Read original article


332. How Much Regularization Survives Averaging? Update Masking in Federated Learning ​

Author: Wenhao Yan, Fu Kuroda, Yucheng Jin, Zhenke Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23286v1 Announce Type: cross Abstract: Federated learning on non-IID data seeks flat minima to generalize across clients, and existing methods borrow sharpness-aware minimization from centralized training. There is a second way to reach flat minima, in which the regularization comes for f...

📖 Read original article


333. Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction ​

Author: Isaac (Rucheng), Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23296v1 Announce Type: cross Abstract: Learned KV-cache eviction often faces a soft-to-hard mismatch: during training, differentiable gates typically attenuate token contributions, whereas inference saves memory only when KV entries are physically removed. We ask whether the attention sub...

📖 Read original article


334. Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty ​

Author: Changdao He, Yibing Ju, Jonathan Calver, Alice Gao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.23300v1 Announce Type: cross Abstract: Algorithmic solver effort is often assumed to align with perceived puzzle difficulty, but this assumption is rarely tested against human solving data. We evaluate this assumption for Nonograms, a popular logic puzzle similar to Sudoku in which numeri...

📖 Read original article


335. FIDES: A Concordance Protocol for LLM-Generated Trading Strategies ​

Author: Arther Tian, Alex Ding, Simon Wu, Aaron Chan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23308v1 Announce Type: cross Abstract: An LLM asked for a trading strategy returns three artifacts at once: a natural-language rationale, an executable implementation, and once run, a track record. Whether these are the same object is rarely checked. We present FIDES, a measurement protoc...

📖 Read original article


336. Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation ​

Author: Mohammad Mahdi Dehshibi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2608.23323v1 Announce Type: cross Abstract: Continuous optimisation methods need to balance sharing information and maintaining alternative search directions. In this paper, we introduce Mycelial Search (Myco), a graph-structured metaheuristic designed around active tips, community-weighted fl...

📖 Read original article


337. Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents ​

Author: Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23329v1 Announce Type: cross Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Rese...

📖 Read original article


338. The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning ​

Author: Matthew Perlman, Atharva Nijasure, James Allan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.23338v1 Announce Type: cross Abstract: LoRA fine-tuning is standard for adapting LLMs to reranking, but it remains unclear where in the network task-specific relevance behavior is learned and what attention-level changes accompany that learning. Through ablation and attention experiments,...

📖 Read original article


339. DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts ​

Author: Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke, Radu Tudor Ionescu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23363v1 Announce Type: cross Abstract: Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-...

📖 Read original article


340. Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ​

Author: Nikita Kezins
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23375v1 Announce Type: cross Abstract: Gumbel-based inference verification bounds LLM weight exfiltration by only forgiving token choices that plausibly arise from honest GPU nondeterminism, reporting a >200x slowdown for a steganographic adversary under benign prompt traffic. This bound ...

📖 Read original article


341. Cross-lingual Biography Enrichment via Claim Extraction and Alignment ​

Author: Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23390v1 Announce Type: cross Abstract: English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures. We study cross-lingual biography enrichment: enriching an existing Englis...

📖 Read original article


342. Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data ​

Author: Yifei Song, Kun Efimov-Zhang, Claire Gardent
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23391v1 Announce Type: cross Abstract: Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different generation tasks. However, most prior work on data-to-text (D2T) generation has focused on specific tasks and ...

📖 Read original article


343. Towards a Densing Law for User Representation Learning at Billion-Scale Capacity ​

Author: Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang, Letian Gong, Baokun Wang, Huan Li, Yu Cheng, Weiqiang Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.23392v1 Announce Type: cross Abstract: User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-sca...

📖 Read original article


344. Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep ​

Author: Pedro Santos
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.SE

arXiv:2608.23395v1 Announce Type: cross Abstract: Recent LLM-agent systems make conflicting design bets: decompose work across many narrow agents, or use one strong tool-using agent. This pilot studies that choice on bounded cross-border VAT determination with reverse charge, where every case has an...

📖 Read original article


345. Adaptive Item-based Collaborative Structures via Noise Rescheduling in Diffusion for Generative Recommendation ​

Author: Jiaqi Wang, Tianying Liu, Heng Chang, Jihong Guan, Wengen Li, Shuigeng Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.23400v1 Announce Type: cross Abstract: Discrete Diffusion Models (DDMs) have recently been introduced to recommendation systems, modeling user history as a token generation process via iterative denoising. However, while effective at capturing user-level sequential patterns, these methods...

📖 Read original article


346. ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation ​

Author: Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo, Xueliang Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23429v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have shown strong performance in high-fidelity image generation, but their sampling process remains computationally intensive due to full model execution at every timestep. While cache-based acceleration has been explore...

📖 Read original article


347. Towards Comprehensive Basketball Understanding ​

Author: Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, Weidi Xie
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23435v1 Announce Type: cross Abstract: Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions amon...

📖 Read original article


348. Reward-Free Continual Adaptation for Resilient Space Robots ​

Author: Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.23452v1 Announce Type: cross Abstract: Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access...

📖 Read original article


349. Machine Learning Assisted Inverse Design of Pixelated mmWave Patch Antennas ​

Author: Nadeem Rather, Holger Claussen, Lester Ho
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2608.23469v1 Announce Type: cross Abstract: A machine learning-assisted framework for the inverse design of pixelated millimetre-wave patch antennas targeting the 22--30 GHz band is presented. The antenna surface is represented as a 19x23 binary pixel grid on a Rogers RT/duroid 5880 substrate,...

📖 Read original article


350. InjecMEM: Memory Injection Attack on LLM Agent Memory Systems ​

Author: Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang, Yuhang Liu, Zhehao Huang, Kun Yang, Xiaolin Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23471v1 Announce Type: cross Abstract: Memory is becoming a default subsystem in deployed LLM agents to provide persistent personalization and continuity. This naturally prompts a question: will memory system introduce new vulnerabilities into agents? Thus we propose InjecMEM, a novel mem...

📖 Read original article


351. MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters ​

Author: ChengAo Shen, Wenchao Yu, Fangyu Wu, Dongjin Song, Hanghang Tong, Dongsheng Luo, Wei Cheng, Haifeng Chen, Jingchao Ni
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23473v1 Announce Type: cross Abstract: Time series forecasting (TSF) is evolving toward multimodal and agentic settings, yet using foundation models remains uneconomical in resource-constrained scenarios, where compact, specialized forecasters are more desirable. However, lightweight fore...

📖 Read original article


352. What's the Catch? Evaluating Temporal Consistency in Vision-Language Models ​

Author: Marek Hradil, Danae S'anchez Villegas
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2608.23474v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, ...

📖 Read original article


353. Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models ​

Author: Sangoh Lee, Sangwoo Mo, Wook-Shin Han
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.23478v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objecti...

📖 Read original article


354. When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World ​

Author: Xiang Chen, Zeyu Zhang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23507v1 Announce Type: cross Abstract: Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identity reconciliation more than a problem of string ma...

📖 Read original article


355. The Measurement Revolution? Credible Measurement and Inference in the Age of AI ​

Author: Melissa Dell, Ashesh Rambachan
Published: 8/25/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, q-fin.EC, stat.AP

arXiv:2608.23524v1 Announce Type: cross Abstract: Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the b...

📖 Read original article


356. Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition ​

Author: Kyle Stein, Guillermo Francia, III Eman El-Sheikh, Andrew Arash Mahyari
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23536v1 Announce Type: cross Abstract: The continual evolution of malware variants necessitates detection systems that can adapt to new threats without retraining from scratch. However, continually updating models on new data often leads to catastrophic forgetting, where previously learne...

📖 Read original article


357. The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams ​

Author: Summer Eunhyung Ann, Haokun Liu, Chenhao Tan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.23541v1 Announce Type: cross Abstract: Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops (Chen et al., 2025), and mixture-of-agents synthesis (Wang et al., 2025), while other work finds that interaction adds cost without i...

📖 Read original article


358. ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings ​

Author: Na Li, Yuchen Jiao, Changxiao Cai, Gen Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, stat.ML

arXiv:2608.23551v1 Announce Type: cross Abstract: Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow t...

📖 Read original article


359. Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography ​

Author: Yuanyuan Zhang, Yida Zhang, Jiahui Li, Yuyan Wu, Fei Dou, Xiao Yin, Zhenlin An, Hae Young Noh, Wenzhan Song
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, physics.bio-ph

arXiv:2608.23562v1 Announce Type: cross Abstract: Ballistocardiography (BCG) is promising for unobtrusive long-term blood pressure (BP) monitoring in laboratory settings, but traditional BCG signals are vulnerable to the variations in body-bed interaction with shifted fiducial points in temporal or ...

📖 Read original article


360. EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings ​

Author: Md Thamed Bin Zaman Chowdhury, Moazzem Hossain
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23563v1 Announce Type: cross Abstract: Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To a...

📖 Read original article


361. SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? ​

Author: Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SE

arXiv:2608.23564v1 Announce Type: cross Abstract: Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing ben...

📖 Read original article


362. How to Train a Critic Stably and Efficiently ​

Author: Penghui Qi, Xiangxin Zhou, Wee Sun Lee
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.23566v1 Announce Type: cross Abstract: Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard ...

📖 Read original article


363. A Survey on Human-AI Collaboration with Large Foundation Models ​

Author: Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC

arXiv:2403.04931v4 Announce Type: replace Abstract: As the capabilities of artificial intelligence (AI) continue to expand rapidly, Human-AI (HAI) Collaboration, combining human intellect and AI systems, has become pivotal for advancing problem-solving and decision-making processes. The advent of La...

📖 Read original article


364. Memory-Enhanced Neural Solvers for Routing Problems ​

Author: Felix Chalumeau, Refiloe Shabe, Noah De Nicola, Arnu Pretorius, Thomas D. Barrett, Nathan Grinsztajn
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2406.16424v4 Announce Type: replace Abstract: Routing Problems are central to many real-world applications, yet remain challenging due to their (NP-)hard nature. Amongst existing approaches, heuristics often offer the best trade-off between quality and scalability, making them suitable for ind...

📖 Read original article


365. Evaluating Large Language Models for automatic analysis of teacher simulations ​

Author: David de-Fitero-Dominguez, Mariano Albaladejo-Gonz'alez, Antonio Garcia-Cabot, Eva Garcia-Lopez, Antonio Moreno-Cediel, Erin Barno, Justin Reich
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2407.20360v2 Announce Type: replace Abstract: Digital Simulations (DS) provide safe environments where users interact with an agent through conversational prompts, providing engaging learning experiences that can be used to train teacher candidates in realistic classroom scenarios. These simul...

📖 Read original article


366. Online design of dynamic networks ​

Author: Duo Wang, Andrea Araldo, Mounim El Yacoubi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SI, physics.soc-ph

arXiv:2410.08875v4 Announce Type: replace Abstract: Designing a network (e.g., a telecommunication or transport network) is mainly done offline, in a planning phase, prior to the operation of the network. On the other hand, a massive effort has been devoted to characterizing dynamic networks, i.e., ...

📖 Read original article


367. Neural-Symbolic Reasoning over Knowledge Graphs: A Survey from a Query Perspective ​

Author: Lihui Liu, Zihao Wang, Hanghang Tong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2412.10390v2 Announce Type: replace Abstract: Knowledge graph reasoning is pivotal in various domains such as data mining, artificial intelligence, the Web, and social sciences. These knowledge graphs function as comprehensive repositories of human knowledge, facilitating the inference of new ...

📖 Read original article


368. Practical Principles for AI Cost and Compute Accounting ​

Author: Stephen Casper, Luke Bailey, Tim Schreier
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2502.15873v5 Announce Type: replace Abstract: Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory requirements for models or developers that are contingent on specific thresholds. However, technical ambigui...

📖 Read original article


369. Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction ​

Author: Yiming Xu, Junfeng Jiao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2508.17527v2 Announce Type: replace Abstract: Accurately predicting travel mode choice is essential for effective transportation planning, yet traditional statistical and machine learning models are constrained by rigid assumptions, limited contextual reasoning, and reduced transferability. Th...

📖 Read original article


370. Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap ​

Author: Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2508.18646v3 Announce Type: replace Abstract: Despite their rapid advancement, large language models (LLMs) suffer from a critical disconnect between benchmark scores and real-world utility. Current evaluation remains fragmented, prioritizing isolated technical metrics over the holistic, devel...

📖 Read original article


371. Beyond Benchmarking: Scenario-Based Evaluation of Large Language Models for Personalized Learning ​

Author: Bo Yuan, Jiazi Hu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.05346v3 Announce Type: replace Abstract: While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of how their pedagogical behaviors differ in authentic learning scenarios. Existing evaluation practices often e...

📖 Read original article


372. MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM ​

Author: Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.20067v5 Announce Type: replace Abstract: Large language models (LLMs) have shown promise in supporting medical diagnosis, with prompting-based methods offering a flexible and deployable means of capability enhancement. However, existing prompt engineering and multi-agent approaches often ...

📖 Read original article


373. AdaR: A Framework for Equipping LLMs with Adaptive Reasoning ​

Author: Zhejian Lai, Xiang Geng, Zhijun Wang, Yang Bai, Jiahuan Li, Rongxiang Weng, Jingang Wang, Xuezhi Cao, Xunliang Cai, Shujian Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.04617v3 Announce Type: replace Abstract: Mathematical reasoning is a primary indicator of large language models (LLMs) intelligence. However, existing LLMs exhibit failures in robustness and generalization. This paper attributes these deficiencies to spurious reasoning, wherein generated ...

📖 Read original article


374. Training Proactive and Personalized LLM Agents ​

Author: Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, Yiming Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward training agents as collaborators that communicate and adapt to people. To facilitate this shift in real-world compl...

📖 Read original article


375. SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation ​

Author: Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.04529v3 Announce Type: replace Abstract: Creating presentation slides from scientific papers is not simply a matter of summarizing paragraphs. A presenter is required to decide what story to tell, which figures and equations to highlight, and how to arrange them into pages that are visual...

📖 Read original article


376. VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection ​

Author: Jiahao Xie, Guangmo Tong
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2602.13880v2 Announce Type: replace Abstract: Graph property detection aims to determine whether a graph exhibits certain structural properties, such as being Hamiltonian. Recently, learning-based approaches have shown great promise by leveraging data-driven models to detect graph properties e...

📖 Read original article


377. SkillNet: Create, Evaluate, and Connect AI Skills ​

Author: Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Yida Xue, Xin Xu, Tongtong Wu, Kun Wang, Yang Liu, Zhen Bi, Jungang Lou, Yuchen Eleanor Jiang, Hangcheng Zhu, Gang Yu, Haiwen Hong, Longtao Huang, Hui Xue, Chenxi Wang, Yijun Wang, Zifei Shan, Xi Chen, Zhaopeng Tu, Feiyu Xiong, Xin Xie, Peng Zhang, Zhengke Gui, Lei Liang, Jun Zhou, Chiyu Wu, Jin Shang, Yu Gong, Junyu Lin, Changliang Xu, Hongjie Deng, Wen Zhang, Keyan Ding, Qiang Zhang, Fei Huang, Ningyu Zhang, Jeff Z. Pan, Guilin Qi, Haofen Wang, Huajun Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG, cs.MA

arXiv:2603.04448v3 Announce Type: replace Abstract: Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Without a unified mechanism for skill consolidation, agents frequently...

📖 Read original article


378. Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary ​

Author: Dane Malenfant
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.06813v2 Announce Type: replace Abstract: In a stationary decentralized Markov game, learning peers generate an episode-indexed sequence of induced MDPs for any focal agent. The joint game remains stationary while the focal agent's rewards and dynamics drift, forming an agent-centric conti...

📖 Read original article


379. ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation ​

Author: Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.29902v2 Announce Type: replace Abstract: Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation or retrieval augmen...

📖 Read original article


380. Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints ​

Author: Minh-Khoi Pham, Thang-Long Nguyen Ho, Thao Thi Phuong Dao, Tai Tan Mai, Minh-Triet Tran, Marie E. Ward, Una Geary, Rob Brennan, Nick McDonald, Martin Crane, Marija Bezbradica
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.01841v3 Announce Type: replace Abstract: Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbalance, and distribution shift. While tabular in-context learning (TICL) and retrieval-augmented methods perform...

📖 Read original article


381. Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence ​

Author: Niklas Herbster, Martin Zborowski, Alberto Tosato, Gauthier Gidel, Tommaso Tosato
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.08169v3 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization. Recent evidence suggests that some misalignment behaviors are encode...

📖 Read original article


382. MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction ​

Author: Wenchang Duan, Zhenguo Gao, Jinguo Xian, Yi Shi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.10169v4 Announce Type: replace Abstract: Trajectory prediction is a key component of autonomous driving systems because future motions directly affect collision checking, behavior planning, and control. The task remains challenging under dense interactions, heterogeneous behaviors, multim...

📖 Read original article


383. Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling ​

Author: Ocean Monjur, Shahriar Kabir Nahin, Anshuman Chhabra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2604.25098v3 Announce Type: replace Abstract: Large Language Models (LLMs) now exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), with impressive performance across math and coding benchmarks. In parallel, research in model compression has developed pruning meth...

📖 Read original article


384. FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models ​

Author: Seunghan Lee, Jun Seo, Jaehoon Lee, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2605.03460v5 Announce Type: replace Abstract: Time series (TS) reasoning models (TSRMs) have shown promising capabilities in general domains, yet they consistently fail on financial domain, which exhibit unique characteristics. We propose a general 2 x 2 capability taxonomy for TSRMs by crossi...

📖 Read original article


385. Reconciling Consistency-Based Diagnosis with Actual-Causality-Based Explanations ​

Author: Leopoldo Bertossi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.LO

arXiv:2605.08688v2 Announce Type: replace Abstract: We establish, from the point of view of Explainable AI (XAI), connections between Consistency-Based Diagnosis (CBD), on one side, and Actual Causality and Causal Responsibility, on the other. CBD has received little attention from the XAI community...

📖 Read original article


386. CogniFold: Always-On Proactive Memory via Cognitive Folding ​

Author: Suli Wang, Yiqun Duan, Yu Deng, Rundong Zhao, Dai Shi, Minghua Deng, Chen Chen, Yiqi Wang, Xinliang Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2605.13438v5 Announce Type: replace Abstract: Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "al...

📖 Read original article


387. GIM: Evaluating models via tasks that integrate multiple cognitive domains ​

Author: Rohit Patel, Alexandre Rezende, Steven McClain
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2605.18663v2 Announce Type: replace Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates memorizati...

📖 Read original article


388. TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization ​

Author: Isabella A. Stewart, Hongrui Chen, Faez Ahmed
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.21622v2 Announce Type: replace Abstract: Topology optimization can generate efficient structures, but designers often must manually translate qualitative intent, such as desired visual style, product experience, or manufacturability into solver settings that are not directly tied to those...

📖 Read original article


389. Designing Benchmarks for Knowledge Work ​

Author: Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.23262v2 Announce Type: replace Abstract: AI agents are moving quickly from answering isolated questions toward completing work through tools, software environments, and multi-step workflows. Much of what these systems are now asked to do is knowledge work, where information and expertise ...

📖 Read original article


390. MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition ​

Author: Yifan Bao, Xinyu Xi, Xinyu Liu, Wen Ge, Lei Jiang, Kevin Zhang, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.00708v2 Announce Type: replace Abstract: Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation protocol, and refinement strategy for a task. AutoML systems automate...

📖 Read original article


391. Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs ​

Author: Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang, Dexu Yu, Ronghao Chen, Yang Zhou, Hongwu Peng, Xuanqi Lan, Dimitris N. Metaxas, Youhua Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.00726v3 Announce Type: replace Abstract: Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures a...

📖 Read original article


392. MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention ​

Author: Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang, Chenghao Yu, Hongxia Xie, Wen-Huang Cheng, Jianlong Fu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.01063v3 Announce Type: replace Abstract: Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate this ability through offline question answering or scenario-level action prediction. MindPower advanc...

📖 Read original article


393. Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification ​

Author: Shihao Ji, Haotao Tan, Zihui Song, Mingyu Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.01160v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used with formal interactive theorem provers such as Lean 4. Scaling these systems with reinforcement learning or search methods requires process reward models (PRMs) that can evaluate intermediate reas...

📖 Read original article


394. Zero knowledge verification for frontier AI training is possible ​

Author: Pierre Peign'e, Ky Nguyen, Paul Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2606.05433v2 Announce Type: replace Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists. Any...

📖 Read original article


395. How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions ​

Author: Donghao Huang, Tomas Drietomsky, Benjamin Barrett, Zhaoxia Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.08051v2 Announce Type: replace Abstract: Merchant information extraction turns noisy financial transaction descriptors into structured fields at production scale. Our deployed LoRA-fine-tuned LLaMA~3.1-8B reaches 96.95% F1, but its memory and throughput motivate smaller replacements. We ...

📖 Read original article


396. Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control ​

Author: Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, physics.flu-dyn

arXiv:2606.08405v2 Announce Type: replace Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific control design in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control a...

📖 Read original article


397. Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models ​

Author: Shelly Bensal, Axel Magnuson, Aparna Balagopalan, Daniel M. Bikel
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.10949v2 Announce Type: replace Abstract: Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by amplifying sycophancy, wherein models prioritize agreement with users over accuracy. We conduct the first s...

📖 Read original article


398. Relational Structural Causal Models ​

Author: Adiba Ejaz, Elias Bareinboim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SI, stat.ML

arXiv:2606.14892v2 Announce Type: replace Abstract: An artificial intelligence must have a model of its environment that is causal, supporting reasoning about interventions and counterfactuals, and also combinatorial, supporting generalization to unseen combinations of objects. In this work, we form...

📖 Read original article


399. Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems ​

Author: Xi Chu, Yupeng Hou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY

arXiv:2606.17443v2 Announce Type: replace Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where co...

📖 Read original article


400. FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness ​

Author: Pianran Guo, Pengcheng Zhou, Yucheng Jian, Shuhua Chen, Zhongliang Yang, Linna Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.17642v3 Announce Type: replace Abstract: Financial multimodal reasoning requires agents to coordinate numerical computation, retrieval, visual interpretation, and temporal grounding across heterogeneous evidence sources. Existing tool-augmented agents improve execution fidelity, yet remai...

📖 Read original article


401. DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue ​

Author: Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.17904v2 Announce Type: replace Abstract: Language models increasingly serve as advisory systems in maintenance operations. To prevent hallucination, recent systems ground these models in procedural documentation to constrain them to approved steps. In practice, however, operator queries f...

📖 Read original article


402. MoCo-AIS: A Contrastive Learning Framework for Similarity Computation of Vessel Trajectories ​

Author: Ruixin Song, Md Mahbub Alam, Zahra Sadeghi, Amilcar Soares, Jos'e F. Rodrigues-Jr, Gabriel Spadon
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.17978v2 Announce Type: replace Abstract: Trajectory similarity is a fundamental task in analyzing mobility patterns, essential for applications such as route pattern extraction, mobility prediction, and anomaly detection. Traditional distance-based measures for computing similarity incur ...

📖 Read original article


403. Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving ​

Author: Shihao Ji, HongXi Li, Zihui Song, Mingyu Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.20274v2 Announce Type: replace Abstract: Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy bet...

📖 Read original article


404. Scaling Laws for Task-Specific LLM Distillation ​

Author: Lavinia Ghita, Dhruv Desai, Ioana Boier
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CE

arXiv:2606.24747v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for do...

📖 Read original article


405. The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing ​

Author: Eileanor LaRocco, Sarah Tan, Adarsh Subbaswamy, Anne Andrews, Andrew Taylor, Cree Gaskin, Chirag Agarwal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2606.25108v2 Announce Type: replace Abstract: Autonomous AI systems are transitioning from advisory roles to autonomous ones for medication prescriptions. Recent U.S. bill H.R. 238 and Utah's prescription-renewal pilot program both authorize AI to prescribe medications in an agentic capacity. ...

📖 Read original article


406. PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs ​

Author: Ying Liu, Yi Ye, Quanyu Feng, Mingxi Ye, Mingtao Zhang, Haoyang Li, Chen Jason Zhang, Qing Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.08269v2 Announce Type: replace Abstract: Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in HTML. We present PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph th...

📖 Read original article


407. Evidence-Aware MapReduce for Forkable Compute ​

Author: Yossi Eliaz
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, math.PR, math.ST, stat.TH

arXiv:2607.09689v4 Announce Type: replace Abstract: Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository, tests, observations, or execution ancestor, so counting outputs can amplify one repeated error into high-con...

📖 Read original article


408. OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets ​

Author: Haolin Xue
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13037v2 Announce Type: replace Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or datase...

📖 Read original article


409. CacheSpec: Finding the Sweet Spot for Small Models in Large Language Models ​

Author: Jingquan Chen, Jie Feng, Jinghua Piao, Shaogang Hu, Yong Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.20507v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these settings often incur substantial inference cost. Many such requests share similar computational struct...

📖 Read original article


410. Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain ​

Author: Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, Bernard Ghanem, David R. Pugh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.20510v2 Announce Type: replace Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each...

📖 Read original article


411. CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning ​

Author: Yubo Wang, Qiuyu Zhao, Zenghui Sun, Shichao Dong, Jinsong Lan, Xiaoyong Zhu, Haoyang Li, Bo Zheng, Lei Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.20553v2 Announce Type: replace Abstract: Memory Manager models are pivotal in agent systems. Existing reinforcement-learning methods commonly use LLM-judged synthetic question-answer (QA) pairs: this provides useful downstream task grounding, but values memory through a sampled query dist...

📖 Read original article


412. RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation ​

Author: Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang, Chi Chen, Ling Yao, Maosong Sun
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.24772v2 Announce Type: replace Abstract: Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-...

📖 Read original article


413. Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration ​

Author: Xiangyu Yin, Jiaxu Liu, Zhen Chen, Chih-Hong Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.27836v2 Announce Type: replace Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for every method we evaluate, and we trace this fragility to optimization...

📖 Read original article


414. A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models ​

Author: Xiangyu Yin, Tora Bodin, Rohan Menon, Chih-Hong Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.27910v2 Announce Type: replace Abstract: Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence candidates across 15 model and layer cells from four architectural fa...

📖 Read original article


415. Nova: An End-to-End MLIR Compiler for Deep Learning ​

Author: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, cs.LG, cs.PL

arXiv:2608.00029v2 Announce Type: replace Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions, their execution mode...

📖 Read original article


416. Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language ​

Author: Shengxin Kong, Liwen Xu, Jingwen Fu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.04384v2 Announce Type: replace Abstract: Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid solvers form an extremely sparse subset: most candidate programs are syntactically incorrect, semantically inc...

📖 Read original article


417. NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation ​

Author: Yuchen Zhou, Niels Bobet, Maribel Acosta
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB

arXiv:2608.07530v2 Announce Type: replace Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would l...

📖 Read original article


418. Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding ​

Author: Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08512v2 Announce Type: replace Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates. In contrast to encyclopedic knowledge, where an old fa...

📖 Read original article


419. The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing ​

Author: Yifeng He, Jicheng Wang, Yinzhe Zhao, Chengyang Shi, Jiachen Liu, Hao Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09855v2 Announce Type: replace Abstract: Agentic auto-research is emerging, but most systems treat scientific discovery as goal-oriented optimization against a final benchmark. This paradigm rewards a sparse final verdict and ignores the exploration that precedes it. When agents optimize ...

📖 Read original article


420. Hierarchical Compositionality for An Assistive AI Agent ​

Author: Tianyi Fu, Mohan Sridharan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.10330v2 Announce Type: replace Abstract: AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predi...

📖 Read original article


421. Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects ​

Author: Xin Xu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.10420v2 Announce Type: replace Abstract: Reasoning shortcuts are rule solutions that reach correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings, asking when rules pin concepts d...

📖 Read original article


422. InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk ​

Author: Yuan Gao (Wanxiang), Zeren Yang (Wanxiang), Junnan Li (Wanxiang), Shawn (Wanxiang), Zhong, Ahmed Dajani, Mai Zheng, Andrea Arpaci-Dusseau, Remzi Arpaci-Dusseau
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.OS

arXiv:2608.11234v2 Announce Type: replace Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how we...

📖 Read original article


423. Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence ​

Author: Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.12645v2 Announce Type: replace Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challeng...

📖 Read original article


424. ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond ​

Author: Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang, Shuai Mao, Yating Ling, Linyi Zou, Tailin Zhou, Yun Hin Chan, Wenkai Zhang, Zhanhong Zhou, Guowei Huang, Hongliang Li, Wenjing Cun, Zhitang Chen, Mingxuan Yuan, Yanhui Geng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.14354v2 Announce Type: replace Abstract: Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, explora...

📖 Read original article


425. Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot ​

Author: Ummara Mumtaz, Aimen Noor, Awais Ahmed
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, q-bio.QM

arXiv:2608.15382v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a ...

📖 Read original article


426. Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL ​

Author: Yiyun Su, Zujun Peng, Yu Tian, Yuting Liu, Changruo Zhao, Huiying Zhu, Luyan Zhang, Heming Zeng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15389v2 Announce Type: replace Abstract: LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the metrics authors themselves rep...

📖 Read original article


427. Dynamic Multi-Byte Prediction With Hierarchical Language Models ​

Author: Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15454v2 Announce Type: replace Abstract: Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address thi...

📖 Read original article


428. When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction ​

Author: Feiyang Ren, Shengtao Wen, Lingbing Guo, Yu Tian, Yuanning Cui, Xiang Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.15592v2 Announce Type: replace Abstract: Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this r...

📖 Read original article


429. HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents ​

Author: Shen Liu, Zhenguo Xu, Shaopu Wang, Yike Gao, Chunlei Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.16447v2 Announce Type: replace Abstract: Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management methods such as...

📖 Read original article


430. SkillEffect: Checked Lowering for Memory-Bounded Agent Tools ​

Author: Yinuo Wang, Yiyu Shi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.17007v2 Announce Type: replace Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct progra...

📖 Read original article


431. AutoResearch: Insight In, Hallucination Out ​

Author: Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.17906v2 Announce Type: replace Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects ...

📖 Read original article


432. Towards Zero-Shot Task Transfer with Neurosymbolic World Models ​

Author: Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.17959v2 Announce Type: replace Abstract: State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models ar...

📖 Read original article


433. What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies ​

Author: Narcis Marincat
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2608.20054v2 Announce Type: replace Abstract: Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-...

📖 Read original article


434. Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation ​

Author: Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.20316v2 Announce Type: replace Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires ...

📖 Read original article


435. ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation ​

Author: Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.20735v2 Announce Type: replace Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a vi...

📖 Read original article


436. Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization ​

Author: Praphul Singh, Shanu Kumar, Akshat Agarwal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.20768v2 Announce Type: replace Abstract: Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexami...

📖 Read original article


437. Why we need an AI-resilient society- Profiling Large Language Models ​

Author: Thomas Bartz-Beielstein, Eva Bartz
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:1912.08786v4 Announce Type: replace-cross Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn...

📖 Read original article


438. What is an intelligent system? ​

Author: Martin Molina
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2009.09083v4 Announce Type: replace-cross Abstract: The term intelligent system has emerged in the field of information technology as a category of computer systems derived from successful applications of artificial intelligence. This paper proposes a general description that identifies the ma...

📖 Read original article


439. Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation ​

Author: Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2011.03783v3 Announce Type: replace-cross Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with annotations of multi-word expressions (MWEs), named AlphaMWE. The MWEs include verbal MWEs (vMWEs) defi...

📖 Read original article


440. Evaluating the Efficacy of LLMs to Emulate Realistic Human Personalities ​

Author: Lawrence J. Klinkert, Stephanie Buongiorno, Corey Clark
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2402.14879v2 Announce Type: replace-cross Abstract: To enhance immersion and engagement in video games, the design of Affective Non-Player Characters (ANPCs) is a key focus for researchers and practitioners. Affective Computing frameworks improve Non-player characters (NPC) by providing person...

📖 Read original article


441. Image-Conditional Diffusion Transformer for Underwater Image Enhancement ​

Author: Xingyang Nie, Caoliang Zhang, Xiaoyu Zhai, Fengzhong Qu, Biao Wang, Huilin Ge
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2407.05389v2 Announce Type: replace-cross Abstract: Underwater image enhancement (UIE) has attracted much attention owing to its importance for underwater operation and marine engineering. Motivated by the recent advance in generative models, we propose a novel UIE method based on image-condit...

📖 Read original article


442. LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding ​

Author: Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2409.14644v5 Announce Type: replace-cross Abstract: The advent of large language models (LLMs) has significantly advanced artificial intelligence in software engineering, with source code embeddings playing a crucial role in tasks such as source code clone detection and source code clustering....

📖 Read original article


443. Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations ​

Author: Yonatan Sverdlov, Ido Springer, Nadav Dym
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.06665v5 Announce Type: replace-cross Abstract: This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike traditional approaches, which address these problems using parameter-sharing, we consider an alternative meth...

📖 Read original article


444. NeST: Neighborhood-aware semantic alignment and temporal modulation for LLM based time series forecasting ​

Author: Jayanie Bogahawatte, Sachith Seneviratne, Maneesha Perera, Saman Halgamuge
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2412.04806v2 Announce Type: replace-cross Abstract: Adapting Large Language Models (LLMs) trained on discrete text data, to forecast continuous time series signals is challenging. While finetuning the LLMs enables such adaptation, effectively integrating both textual and time series informatio...

📖 Read original article


445. SRMT: Shared Memory for Multi-agent Lifelong Pathfinding ​

Author: Alsu Sagirova, Yuri Kuratov, Mikhail Burtsev
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2501.13200v2 Announce Type: replace-cross Abstract: Coordination in decentralized multi-agent reinforcement learning (MARL) necessitates that agents share information about their behavior and intentions. Existing approaches rely on communication protocols with domain or resource constraints or...

📖 Read original article


446. Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models ​

Author: Akash Bonagiri, Lucen Li, Rajvardhan Oak, Zeerak Babar, Magdalena Wojcieszak, Anshuman Chhabra
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.SI

arXiv:2501.13976v2 Announce Type: replace-cross Abstract: The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classif...

📖 Read original article


447. Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching ​

Author: Yuling Jiao, Wensen Ma, Defeng Sun, Hansheng Wang, Yang Wang
Published: 8/25/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, stat.ME

arXiv:2502.14424v4 Announce Type: replace-cross Abstract: Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate representation learning as Distribution Matching (DM), learning an augmentation-invariant encoder whose induced...

📖 Read original article


448. SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging ​

Author: Danielle L. Ferreira, Ahana Gangopadhyay, Hsi-Ming Chang, Ravi Soni, Gopal Avinash
Published: 8/25/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2503.05916v2 Announce Type: replace-cross Abstract: Accurate segmentation of anatomical structures in ultrasound (US) images, particularly small ones, is challenging due to noise and variability in imaging conditions (e.g., probe position, patient anatomy, tissue characteristics and pathology)...

📖 Read original article


449. Deep Contrastive Unlearning for Language Models ​

Author: Estrid He, Tabinda Sarwar, Ibrahim Khalil, Xun Yi, Ke Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2503.14900v2 Announce Type: replace-cross Abstract: The past a few years have witnessed the great success of large language models, demonstrating powerful capabilities in comprehending textual data and generating human-like languages. Large language models achieve success by being trained on v...

📖 Read original article


450. Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling ​

Author: Hengran Zhang, Keping Bi, Jiafeng Guo, Xiaojie Sun, Shihao Liu, Daiting Shi, Dawei Yin, Xueqi Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2504.05216v4 Announce Type: replace-cross Abstract: Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for downstream tasks such as re-ranking and augmenting generation. Recently, large language models (LLMs) have demonstrated impressive semantic understandin...

📖 Read original article


451. AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study ​

Author: Mostafa Faghih Shojaei, Rahul Gulati, Benjamin A. Jasperson, Shangshang Wang, Simone Cimolato, Manas Vardhan, Dangli Cao, Willie Neiswanger, Krishna Garikipati
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG

arXiv:2504.08846v2 Announce Type: replace-cross Abstract: We introduce AI University (AI-U), a flexible framework for AI-driven course content delivery that adapts to a course's instructional style. AI-U combines a fine-tuned large language model (LLM) with retrieval-augmented generation (RAG) and a...

📖 Read original article


452. ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model ​

Author: Wuyang Lan, Wenzheng Wang, Changwei Ji, Guoxing Yang, Yongbo Zhang, Xiaohong Liu, Luonan Chen, Shengge Li, Song Wu, Guangyu Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2504.09421v3 Announce Type: replace-cross Abstract: Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here, we introduce Clinica...

📖 Read original article


453. COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification ​

Author: Mariano Rivera, Angello Hoyos
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2505.18315v4 Announce Type: replace-cross Abstract: We introduce CoLoRA (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs). CoLoRA extends LoRA to convolutional layers by decomposing kernel updates into lightweight depthwise a...

📖 Read original article


454. Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric ​

Author: Jatin Kumar Arora, Soutrik Bandyopadhyay, Sunil Sulania, Shubhendu Bhasin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2505.23197v4 Announce Type: replace-cross Abstract: Path planning for autonomous robots faces a fundamental trade-off between path length and obstacle clearance. While existing algorithms typically prioritize a single objective, we introduce the Unified Path Planner (UPP), a graph-search algor...

📖 Read original article


455. Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games ​

Author: Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2505.24255v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, crucial for complex social interactions. We investigate ToM reasoning's role in aligning agentic behaviors with hum...

📖 Read original article


456. Time Series Forecasting via Reasoning: A Slow-Thinking Approach with Reinforcement Fine-Tuned LLMs ​

Author: Yitong Zhou, Yucong Luo, Mingyue Cheng, Qi Liu, Jiahao Wang, Daoyu Wang, Enhong Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.10630v4 Announce Type: replace-cross Abstract: To advance time series forecasting (TSF), various methods have been proposed to improve prediction accuracy, evolving from statistical techniques to data-driven deep learning architectures. Despite their effectiveness, most existing methods s...

📖 Read original article


457. Seismic Acoustic Impedance Inversion Framework Based on Conditional Latent Generative Diffusion Model ​

Author: Jie Chen, Hongling Chen, Jinghuai Gao, Chuangji Meng, Tao Yang, XinXin Liang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.13529v2 Announce Type: replace-cross Abstract: Seismic acoustic impedance plays a crucial role in lithological identification and subsurface structure interpretation. However, due to the inherently ill-posed nature of the inversion problem, directly estimating impedance from post-stack se...

📖 Read original article


458. From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment ​

Author: Hexiang Gu, Qifan Yu, Yuan Liu, Zikang Li, Saihui Hou, Jian Zhao, Zhaofeng He
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2506.18919v5 Announce Type: replace-cross Abstract: As a multimodal communication medium that integrates images and text, memes often convey implicit harmful content through metaphors, satire, and humor, making harmful meme detection a complex and challenging task. Although recent studies have...

📖 Read original article


459. A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs ​

Author: Kethmi Hirushini Hettige, Jiahao Ji, Cheng Long, Shili Xiang, Gao Cong, Jingyuan Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2506.20073v2 Announce Type: replace-cross Abstract: Spatio-temporal data mining plays a pivotal role in informed decision making across diverse domains. However, existing models are often restricted to narrow tasks, lacking the capacity for multi-task inference and complex long-form reasoning ...

📖 Read original article


460. Mission-Aligned Learning-Informed Control of Autonomous Systems: Formulation and Foundations ​

Author: Vyacheslav Kungurtsev, Alessandro Di Frenna, Gustav Sir, Monicah Cherop Naibei, Haozhe Tian, Homayoun Hamedmoghadam, Akhil Anand, Sebastien Gros
Published: 8/25/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.RO

arXiv:2507.04356v3 Announce Type: replace-cross Abstract: Research, innovation and practical capital investment have been increasing rapidly toward the realization of autonomous physical agents. This includes industrial and service robots, unmanned aerial vehicles, embedded control devices, and a nu...

📖 Read original article


461. MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora ​

Author: Tuan-Luc Huynh, Thuy-Trang Vu, Weiqing Wang, Trung Le, Dragan Ga\v{s}evi'c, Yuan-Fang Li, Thanh-Toan Do
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.LG

arXiv:2507.09924v2 Announce Type: replace-cross Abstract: Continually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationally expensive and impractical under resource constraints. We propose MixLoRA-DSI, a novel framework tha...

📖 Read original article


462. Text-ADBench: Text Anomaly Detection Benchmark Based on LLM Embeddings ​

Author: Feng Xiao, Jicong Fan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2507.12295v2 Announce Type: replace-cross Abstract: Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation identification, spam detection and content moderation, etc. Despite significant advances in large langu...

📖 Read original article


463. RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning ​

Author: Situo Zhang, Hanqi Li, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Danyu Luo, Bo Chen, Xin Chen, Kai Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CE, cs.AI, physics.chem-ph

arXiv:2507.17448v2 Announce Type: replace-cross Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existing AI methods often rely on pattern matching rather than transferable chemical reasoning, limiting both generalizability and interpretability. Here we...

📖 Read original article


464. TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios ​

Author: Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD, eess.AS

arXiv:2507.18061v4 Announce Type: replace-cross Abstract: Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interac...

📖 Read original article


465. Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads ​

Author: Jiayin Jin, Erika Sun, Zhimeng Pan, Yang Tang, Jiarui Feng, Kungang Li, Chongyuan Xiang, Jiacheng Li, Runze Su, Siping Ji, Han Sun, Ling Leng, Prathibha Deshikachar
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE

arXiv:2508.02609v3 Announce Type: replace-cross Abstract: Graph Neural Networks (GNN) have been extensively applied to industry recommendation systems, as seen in models like GraphSage\cite{GraphSage}, TwHIM\cite{TwHIM}, LiGNN\cite{LiGNN} etc. In these works, graphs were constructed based on users' ...

📖 Read original article


466. From Isolation to Alignment: Unified LoRA for Efficient Multi-Task Learning ​

Author: Jinda Liu, Yi Chang, Yuan Wu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.05078v2 Announce Type: replace-cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is essential for adapting Large Language Models (LLMs) to multi-task scenarios. A prevailing trend in this field involves complex LoRA variants with multiple adapters or heads, which rely on the premise ...

📖 Read original article


467. An Information-Flow Perspective on Explainability Requirements: Specification and Verification ​

Author: Bernd Finkbeiner, Hadar Frenkel, Julian Siber
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2509.01479v3 Announce Type: replace-cross Abstract: Explainable systems expose information about why certain observed effects are happening to the agents interacting with them. We argue that this constitutes a positive flow of information that needs to be specified, verified, and balanced agai...

📖 Read original article


468. ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting ​

Author: Jiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long, Ge Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.22225v3 Announce Type: replace-cross Abstract: Lifting 2D open-vocabulary understanding into 3D Gaussian Splatting (3DGS) scenes is a critical challenge. Mainstream methods, built on an embedding paradigm, suffer from three key flaws: (i) geometry-semantic inconsistency, where points, rat...

📖 Read original article


469. HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models ​

Author: Zhinan Xie, Peisong Wang, Shuang Qiu, Jian Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.23928v3 Announce Type: replace-cross Abstract: Speculative decoding has proven effective for accelerating inference in Large Language Models (LLMs), yet its extension to Vision-Language Models (VLMs) remains limited by the computational burden and semantic inconsistency introduced by visu...

📖 Read original article


470. SLogic: Subgraph-Informed Logical Rule Learning for Knowledge Graph Completion ​

Author: Trung Hoang Le, Tran Cao Son, Ishtiaq Ahmed, Huiping Cao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.00279v3 Announce Type: replace-cross Abstract: Logical rule-based methods offer an interpretable approach to knowledge graph completion (KGC) by capturing compositional relationships in the form of human-readable inference rules. While existing logical rule-based methods learn rule confid...

📖 Read original article


471. RSTGCN: Railway-centric Spatio-Temporal Graph Convolutional Network for Train Delay Prediction ​

Author: Koyena Chowdhury, Paramita Koley, Abhijnan Chakraborty, Saptarshi Ghosh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.01262v2 Announce Type: replace-cross Abstract: Accurate prediction of train delays is critical for efficient railway operations. While earlier approaches have largely focused on forecasting the exact delays of individual trains, studies on station-level delay prediction are somewhat spars...

📖 Read original article


472. GraphMed-LT: Patient-Specific Graph Memory with Latent Clinical Thought Refinement for Multi-Turn Medical Conversations ​

Author: Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.03536v3 Announce Type: replace-cross Abstract: Multi-turn medical question answering (QA) aims to model realistic clinical diagnosis, where a doctor gathers patient information across multiple turns of conversation. Existing multi-turn medical conversation systems have shown promising pro...

📖 Read original article


473. LLM-Specific Utility for Retrieval-Augmented Generation ​

Author: Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2510.11358v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved passages are useful for a large language model (LLM) to generate correct and complete answers. We argue...

📖 Read original article


474. A New Type of Adversarial Examples ​

Author: Xingyang Nie, Caoliang Zhang, Su Pan, Biao Wang, Huilin Ge, Tao Fang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GR

arXiv:2510.19347v2 Announce Type: replace-cross Abstract: Most machine learning models are vulnerable to adversarial examples, which poses security concerns on these models. Adversarial examples are crafted by applying subtle but intentionally worst-case modifications to examples from the dataset, l...

📖 Read original article


475. Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion ​

Author: Zhiwen Yu, Zhaocheng Liu, Xiaoqing Liu, Huanqiang Zeng, C. L. Philip Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD, eess.AS

arXiv:2510.21797v4 Announce Type: replace-cross Abstract: Multimodal learning faces modality imbalance, where dominant modalities suppress weaker ones due to inconsistent convergence rates. Existing static or heuristic methods overlook sample-level variations in prediction bias and fail to isolate l...

📖 Read original article


476. One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing ​

Author: Xu Yang, Chenhui Lin, Haotian Liu, Qi Wang, Yue Yang, Wenchuan Wu
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2511.12484v2 Announce Type: replace-cross Abstract: With the integration of massive distributed energy resources and the widespread participation of novel market entities, the operation of active distribution networks (ADNs) is progressively evolving into a complex, multi-scenario, and multi-o...

📖 Read original article


477. Radial Compensation: The Inverse Base-Distribution Problem for Chart-Based Generative Models on Riemannian Manifolds ​

Author: Marios Papamichalis, Regina Ruane
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.DG, math.IT, stat.ML

arXiv:2511.14056v3 Announce Type: replace-cross Abstract: Latent-variable models on spheres and hyperbolic spaces usually draw a Gaussian in the tangent space at a base point and push it onto the manifold. On these spaces the distance from the base point is the coordinate that carries meaning: depth...

📖 Read original article


478. HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation ​

Author: Nguyen Tri Nghia, Nguyen Van Son, Nguyen Thi Hanh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2511.17537v5 Announce Type: replace-cross Abstract: Wireless Sensor Networks (WSN) are the backbone of essential monitoring applications, but their deployment in unfavourable conditions increases the risk to data integrity and system reliability. Traditional fault detection methods often strug...

📖 Read original article


479. MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design ​

Author: S. Rodionov, A. Burguete-Lopez, M. Makarenko, Q. Wang, F. Getman, A. Fratalocchi
Published: 8/25/2026, 4:00:00 AM
Categories: physics.optics, cs.AI

arXiv:2511.18980v2 Announce Type: replace-cross Abstract: Foundation models (FM) are transforming artificial intelligence by enabling generalizable, data-efficient solutions across different domains for a broad range of applications. However, the lack of large and diverse datasets limits the develop...

📖 Read original article


480. Multi-Context Fusion Transformer for Pedestrian Crossing Intention Prediction in Urban Environments ​

Author: Yuanzhe Li, Hang Zhong, Steffen M"uller
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2511.20011v3 Announce Type: replace-cross Abstract: Pedestrian crossing intention prediction is essential for autonomous vehicles to improve pedestrian safety and reduce traffic accidents. However, accurate pedestrian intention prediction in urban environments remains challenging due to the mu...

📖 Read original article


481. MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving ​

Author: Jia Hu, Zhexi Lian, Xuerun Yan, Ruiang Bi, Dou Shen, Yu Ruan, Chunlong Xia, Haoran Wang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2512.03795v3 Announce Type: replace-cross Abstract: Autonomous Driving (AD) vehicles still struggle to exhibit human-like behavior in highly dynamic and interactive traffic scenarios. The key challenge lies in AD's limited ability to interact with surrounding vehicles, largely due to a lack of...

📖 Read original article


482. Diagnosing Capability Preservation and Task Sensitivity in Memory Augmented Document Classifiers ​

Author: Isaac Kofi Nti
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2512.06582v2 Announce Type: replace-cross Abstract: End task accuracy alone cannot determine whether a memory mechanism preserves an acquired capability, exposes sample-specific stored information, or contributes measurably to downstream performance. This study introduces Protected QL Memory a...

📖 Read original article


483. Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs ​

Author: Pius Horn, Janis Keuper
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR

arXiv:2512.09874v3 Announce Type: replace-cross Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academic literature, yet existing benchmarks either exclude formulas entirely or lack semantically-a...

📖 Read original article


484. An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift ​

Author: Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2601.05882v2 Announce Type: replace-cross Abstract: Preference tuning aligns base language models to human judgments of quality, helpfulness, or safety by optimizing over explicit preference signals rather than likelihood alone. Prior work has shown that preference tuning degrades performance ...

📖 Read original article


485. LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems ​

Author: Jo~ao A. Leite, Olesya Razuvayevskaya, Kalina Bontcheva, Carolina Scarton
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2601.16890v2 Announce Type: replace-cross Abstract: Automated fact-checking (AFC) systems are susceptible to adversarial attacks, enabling false claims to evade detection. Existing adversarial frameworks typically rely on injecting noise or altering semantics, yet no existing framework exploit...

📖 Read original article


486. PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation ​

Author: Qingyu Fan, Zhaoxiang Li, Jinrui Hu, Yi Lu, Wang Chen, Qiu Shen, Xiao-xiao Long, Yinghao Cai, Tao Lu, Shuo Wang, Xun Cao
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2601.17885v2 Announce Type: replace-cross Abstract: Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint changes and scene variations. Existing vision-language-action models often lack such robustness because (i) multi-view features are fus...

📖 Read original article


487. Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair ​

Author: Runxiang Cheng, Michele Tufano, Jos'e Cambronero, Renyao Wei, Sherry Shi, Grant Uy, Pat Rondon, Franjo Ivan\v{c}i'c
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2601.19066v3 Announce Type: replace-cross Abstract: Bug Reproduction Tests (BRTs) have been used in many Automated Program Repair (APR) systems, primarily for validating fixes and aiding fix generation. In practice, when developers submit a patch, they often implement the BRT alongside the fix...

📖 Read original article


488. Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks ​

Author: Luwei Sun, Dongrui Shen, Feng Chuanwen, Jianfe Li, Yulong Zhao, Han Feng
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.21242v2 Announce Type: replace-cross Abstract: Motivated by challenges in conditional generative modeling, where the target conditional density takes the form of a ratio f1 over f2, this paper develops a theoretical framework for approximating such ratio-type functionals. Here, f1 and f2 ...

📖 Read original article


489. Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models ​

Author: Guillermo Gil de Avalle, Laura Maruster, Christos Emmanouilidis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.22754v2 Announce Type: replace-cross Abstract: Industrial troubleshooting guides encode diagnostic procedures in flowchart-like diagrams where spatial layout and technical language jointly convey meaning. To integrate this knowledge into operator support systems, which assist shop-floor p...

📖 Read original article


490. Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency ​

Author: Bingzheng Wang, Xiaoyan Gu, Hongbo Xu, Hongcheng Li, Zimo Yu, Jiang Zhou, Weiping Wang, Wu Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2602.01765v2 Announce Type: replace-cross Abstract: Diffusion models have been widely deployed in AIGC services, but their reliance on opaque training data exposes them to backdoor attacks. In practical auditing scenarios, auditors are typically unable to access model parameters due to intelle...

📖 Read original article


491. ReasonEdit: Editing Vision-Language Models using Human Reasoning ​

Author: Jiaxing Qiu, Kaihua Hou, Roxana Daneshjou, Ahmed Alaa, Thomas Hartvigsen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.02408v5 Announce Type: replace-cross Abstract: Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically requir...

📖 Read original article


492. First-Principles AI finds crystallization of fractional quantum Hall liquids ​

Author: Ahmed Abouelkomsan, Liang Fu
Published: 8/25/2026, 4:00:00 AM
Categories: cond-mat.mes-hall, cond-mat.str-el, cs.AI

arXiv:2602.03927v2 Announce Type: replace-cross Abstract: When does a fractional quantum Hall (FQH) liquid crystallize? Addressing this question requires a framework that treats fractionalization and crystallization on equal footing, especially in strong Landau-level mixing regime. Here, we introduc...

📖 Read original article


493. Mode-Dependent Rectification for Stable PPO Training ​

Author: Mohamad Mohamad, Francesco Ponzio, Xavier Descombes
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.05619v2 Announce Type: replace-cross Abstract: Mode-dependent architectural components (layers that behave differently during training and evaluation, such as Batch Normalization or dropout) are commonly used in visual reinforcement learning but can destabilize on-policy optimization. We ...

📖 Read original article


494. iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems ​

Author: Yi-Xiang Hu, Yuke Wang, Feng Wu, Zirui Huang, Shuli Zeng, Xiang-Yang Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2602.06064v2 Announce Type: replace-cross Abstract: Scheduling precedence-constrained tasks under shared renewable resources is critical to modern computing platforms. It is often modeled as the Resource Investment Problem (RIP) by minimizing the cost of provisioned renewable resources under p...

📖 Read original article


495. Which Algorithms Can Graph Neural Networks Learn? ​

Author: Solveig Wittig, Antonis Vasileiou, Robert R. Nerem, Timo Stoll, Floris Geerts, Yusu Wang, Christopher Morris
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, cs.NE

arXiv:2602.13106v2 Announce Type: replace-cross Abstract: In recent years, there has been growing interest in understanding neural architectures' ability to learn to execute discrete algorithms, a line of work often referred to as neural algorithmic reasoning. The goal is to integrate algorithmic re...

📖 Read original article


496. ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies ​

Author: Xingjian Wu, Xvyuan Liu, Junkai Lu, Siyuan Wang, Xiangfei Qiu, Yang Shu, Jilin Hu, Chenjuan Guo, Bin Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2602.14681v4 Announce Type: replace-cross Abstract: LLM-powered Multi-Agent Systems (MAS) have emerged as an effective approach towards collaborative intelligence, and have attracted wide research interests. Among them, ``self-evolving'' MAS, treated as a more flexible and powerful technical r...

📖 Read original article


497. Continual Uncertainty Learning for Robust Control of Nonlinear Systems with Multiple Heterogeneous Uncertainties ​

Author: Heisei Yonezawa, Ansei Yonezawa, Itsuro Kajiwara
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY

arXiv:2602.17174v3 Announce Type: replace-cross Abstract: Robust control of mechanical systems with multiple uncertainties remains a fundamental challenge, particularly when nonlinear dynamics and operating-condition variations are intricately intertwined. Although deep reinforcement learning combin...

📖 Read original article


498. Learning with Boolean threshold functions ​

Author: Veit Elser, Manish Krishan Lal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17493v2 Announce Type: replace-cross Abstract: We develop a method for training neural networks on Boolean data in which the values at all nodes are strictly $\pm 1$, and the resulting models are typically equivalent to networks whose nonzero weights are also $\pm 1$. The method replaces ...

📖 Read original article


499. Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries ​

Author: Muhammad Aziz Ullah, Abdul Serwadda
Published: 8/25/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.SE

arXiv:2602.18492v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are now good enough at coding that developers can describe intent in plain language and let the tool produce the first code draft, a workflow increasingly built into tools like GitHub Copilot, Cursor, and Replit. ...

📖 Read original article


500. Semantic Substrate Dynamics Theory: An Operator-Theoretic Framework for Geometric Semantic Drift ​

Author: Stephen Russell
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2602.18699v2 Announce Type: replace-cross Abstract: Studies of semantic drift report heterogeneous signals, including embedding displacement, neighbor change, distributional divergence, and recursive trajectory instability, without a shared account that relates them. Semantic Substrate Dynamic...

📖 Read original article


501. VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting ​

Author: Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.14659v2 Announce Type: replace-cross Abstract: Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still struggles to achieve reliable spatio-temporal grounding during...

📖 Read original article


502. When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making ​

Author: Jun Liu, Pu Zhao, Zhenglun Kong, Xuan Shen, Peiyan Dong, Fan Yang, Lin Cui, Hao Tang, Geng Yuan, Wei Niu, Wenbin Zhang, Xue Lin, Gaowen Liu, Yanzhi Wang, Dong Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2603.16673v5 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the environment. However, invoking LLM reasoning introduces substant...

📖 Read original article


503. Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing ​

Author: Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2603.28900v2 Announce Type: replace-cross Abstract: We address robust separation assurance for small Unmanned Aircraft Systems (sUAS) under GPS degradation and spoofing via Multi-Agent Reinforcement Learning (MARL). In cooperative surveillance, each aircraft (or agent) broadcasts its GPS-deriv...

📖 Read original article


504. SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning ​

Author: Daeyong Kwon, Soyoung Yoon, Seung-won Hwang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.01993v3 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid intermediate reasoning. We propose SAFE, an LLM-as-verifier framework for evidence-grounded multi-h...

📖 Read original article


505. Verbalizing LLMs' assumptions to explain and control sycophancy ​

Author: Myra Cheng, Isabel Sieh, Humishka Zope, Sunny Yu, Lujain Ibrahim, Aryaman Arora, Jared Moore, Desmond Ong, Dan Jurafsky, Diyi Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2604.03058v3 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment. We hypothesize that this behavior arises from LLMs' incorrect assumptions about the user, like under...

📖 Read original article


506. What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know" ​

Author: Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.05779v2 Announce Type: replace-cross Abstract: While large language models (LLMs) demonstrate strong capabilities across diverse user queries, they still suffer from hallucinations, often arising from knowledge misalignment between pre-training and fine-tuning. To address this misalignmen...

📖 Read original article


507. FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts ​

Author: Guillermo Gil de Avalle, Laura Maruster, Eric Sloot, Christos Emmanouilidis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.06770v2 Announce Type: replace-cross Abstract: Maintenance procedures in manufacturing facilities are often documented as flowcharts in static PDFs or scanned images. They encode procedural knowledge essential for asset lifecycle management, yet inaccessible to modern operator support sys...

📖 Read original article


508. Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework ​

Author: Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evidence beyond the visible range remains largely unexplored. Hyperspectral imagery (HSI) provides dense s...

📖 Read original article


509. Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types ​

Author: Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.09544v3 Announce Type: replace-cross Abstract: Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate how this capability is organized within model parameters. We ...

📖 Read original article


510. CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation ​

Author: Aarush Sinha, Arion Das, Soumyadeep Nag, Charan Karnati, Shravani Nag, Chandra Vadhan Raj, Aman Chadha, Vinija Jain, Suranjana Trivedy, Amitava Das
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL

arXiv:2604.09746v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent environments has become an important alignment challenge. We take a neutral empirical stance and const...

📖 Read original article


511. Alignment midtraining for animals ​

Author: Jasmine Brazilek, Miles Tidmarsh
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.13076v4 Announce Type: replace-cross Abstract: We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reas...

📖 Read original article


512. DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories ​

Author: Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.20443v3 Announce Type: replace-cross Abstract: We introduce DialToM, an annotated Theory of Mind (ToM) benchmark built from naturalistic human-human dialogues using a multiple-choice evaluation framework. Concurrent with recent work showing a gap between explicit mental-state inference an...

📖 Read original article


513. ONOTE: Hypergraph-Grounded Omnimodal Reasoning for Computational Music Science ​

Author: Menghe Ma, Siqing Wei, Yuecheng Xing, Ziyue Zhu, Zhenghong Lin, Yaheng Wang, Fanhong Meng, Peijun Han, Luu Anh Tuan, Haoran Luo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM, eess.AS

arXiv:2604.20719v2 Announce Type: replace-cross Abstract: Omnimodal notation processing, centered on sheet music, is a controlled scientific setting in which auditory, visual, symbolic, and physical representations must encode the same musical events. Yet existing work remains fragmented across reco...

📖 Read original article


514. DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models ​

Author: Dake Bu, Wei Huang, Andi Han, Si Wu, Hau-San Wong, Qingfu Zhang, Taiji Suzuki, Atsushi Nitanda
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.24357v3 Announce Type: replace-cross Abstract: Discrete diffusion models admit many token orders, yet most systems rely on confidence-based decoding. Confidence is a strong and efficient heuristic, but it can be myopic because local certainty does not measure a position's effect on termin...

📖 Read original article


515. DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning ​

Author: Haoyu Huang, Jiaxin Bai, Shujie Liu, Yang Wei, Huihao Jing, Hong Ting Tsang, Yisen Gao, Zhongwei Xie, Yufei Li, Yangqiu Song
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.10488v2 Announce Type: replace-cross Abstract: External knowledge enables large language model (LLM) agents to ground their actions and decisions beyond intrinsic parametric memory in open-ended, knowledge-intensive downstream tasks. Yet the quality of the underlying knowledge bases is sy...

📖 Read original article


516. Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective ​

Author: Feng Zhang, Xinhong Ma, Ziqiang Dong, Xi Leng, Jianfei Zhao, Xin Sun, Yang Yang, Guanjun Jiang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.12969v4 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admits an equivalent discriminative reformulation, in which policy...

📖 Read original article


517. Edge-AI-Driven Learning-to-Rank for Decentralized Task Allocation in Circular Smart Manufacturing ​

Author: Mohammadhossein Ghahramani, Yan Qiao, Mengchu Zhou
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.16433v2 Announce Type: replace-cross Abstract: Task allocation in smart manufacturing systems must operate under decentralized decision-making, dynamic workloads, and shared-resource constraints. In circular manufacturing settings, these challenges are further intensified because tasks co...

📖 Read original article


518. How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws ​

Author: Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.25698v2 Announce Type: replace-cross Abstract: High-quality data is scarce in large language model (LLM) training, yet how to schedule its use with optimization dynamics lacks theoretical guidance. We extend functional scaling laws with time-varying data quality and derive asymptotically ...

📖 Read original article


519. BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks ​

Author: Tirtharaj Dash
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE, q-bio.QM

arXiv:2605.28739v2 Announce Type: replace-cross Abstract: Tabular data in knowledge-rich domains often carries a latent prior in the form of Boolean implication relationships (BIRs) between pairs of features. We mine such relationships with a sparse-exception binomial test. We encode the resulting t...

📖 Read original article


520. DRIFT: Joint Channel Estimation and Prediction Towards Pilotless 6G Non-Terrestrial Networks ​

Author: Bruno De Filippo, Carla Amatetti, Alessandro Vanelli-Coralli
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2605.31065v2 Announce Type: replace-cross Abstract: Non-terrestrial networks (NTNs) are expected to play a pivotal role in sixth-generation (6G) systems by enabling ubiquitous connectivity and massive communication. In this context, channel prediction emerges as a key technique to improve the ...

📖 Read original article


521. Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context ​

Author: Shihao Ji, Mingyu Li, Zihui Song
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.01101v2 Announce Type: replace-cross Abstract: The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cognitive Engine (NBCE) parallelizes long-context inference by chunking documents and routing to ...

📖 Read original article


522. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments ​

Author: Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2606.03770v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost efficiency, low latency, and optimal resource u...

📖 Read original article


523. Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning ​

Author: Yuhuan Yuan, Zhouliang Yu, Minghao Liu, Weiyang Liu, Ge Lin Kan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.07602v2 Announce Type: replace-cross Abstract: LLM-based LEGO assembly requires both semantic grounding and physical feasibility. In this paper, we identify a data-induced failure mode, physhack, in which generated assemblies satisfy physical-validity constraints while remaining geometric...

📖 Read original article


524. SocraticPO: Policy Optimization via Interactive Guidance ​

Author: Zirui Liu, Tingyue Pan, Jie Ouyang, Qi Liu, Xianquan Wang, Jiayu Liu, Qingchuan Li, Jing Sha, Zhenya Huang, Shijin Wang, Enhong Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2606.09887v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization direction but rarely explain how a model should revise its mistak...

📖 Read original article


525. Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation ​

Author: Yuchen Ling, Shengcheng Yu, Zhenyu Chen, Chunrong Fang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2606.10749v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, and act on external environments. This transition changes the nature of security risk. In age...

📖 Read original article


526. TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning ​

Author: Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2606.11119v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. However, rollout-intensive policy optimization is often limited by insufficient reward contra...

📖 Read original article


527. Chain of Operators: An Inference-Time Harness for In-Context Operator Learning ​

Author: Minghui Yang, Chenghan Wu, Ling Guo, Liu Yang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.12318v2 Announce Type: replace-cross Abstract: While scientific foundation models show immense promise in accelerating physical simulations and numerical forecasting, they remain notoriously brittle when encountering out-of-distribution (OOD) scenarios. Adapting these generalist models to...

📖 Read original article


528. One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders ​

Author: Minghao Luo, Liang Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.13610v2 Announce Type: replace-cross Abstract: Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: LLM recommenders may consume web content that Generative Engine Optimization (GEO) operators have polluted t...

📖 Read original article


529. Deep Learning-Driven Inverse Design of Doherty Power Amplifiers Using Pixelated Combiners and Dual-State Impedance Synthesis ​

Author: Han Zhou, Haojie Chang, David Widen, Christian Fager
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.AR, cs.SY, eess.SY

arXiv:2606.18395v2 Announce Type: replace-cross Abstract: The output combiner of a Doherty power amplifier (PA) integrates load modulation, impedance matching, and phase compensation within a single network, making its design and synthesis highly challenging. In this paper, we propose a three-port D...

📖 Read original article


530. Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements ​

Author: Han Zhou, Richard Bannister, Caspar Pierce, Haojie Chang, David Widen, Ludvig Fornstedt, Gabriel Melin, Alexander Bohlin, Pontus Lindeberg Fredriksson, Dilbagh Singh, Christian Fager, Koen Buisman
Published: 8/25/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.AR, cs.SY, eess.SY

arXiv:2606.18402v2 Announce Type: replace-cross Abstract: Traditional microwave filter design typically relies on iterative parameter tuning and predefined topologies, which limits design space and increases development time. This study uses a deep learning approach combining convolutional neural ne...

📖 Read original article


531. RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation ​

Author: Pengzhi Yang, Xinyu Wang, Pengyu Jing, Kehan Wen, Yiduo Qu, Zhenhao Huang, Minghao Fu, Xin Liu, Yaheng Shen, Fan Shi
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2606.22027v3 Announce Type: replace-cross Abstract: Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poo...

📖 Read original article


532. Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Structure in Passive Object-State World Models ​

Author: Yang Liu, Yuming Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2606.28455v3 Announce Type: replace-cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. We introduce a controlled diagnostic protocol for studying event-cond...

📖 Read original article


533. What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs ​

Author: Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, stat.ML

arXiv:2606.28615v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs. Yet it remains unclear whether these explanations ...

📖 Read original article


534. Compositional Dynamics in Learning and Mechanics ​

Author: David I. Spivak
Published: 8/25/2026, 4:00:00 AM
Categories: math.CT, cs.AI

arXiv:2606.28984v2 Announce Type: replace-cross Abstract: We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The syntax is an operad Arr whose objects are input-output interfaces (pairs of manifolds) and whose morph...

📖 Read original article


535. Group-Equivariant Poincar\'e Convolutional Networks ​

Author: Aiden Durrant, Rahul Baburajan, Georgios Leontidis
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.00556v2 Announce Type: replace-cross Abstract: While recent methods like that of the Poincar'e ResNet have demonstrated the ability to learning visual representations directly in hyperbolic space, their optimisation remains a challenge, primarily due to the parameter redundancy of learni...

📖 Read original article


536. Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies ​

Author: Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.02092v3 Announce Type: replace-cross Abstract: Deploying a pretrained flow-matching vision-language-action (VLA) policy on a particular robot and workspace often calls for task-specific adaptation, while full- policy fine-tuning is costly and changes the base behavior. We present Guided A...

📖 Read original article


537. Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting ​

Author: Disheng Liu, Tuo Liang, Chaoda Song, Yu Yin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.02637v2 Announce Type: replace-cross Abstract: Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or ...

📖 Read original article


538. DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation ​

Author: Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia, Aditya Joshi, Lu Yin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.07669v2 Announce Type: replace-cross Abstract: Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM...

📖 Read original article


539. GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification ​

Author: Yash Shah, Omar Todd, Philipp Seeb"ock, Georg Langs, Ben Glocker, Raghav Mehta
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10357v2 Announce Type: replace-cross Abstract: The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in clinical practice. Recently, a hybrid pipeline (GRC-Net) for CVD classification was proposed, which lev...

📖 Read original article


540. Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift ​

Author: Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.10358v2 Announce Type: replace-cross Abstract: Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and can...

📖 Read original article


541. Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels ​

Author: Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun, Anton Sergeev
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.12792v2 Announce Type: replace-cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to the benchmark's grading procedure and capture only observed behavior on a given s...

📖 Read original article


542. An offline approach to fNIRS-guided reinforcement learning for robot behavior ​

Author: Julia Santaniello, Madelaine Brower, Benson Jiang, Donatello Sassaroli, Chenyuan Zhang, Robert Jacob, Jivko Sinapov
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.14393v2 Announce Type: replace-cross Abstract: Human-in-the-loop Reinforcement Learning has become a popular approach for training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectro...

📖 Read original article


543. Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits ​

Author: Haifeng Li, Mo Hai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.16646v2 Announce Type: replace-cross Abstract: Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, and they fail silently. A generated model often runs and still encodes the wrong problem, while standard evaluation ...

📖 Read original article


544. Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT ​

Author: Tanveer Ahmed, Seyedali Pourmoafil
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY, cs.LG

arXiv:2607.18429v2 Announce Type: replace-cross Abstract: Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. Most reported detection accuracies, however, are measured on clean, in-distribution test data rather tha...

📖 Read original article


545. CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders ​

Author: Soroosh Tayebi Arasteh, Sven Nebelung, Daniel Truhn
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV

arXiv:2607.18451v2 Announce Type: replace-cross Abstract: Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Nearest-neighbor discordance does, but with unequal banks the opposite-label neighbor wins on density, not g...

📖 Read original article


546. PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image ​

Author: Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.19261v3 Announce Type: replace-cross Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped pat...

📖 Read original article


547. CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference ​

Author: Jiyuan Tan, Vasilis Syrgkanis
Published: 8/25/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, econ.EM

arXiv:2607.22511v3 Announce Type: replace-cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such r...

📖 Read original article


548. A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs ​

Author: Shihao Ji, Mingyu Li, Zihui Song
Published: 8/25/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2607.22567v2 Announce Type: replace-cross Abstract: Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order gradient-free methods. We develop a kinetic framework for algorithms that estimate both gradient and H...

📖 Read original article


549. A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks ​

Author: Du Yin, Xiachong Lin, Yue Tan, Jinliang Deng, Estrid He, Hao Xue, Flora D. Salim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.25875v3 Announce Type: replace-cross Abstract: Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies typically assume fixed sensor graphs, overlooking the continuous evolution of real-world traffic networ...

📖 Read original article


550. IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations ​

Author: David Kaleko, Sergey Ivanov, Md Mofijul Islam
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.26075v2 Announce Type: replace-cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domain specialist...

📖 Read original article


551. What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs ​

Author: Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2608.00013v2 Announce Type: replace-cross Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model famil...

📖 Read original article


552. Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization ​

Author: Daojie Peng, Fulong Ma, Bingtao Wang, Sheng Wang, Jun Ma
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2608.00569v2 Announce Type: replace-cross Abstract: Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Exi...

📖 Read original article


553. CallScreenBench: Benchmarking Small Language Models as Phone Secretaries ​

Author: Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.01033v2 Announce Type: replace-cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secre...

📖 Read original article


554. Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal reference ​

Author: Luan Ozelim, Hector Zenil
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.01575v2 Announce Type: replace-cross Abstract: Whether large language models perform algorithmic inference or pattern completion is hard to test, because most benchmarks supply answers but no distributional reference for what the shown evidence licenses. F-ICL supplies one exactly: we exh...

📖 Read original article


555. Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory ​

Author: Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang
Published: 8/25/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.NE

arXiv:2608.01947v2 Announce Type: replace-cross Abstract: The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) lik...

📖 Read original article


556. Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States ​

Author: Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.02657v2 Announce Type: replace-cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this threat, little is known about the internals of agentic LLMs whe...

📖 Read original article


557. AI Alignment and Fiduciary Obligation ​

Author: Benjamin Lange
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.02660v2 Announce Type: replace-cross Abstract: Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collaboration, learning, emotional support, and companionship among others. Current alignment efforts consider ...

📖 Read original article


558. Breadcrumbing Search Agents ​

Author: Xuebin Li, Hanqing Zhao, Siyuan Liang, Kejiang Chen, Weiming Zhang, Dacheng Tao, Nenghai Yu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2608.04565v2 Announce Type: replace-cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to prompt injection a...

📖 Read original article


559. SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery ​

Author: Shrenik Zinage
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.04930v2 Announce Type: replace-cross Abstract: Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graphs (DAGs) that explain the observed data. The resulting posterior allows systematic reasoning regarding ...

📖 Read original article


560. Answer First, Reason Later: When Commitment Order Costs Accuracy in Diffusion Language Models ​

Author: Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Hwiyeong Lee, Taesup Kim
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05687v2 Announce Type: replace-cross Abstract: Masked diffusion language models revise many masked output positions in parallel. We call a token committed once it becomes visible and is never masked again, and call a response answer-first when the final answer commits before the reasoning...

📖 Read original article


561. Hidden Language Consistency Phenomena in Reasoning LLMs ​

Author: Muhammad Ali Shafique, Kelly Marchisio
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08447v2 Announce Type: replace-cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors ...

📖 Read original article


562. Population-Scalable Multi-Agent World Modeling ​

Author: Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li, He Li, Yimin Sheng, Tianxi Tan, Zhenkai Zhang, Jiao Liang, Jianyi Zhu, Yong-Lu Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08600v3 Announce Type: replace-cross Abstract: World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed ...

📖 Read original article


563. Withholding the Completing Chunk: Exact Release-Boundary Equivalence for Production Streaming Guardrails ​

Author: Christopher M. Frost
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2608.10279v2 Announce Type: replace-cross Abstract: Streaming language-model output creates an enforcement boundary: a control that detects a prohibited pattern after releasing its completing chunk cannot recall it. We study a production policy in which each ordered family is the conjunction o...

📖 Read original article


564. TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation ​

Author: Pengyu Zhang, Yangqin Jiang, Klim Zaporojets, Congfeng Cao, Paul Groth
Published: 8/25/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.10983v2 Announce Type: replace-cross Abstract: Multi-modal recommenders fuse user-item interaction signals with item modalities such as text, images, and audio, but the usefulness of each drifts over time and at different rates. For example, around Valentine's Day, chocolate purchases bec...

📖 Read original article


565. PatientAct: Theory-Grounded Mental Health Client Simulation ​

Author: Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2608.12750v2 Announce Type: replace-cross Abstract: LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic r...

📖 Read original article


566. Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce ​

Author: Zeyuan Li, Lukas Petersson, Alessandro Acquisti, Michiel A. Bakker
Published: 8/25/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.14825v3 Announce Type: replace-cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on...

📖 Read original article


567. GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix ​

Author: Jinhyun Jeon, Sungjoo Yoo
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15584v2 Announce Type: replace-cross Abstract: Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix...

📖 Read original article


568. PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data ​

Author: Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.QM

arXiv:2608.16419v2 Announce Type: replace-cross Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environ...

📖 Read original article


569. Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors ​

Author: Hang Zhang, Kaifeng Zhang, Yixiao Ma, Weijie Xu, Ye Zhu, Kai Ming Ting
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.16700v2 Announce Type: replace-cross Abstract: Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise their legal right to have their data $D_f$ removed from a machine learning model. This process is typi...

📖 Read original article


Author: Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.17536v2 Announce Type: replace-cross Abstract: Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answ...

📖 Read original article


571. GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction ​

Author: Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.18234v2 Announce Type: replace-cross Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible...

📖 Read original article


572. Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas ​

Author: Dmitry V. Alexandrov
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CC, cs.PL

arXiv:2608.18445v2 Announce Type: replace-cross Abstract: We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a combinatorial framework originally motivated by Boolean satisfiability, based on triplet structures and a filter that we call...

📖 Read original article


573. Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling ​

Author: Tajwaar Shafiq, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.18689v2 Announce Type: replace-cross Abstract: We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate four omni LLMs in a zero-shot setting and compare them with fine-tuned models. Our results show that f...

📖 Read original article


574. AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL ​

Author: Daniele Lizzio Bosco, Jacopo Cossio, Carla Piazza, Giuseppe Serra
Published: 8/25/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2608.18946v2 Announce Type: replace-cross Abstract: Clifford circuits play a foundational role in quantum computing, particularly due to their importance in quantum error correction and fault-tolerant logical synthesis. While these circuits can be efficiently simulated and represented as sympl...

📖 Read original article


575. Interpretable AI predicts a 2026 summer dry anomaly in central China ​

Author: Anran Wang, Wen Shi, Yong Luo, Jianbin Huang, Lijuan Chen, Junhu Zhao, Weixin Jin, Huihui Yuan
Published: 8/25/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI

arXiv:2608.19163v2 Announce Type: replace-cross Abstract: Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability than precipitation itself. Here, we employ a deep learning model that translates dynamical circulation ...

📖 Read original article


576. SPADE: Self-Play in Adaptive Synthetic Executable Environments ​

Author: Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.19197v2 Announce Type: replace-cross Abstract: Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribu...

📖 Read original article


577. Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life Prediction ​

Author: Valeriu Dimidov, Rapha"el Frank
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.19218v2 Announce Type: replace-cross Abstract: Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). I...

📖 Read original article


578. Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition ​

Author: Akarsh Jain, Arya Pawa, Ayush Debnath, Smera Rawal, Sayeed Shafayet Chowdhury
Published: 8/25/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG

arXiv:2608.19232v2 Announce Type: replace-cross Abstract: Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of decision-making. Activ...

📖 Read original article


579. VGI-Bench: Probing Visual Intelligence in Video Generation Models ​

Author: Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.19583v2 Announce Type: replace-cross Abstract: Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors o...

📖 Read original article


580. Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis ​

Author: Xuan Yang, Xiaohan Yuan, Hao Li, Lingyu Chen, Yanan Liu, Qingya Li, Lei Li
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.19738v2 Announce Type: replace-cross Abstract: Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventricular meshes are not routinely available, whereas end-diastolic (ED) anatomy can often be obtained relia...

📖 Read original article


581. Question-Guided Evidence Acquisition for Multimodal Visual Question Answering ​

Author: Alin-Ionut Popa
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.19739v2 Announce Type: replace-cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting in the model's contex...

📖 Read original article


582. MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents ​

Author: Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan, Dalin Zhang, Jiqiang Liu
Published: 8/25/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.19803v2 Announce Type: replace-cross Abstract: Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping or graph-ba...

📖 Read original article


583. Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection ​

Author: Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.20169v2 Announce Type: replace-cross Abstract: We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains w...

📖 Read original article


584. AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images ​

Author: Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Prifti
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21067v2 Announce Type: replace-cross Abstract: Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leadin...

📖 Read original article


585. Atom Learning Model (ALM): how a real classroom got tokenised ​

Author: Philipp Bogdan
Published: 8/25/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.21106v2 Announce Type: replace-cross Abstract: The Atom Learning Model (ALM) tokenises a school curriculum. Two secondary mathematics textbooks were read by machine into 1,934 atoms, each one thing a learner can do in a single step, ordered by 4,616 machine-written prerequisite links. Bot...

📖 Read original article