Skip to content

arXiv cs.AI - 2026-08-26 ​

352 items collected.


1. RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation ​

Author: Yuan Si, Simeng Han, Daming Li, Jialu Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fixe...

📖 Read original article


2. ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence ​

Author: Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that d...

📖 Read original article


3. LLM Agents Perform Controlled Experiments Using Simulation Models ​

Author: Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St"umpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA, cs.SE

arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to interv...

📖 Read original article


4. A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts ​

Author: Ihor Kendiukhov
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, astro-ph.IM

arXiv:2608.23626v1 Announce Type: new Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic....

📖 Read original article


5. TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery ​

Author: Kang Zhou, Yujia Tong, Yong Tao, Jingling Yuan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci

arXiv:2608.23631v1 Announce Type: new Abstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each costly property evaluation informs the next search step. Existing agents mainly store evaluated candidates...

📖 Read original article


6. Function-Level Execution Feedback for Code Preference Optimization ​

Author: Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Superv...

📖 Read original article


7. Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes ​

Author: Heather Renze
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY

arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-t...

📖 Read original article


8. How much of a measured AI preference is the model, and how much is the instrument? ​

Author: Jason Hung
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built ...

📖 Read original article


9. AI Agents Push Humans Out of the Loop ​

Author: Margaret Mitchell, Avijit Ghosh, Samir Passi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.23642v1 Announce Type: new Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ''human in the loop'', but this is not a simple solution: Not only do current approaches to AI agent design impede e...

📖 Read original article


10. FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare ​

Author: Jacob Idoko, Siddhartha Paudel, Mariana Bento, Roberto Souza, Gouri Ginde
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23643v1 Announce Type: new Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether adoption is economically worthwhile in real clinical settings. This study proposes FLARE, a systematic...

📖 Read original article


11. Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value ​

Author: Kalin Stoyanov
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23644v1 Announce Type: new Abstract: Large language models (LLMs) are becoming routine instruments of scientific research, assisting with literature synthesis, hypothesis development, coding, and formal reasoning. Their use raises a central epistemic question: when parts of scientific rea...

📖 Read original article


12. MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models ​

Author: Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.23646v1 Announce Type: new Abstract: Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discovery, where reusable vector representations support property prediction, virtual screening, and retrieval. Most molecular encoders are special...

📖 Read original article


13. Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering ​

Author: Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.23666v1 Announce Type: new Abstract: Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains. However, it becomes particularly consequential in clinical question answering, where responses must remain grounded in the provided context and ro...

📖 Read original article


14. Automata from Agent Traces: Failure and Next-Step Prediction ​

Author: Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu, Kleyton Da Costa, Ilham Wicaksono, Adriano Koshiyama
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.23670v1 Announce Type: new Abstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so the...

📖 Read original article


15. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment ​

Author: Stephen Chung, Wenyu Du, William J. Wesley
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.DM, cs.MA

arXiv:2608.23691v1 Announce Type: new Abstract: We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own ...

📖 Read original article


16. Do LLMs Understand Limit Order Book Dynamics? ​

Author: Junxiao Chen, Paul Glasserman
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23706v1 Announce Type: new Abstract: A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect scores in generating valid sequences of LOB events. However, the LLM's implicit world model fails to learn the state of the LOB. This deficiency leads t...

📖 Read original article


17. AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace ​

Author: Seonglae Cho, Donghyun Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.23740v1 Announce Type: new Abstract: Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborative editing protocols solve this coordination problem...

📖 Read original article


18. Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware ​

Author: Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23807v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR) models, since they denoise many tokens at once. Recent systems have begun building serving infrastructure for dLLMs, but none first measure how thes...

📖 Read original article


Author: Jiongxiao Wang, Dingli Ma, Chaoqun Ni
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23811v1 Announce Type: new Abstract: Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedical domain poses unique challenges. Validating biomedical claims requires rigorous interpretation of scientific literature, assessment of ret...

📖 Read original article


20. A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification ​

Author: Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha"el Ralaivao, Alain Josu'e Ratovondrahona, Thomas Mahatody
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2608.23817v1 Announce Type: new Abstract: SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary substantially when the input is perturbed by small amounts of noise--a problem we observed firsthand in our previous work on food security in Madaga...

📖 Read original article


21. Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention ​

Author: Sergii Kozyrev (Minima AI, Inc), Davyd Maiboroda (Minima AI, Inc)
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present Minima-KV, a retention-preserving hierarchy for mixed-format paged attention. Recent and protected Anchor pages remain in FP8, while older n...

📖 Read original article


22. SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models ​

Author: Lijia Huang, Yao Fu, Sihao Ren
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether suc...

📖 Read original article


Author: Haoyang Fang, Bernie Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.23848v1 Announce Type: new Abstract: Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. In this regime, standard MCTS allocates budget poorly...

📖 Read original article


24. In-Context Inpainting for Time Series Forecasting ​

Author: Thang Nguyen, Dung Nguyen, Romero Morais, Truyen Tran
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23855v1 Announce Type: new Abstract: We propose ICI-Time, a novel framework that reframes time series forecasting as a visual inpainting task, leveraging the generalisation power of large vision models (LVMs). Unlike methods that require specialised temporal architectures and extensive do...

📖 Read original article


25. Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications ​

Author: Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23870v1 Announce Type: new Abstract: When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context, regulatory environment, organizational values, and user personas. Yet,...

📖 Read original article


26. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors ​

Author: Joshua Penman
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.LG

arXiv:2608.23873v1 Announce Type: new Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and it can lose track or be confused: text can be written to read like anyth...

📖 Read original article


27. AI Finds A Way ​

Author: Aaron Dharna, Cong Lu, Ryan Sullivan, Joel Lehman, Victoria Krakovna, Jeff Clune
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23875v1 Announce Type: new Abstract: Artificial Intelligence (AI) algorithms frequently learn creative and unexpected solutions, surprising even expert researchers who develop and study them. They often astonish practitioners by discovering unanticipated behavior, exploiting loopholes in ...

📖 Read original article


28. Provenance Guided Incremental Learning Under Evolving Concept Definitions ​

Author: Ismail Lamaakal
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.23893v1 Announce Type: new Abstract: Learning systems deployed over long periods must adapt not only to statistical changes in incoming data, but also to revisions of the definitions that generate their prediction targets. Conventional concept-drift methods typically infer such changes fr...

📖 Read original article


29. BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification ​

Author: Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, Nithin Parsan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recovered from modifications that scientists made to published protocols during real experimental work. Adapting a published protocol to a new ex...

📖 Read original article


30. Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems ​

Author: Paul Vautravers, Oliver Chalkley, Gabriel Downer, Kate S, Damian Ruck
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23906v1 Announce Type: new Abstract: Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI ev...

📖 Read original article


31. Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment ​

Author: Aryan Brar, Justin Du, Avery Lor, Kylie Seto, Eric Taylor
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23908v1 Announce Type: new Abstract: Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing it efficiently often involves complex considerations that are specific to the holdings within that portfolio and the individual who owns it. We introdu...

📖 Read original article


32. PROOF-Gen: From Optimized Data to Better Distillation ​

Author: Anh Ta, Junjie Zhu, Shahin Shayandeh
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.23911v1 Announce Type: new Abstract: Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly ...

📖 Read original article


33. MARS: Multi-Specialist LLM Relay System for Competitive Programming ​

Author: Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.PL

arXiv:2608.23918v1 Announce Type: new Abstract: Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic techniq...

📖 Read original article


34. Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining ​

Author: Yicheng Mao, Hongru Du
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, stat.ML

arXiv:2608.23922v1 Announce Type: new Abstract: Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on can...

📖 Read original article


35. Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors ​

Author: Andrew Hu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23932v1 Announce Type: new Abstract: This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforcement learning framework designed to examine how evolutionary mismatch, bounded rationality, and satisficing contribute to adaptive and maladaptive behavi...

📖 Read original article


36. More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight ​

Author: Yuchen Han, Cheng Yan, Wuyang Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23941v1 Announce Type: new Abstract: Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets planned actions before irreversible execution. Over-blocking forfeits usefulness and pressures deployers to disable it. Every protocol must fix a unit of v...

📖 Read original article


37. Recursive Agentic Reasoning ​

Author: Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23956v1 Announce Type: new Abstract: Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introduce a unified view of...

📖 Read original article


38. More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving ​

Author: Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal, Jay Iyer, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly, Raj Dandekar
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23962v1 Announce Type: new Abstract: When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight devices, buying memory headroom at the price of an all-reduce on every l...

📖 Read original article


39. Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design ​

Author: Nejla Ghaboosi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.23970v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in understanding and interpreting mul- timedia content. However, their ability to generate me- dia remains limited. Recent approaches have attempted to bridge this gap by translati...

📖 Read original article


40. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs ​

Author: Zhengxiang Wang, Owen Rambow
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.23978v1 Announce Type: new Abstract: Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: target information is often incomplete, ambiguous, and establ...

📖 Read original article


41. Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling ​

Author: Muntaser Syed, Markus Zanker, Marius Silaghi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY, cs.GT, cs.MA

arXiv:2608.23979v1 Announce Type: new Abstract: In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which arguments each voter sees, acquiring much of the decision; practice delegates it to opaque learned rankers, so a voter cannot recompute or contest th...

📖 Read original article


42. Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning ​

Author: Zhen Bi, Xueshu Chen, Yan Wang, Zhizhi Peng, Haosen Hong, Zhen Wang, Zhixuan Chu, Bingyu Zhu, Jungang Lou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.23982v1 Announce Type: new Abstract: Scientific reasoning requires language models to retrieve specialized knowledge and incorporate it reliably into multi-step computation. Conditional memory provides an explicit lookup pathway that complements dense neural representations, but its usefu...

📖 Read original article


43. Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction ​

Author: Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24001v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for future prediction, motivating the use of multiple models as a wisdom-of-the-crowd mechanism. However, simply increasing crowd size does not guarantee effective diversity, as different LLMs may exhi...

📖 Read original article


44. Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing ​

Author: Haotian Zhang, Shucun Wang, Jinze Wu, Liang Ding, Shuochen Liu, Zhenya Huang, Jing Sha, Shijin Wang, Qi Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24005v1 Announce Type: new Abstract: Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains si...

📖 Read original article


45. Reflection with Action-Induced Visual Differences for Desktop GUI Agents ​

Author: Yijie Ma, Chaoyue Niu, Fan Wu, Guihai Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24015v1 Announce Type: new Abstract: The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective alignment in complex tasks through modular collaboration. However, desktop GUIs introduce a key challenge: large, dense interfaces often exhibit subtle or...

📖 Read original article


46. Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding ​

Author: Hyunho Kook, Junhyuk So, Tianyu Fu, Haizhong Zheng, Beidi Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24024v1 Announce Type: new Abstract: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning. However, modern LLMs increasingly act as multi-turn search agents...

📖 Read original article


47. Relative Time Intervals Representation for Word-level Timestamping with Masked Training ​

Author: Quanwei Tang, Zhiyu Tang, Xu Li, Dong Zhang, Shoushan, Guodong Zhou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24041v1 Announce Type: new Abstract: Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored. Our work addresses this gap by enabling SpeechLLMs to jointly model spee...

📖 Read original article


48. Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment ​

Author: Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24046v1 Announce Type: new Abstract: When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standa...

📖 Read original article


49. Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems ​

Author: CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee, Jee-Hyong Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CE

arXiv:2608.24069v1 Announce Type: new Abstract: LLM-based multi-agent trading systems, in which specialized agents collaborate through structured communication to produce trading decisions, are moving rapidly from research prototypes to live deployments that control real assets. The same inter-agent...

📖 Read original article


50. Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression ​

Author: Mohammad Mozaffari
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG, cs.PF

arXiv:2608.24070v1 Announce Type: new Abstract: Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an ...

📖 Read original article


51. AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval ​

Author: Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24076v1 Announce Type: new Abstract: Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN)...

📖 Read original article


52. EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals ​

Author: Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.SE

arXiv:2608.24086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis, but their ability to analyze raw physical-layer measurements remains untested. We introduce \textbf{EMRB} (\textbf{E}lectro\textbf{m}agnetic \tex...

📖 Read original article


53. Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments ​

Author: Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang, Lin Fu, Hong Zhou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24099v1 Announce Type: new Abstract: GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehens...

📖 Read original article


54. ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation ​

Author: JooYoung Jang, Taegyeong Lee, Jihyeon Park, Nojun Kwak
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24103v1 Announce Type: new Abstract: Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so agents must rec...

📖 Read original article


55. Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing ​

Author: Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24112v1 Announce Type: new Abstract: Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fixed categories, weak...

📖 Read original article


56. AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL ​

Author: Xiaolong Jin, Dingmin Wang, Vijay Lingam, Varun Kumar
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24114v1 Announce Type: new Abstract: Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards, which assign a uniform advantage to every step and cannot identify which decisions led to success or failure. Self-distillation methods can provide...

📖 Read original article


57. Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping ​

Author: Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui, Zujie Wen, Zhiqiang Zhang, Jun Zhou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the code generation capabilities of Large Language Models (LLMs). However, the efficacy of RLVR in coding implementations is fundamentally limited by...

📖 Read original article


58. OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses ​

Author: Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they sc...

📖 Read original article


59. Task-Adaptive Rubrics for GUI Reward Modeling ​

Author: Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, ho...

📖 Read original article


60. Paritok-4B: Intent-Conditioned Context Compression for Coding Agents ​

Author: Jiayu Shi, Luzhuo Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SE

arXiv:2608.24188v1 Announce Type: new Abstract: Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code poorly: they paraphrase identifiers and drop the ex...

📖 Read original article


61. Preference Data Selection for Mitigating the Alignment Tax in Large Language Models ​

Author: Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is crucial for real-world deployment but frequently incurs an alignment tax, leading to the catastrophic forgetting of pre-trained general capabilities. While previous works primarily frame this probl...

📖 Read original article


62. MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG ​

Author: Qiuyi Qi, Tian Liang, Jiamu Wang, Jinjian Zhang, Wei Zhou, Pengcheng Zhu, Linjian Mo, Ming Kong, Jie Liu, Qiang Zhu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervision and overlook the agent's internal belief about whether the current evi...

📖 Read original article


63. Constraint-Guided Enterprise Data Mapping with Large Language Models ​

Author: Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.24218v1 Announce Type: new Abstract: Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or granularity mismatches. Manual matching is still common in practice, but does not scale as schemas and providers evolve. LLM-only matching improves semant...

📖 Read original article


64. Evaluating Multiple LLM Generations with Validated Task Coverage ​

Author: Florian Le Bronnec, Rio Yokota
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation settings, however, still focus on individual outputs or reduce multiple samples to a single success or ...

📖 Read original article


65. TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models ​

Author: Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan, Talal Hammad, Faisal Wahbo, Aminullah Tora, Mona Alshahrani, Xin Gao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchmarks for unsafe cont...

📖 Read original article


66. STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation ​

Author: Junyeong Maeng, Eunsong Kang, Heung-Il Suk
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their changes relative to a prior study. Existing methods jointly model diagnosis, attribute estimation, temporal comparison, and language generation within ...

📖 Read original article


67. SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction ​

Author: Xue Hu, Zewei Pan, Zeli Su, Zhou Liu, Wentao Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24252v1 Announce Type: new Abstract: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful implementations. We define this failure mode as semantic drift, where generated code silently diverges from the paper's specifications. We introduce SemanticAl...

📖 Read original article


68. Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks ​

Author: Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be justifiable against an applicable legal standard. Existing legal-AI benchmarks and LLM-as-judge protocols provide important infrastructure for me...

📖 Read original article


69. Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing ​

Author: Yaoyi Qi, Xingxing Weng, Chao Pang, Yongkang Cui, Xiangyu Hao, Xiaokang Zhang, Guibo Zhu, Gui-Song Xia
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving the performance of change detection models. However, existing synthesis methods typically rely on handcrafted rules to simulate changes, where limited co...

📖 Read original article


70. Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding ​

Author: Hao Ren, Junbin Gao, Jiaojiao Jiang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.IR

arXiv:2608.24273v1 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing methods primarily address catastrophic forgetting, but entity admission also changes the candidate universe of every compatible query. A historic...

📖 Read original article


71. Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling ​

Author: Ying Jin, Weiqing Min, Mingyu Huang, Shuqiang Jiang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24274v1 Announce Type: new Abstract: A sustainable diet represents a multi-dimensional synergy among four essential pillars: nutrition adequacy, economic affordability, cultural acceptability, and environmental respect. Despite the prevalence of population-level sustainability modeling, p...

📖 Read original article


72. RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards ​

Author: Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.24275v1 Announce Type: new Abstract: Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unsee...

📖 Read original article


73. ReproAgent: Contract-Guided Paper-to-Code Reproduction ​

Author: Xue Hu, Zewei Pan, Zhongyuan Wang, Zhou Liu, Zeli Su, Wentao Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24291v1 Announce Type: new Abstract: Paper-to-code reproduction asks scientific AI agents to turn research papers into executable repositories that preserve the paper's method, protocol and artifacts. This is difficult because the specification is split: explicit paper content such as alg...

📖 Read original article


74. VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models ​

Author: Guoyang Xu, Hao Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much longer video. Existing approaches improve this process through compression, retrieval, memory, and agentic evidence acquisition, but these mechanisms ...

📖 Read original article


75. OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning ​

Author: Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24310v1 Announce Type: new Abstract: Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retr...

📖 Read original article


76. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight ​

Author: Anupam Purwar, Shashank Singh, Kritika Srivastava
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.ET

arXiv:2608.24314v1 Announce Type: new Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by co...

📖 Read original article


77. Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control ​

Author: Francisco M. Arrabal-Campos, Ignacio Fernandez, Francisco G. Montoya, Alfredo Alcayde
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2608.24319v1 Announce Type: new Abstract: An intelligent system does not merely reason: it governs its own reasoning - how much to compute, when to stop, which module to activate. Can that role be played by a dynamic internal field - a low-dimensional homeostatic state with explicit physics an...

📖 Read original article


78. SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception ​

Author: Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24325v1 Announce Type: new Abstract: Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimu...

📖 Read original article


79. Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning ​

Author: Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung, Yi-an Lai, Arshit Gupta
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24338v1 Announce Type: new Abstract: Inference-time decoding methods improve LLM reasoning by exploring multiple candidate trajectories, yet treat each trajectory as atomic: either retaining it whole or discarding it irreversibly. This wastes computation on partially promising candidates ...

📖 Read original article


80. The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents ​

Author: Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24358v1 Announce Type: new Abstract: Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unfold, users face a practical cost-quality trade-off: escalating to a stronger model when a cheaper one struggles, or downshifting once t...

📖 Read original article


81. Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems ​

Author: Yarden Bakish, Amir Dudai, Roy Ganz, Oren Nuriel, Elad Ben Avraham, Mor Shpigel Nacson, Ron Litman
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24361v1 Announce Type: new Abstract: Multi-agent LLM systems are increasingly deployed in real-world applications, where failures can be costly and difficult to localize. Despite growing efforts to automate failure attribution, diagnosing failed runs still largely relies on human engineer...

📖 Read original article


82. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use ​

Author: Rongfeng Guo, Yinxuan Huang, Yusen Wu, Maoqing Zhong, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24368v1 Announce Type: new Abstract: Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each action remains consistent with it. However, direct function-calling and ReAct-style policies learn state tracking and action generation within the sa...

📖 Read original article


83. Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs ​

Author: Lei Jiang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24369v1 Announce Type: new Abstract: While large language models (LLMs) possess vast zero-shot procedural knowledge, their tendency to produce homogenized logic often obscures the unique, idiosyncratic execution processes of individual human creators. In this paper, we investigate the com...

📖 Read original article


84. ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping ​

Author: Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24411v1 Announce Type: new Abstract: The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of autoregressive decoding. Speculative Decoding (SD) mitigates this by using a lightweight draft model to speculate future tokens, which are then va...

📖 Read original article


85. A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation ​

Author: Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24419v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the ...

📖 Read original article


86. Partial Identification under Causal Orders by Linear Programming ​

Author: Eric Rossetto, Alessandro Antonucci
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24427v1 Announce Type: new Abstract: Non-parametric (partial) identification of counterfactual queries typically relies on a fully specified causal graph. Motivated by settings with incomplete domain knowledge, we challenge this requirement by leveraging structural assumptions that are in...

📖 Read original article


87. A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads ​

Author: Chenghan Li, Qingxiang Liu, Yinliang Xu, Yuxuan Liang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24441v1 Announce Type: new Abstract: Electric vehicle (EV) charging loads exhibit strong behavioral heterogeneity and temporal variability, posing significant challenges for online probabilistic forecasting under evolving operating conditions. In particular, persistent charging patterns m...

📖 Read original article


88. Mahalanobis-Based Multi-Head Attention for Complex State Propagation ​

Author: Xiaohe Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24462v1 Announce Type: new Abstract: In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attention mechanism that replaces the standard dot-product with a \textbf{Mahalanobis distance-based RBF kernel}, which effectively computes attention in an in...

📖 Read original article


89. HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning ​

Author: Qiuyu Zhu, Yi Gao, Zhichao Wan, Mingyang Ma
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24467v1 Announce Type: new Abstract: Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. This limitation ...

📖 Read original article


90. Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling ​

Author: He Wang, Junyu Wu, Hui Li, Yanjie Song, Witold Pedrycz, Liang Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24470v1 Announce Type: new Abstract: Heterogeneous agile Earth observation satellite (AEOS) scheduling requires task selection, satellite assignment, and observation sequencing under satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard...

📖 Read original article


91. Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites ​

Author: He Wang, Junyu Wu, Yeye Liu, Yifan Zhou, Jie Zhang, Hui Li, Yanjie Song, Liang Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24471v1 Announce Type: new Abstract: Maritime moving-target observation scheduling with agile Earth observation satellites is a dynamic, sequence-dependent combinatorial optimization problem. Sea-surface targets move continuously, causing feasible observation windows to vary with target m...

📖 Read original article


92. PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents ​

Author: Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24509v1 Announce Type: new Abstract: LLM agents increasingly solve tasks by invoking multiple tools, where parallel execution is essential for low latency but difficult to manage safely. Existing agent benchmarks primarily evaluate tool selection, argument generation, and end-to-end succe...

📖 Read original article


93. Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models ​

Author: Abdulhady Abas Abdullah, Erik Cambria, Milena Zivkovic
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24534v1 Announce Type: new Abstract: Clinical LLMs can generate recommendations that are factually plausible yet physiologically unsafe. We investigate whether safety alignment can be improved by grounding preference optimization in structured physiological knowledge rather than text-only...

📖 Read original article


94. Discovering Adaptive Transmission Programs for Collective Innovation ​

Author: C'edric Colas, J'er'emy Perez, Eleni Nisioti, Akhilesh Mocherla, Pierre-Yves Oudeyer, Cl'ement Moulin-Frier, Maxime Derex
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24545v1 Announce Type: new Abstract: Human collective intelligence depends on transmission processes: who shares what with whom, how, and when. While these processes emerge from individual cognition, they can also be directed by deliberate top-down protocols. Prior work has studied how tr...

📖 Read original article


95. When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows ​

Author: Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.24569v1 Announce Type: new Abstract: Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from...

📖 Read original article


96. EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents ​

Author: Lihang Zeng, Shaoting Zhang, Xiaofan Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24570v1 Announce Type: new Abstract: Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large langua...

📖 Read original article


97. Joint Optimization of Tool Creation and Use for Large Language Model Agents ​

Author: Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24571v1 Announce Type: new Abstract: Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no si...

📖 Read original article


98. PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos ​

Author: Siyao Yan, Bo Han, Jisheng Dang, Bimei Wang, Shude Wang, Hong Peng, Yulan Guo, Jianhuang Lai, Bin Hu, Tat-SengChua
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24574v1 Announce Type: new Abstract: Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches. These failures are more common when targets are partly hidden or when si...

📖 Read original article


99. Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms ​

Author: Andrea Di Nezza, Mihir Patel, Fabio Fagnani, Sara Bernardini
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24585v1 Announce Type: new Abstract: Automated high-density storage systems (warehouses, robotic parking, plant logistics, etc.) require fleets of agents to move through scarce task-critical resources and then park without obstructing future operations. We introduce Pivot-and-Station Mult...

📖 Read original article


100. Causal Modelling of Support Interventions for Student Competency Assessment ​

Author: Francesca Mangili, Alessandro Antonucci, Rafael Caba~nas
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24632v1 Announce Type: new Abstract: Accurate assessment of student competencies is essential for enabling educators to identify individual needs, design targeted interventions, and evaluate the effectiveness of educational strategies. Empirical assessment procedures are typically grounde...

📖 Read original article


101. Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning ​

Author: Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24658v1 Announce Type: new Abstract: Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for difficult tasks (up ...

📖 Read original article


102. The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models ​

Author: Augusto Camargo
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY

arXiv:2608.24662v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated under the assumption that their observable behavior is primarily determined by model weights, training data, alignment procedures, and user prompts. This view is incomplete. Modern inference pipelines...

📖 Read original article


103. Confident at the moment of action: belief miscalibration in LLM play under hidden information ​

Author: Bhushan Kashinath Joshi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.24691v1 Announce Type: new Abstract: Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocat...

📖 Read original article


104. Lifted Model Construction under Approximate Commutativity ​

Author: Malte Luttermann, Jan Speller, Tanya Braun, Marcel Gehrke, Ralf M"oller
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.24713v1 Announce Type: new Abstract: Lifted inference algorithms enable scalable probabilistic inference even for large object domains by leveraging the indistinguishability of objects in a probability distribution. An essential prerequisite for constructing a lifted representation is to ...

📖 Read original article


105. Meta$^n$: Recursive Self-Improvement through Emergent Depth ​

Author: Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SY, eess.SY

arXiv:2608.24735v1 Announce Type: new Abstract: Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level hold that level fixed, and those that edit themselves must leave part of their own editing machinery untouched to stay stable, capping ...

📖 Read original article


106. RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons ​

Author: Runyu Wang, Bo Liu, Xiaxin Zhang, Yu Han, Jiawei Cao, Xiaoye Zhang, Zhe Zhang, Yifan Yang, Peng Ping
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24758v1 Announce Type: new Abstract: Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or computationally expensive procedures, which either obscure population-level v...

📖 Read original article


107. Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav ​

Author: Hongyu Guo, Zhiyu Zheng, Zhao Cao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24764v1 Announce Type: new Abstract: Large language model agents are moving beyond conventional retrieval-augmented generation toward direct interaction with external corpora. Direct Corpus Interaction (DCI) keeps the full corpus accessible, yet reachable evidence can remain unusable unde...

📖 Read original article


108. StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing ​

Author: Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.24777v1 Announce Type: new Abstract: LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed ...

📖 Read original article


109. Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought ​

Author: Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24790v1 Announce Type: new Abstract: Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely tested. General-domain CoT-faithfulness probes ignore clinical cost, and medical LLM evaluations treat the chain...

📖 Read original article


110. CAFE: Self-Improving Search Agents Need Co-Evolving Feedback ​

Author: Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou, Xinbing Liang, Shizheng Zhu, Yuhui Wang, Jingqi Tong, Zhiheng Xi, Jiazheng Zhang, Clive Bai, Clarenceai, Blaze Chen, Tao Gui, Qi Zhang, Xuanjing Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24794v1 Announce Type: new Abstract: Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory...

📖 Read original article


111. StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments ​

Author: Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala, Sagar Davasam
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.24804v1 Announce Type: new Abstract: We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and age...

📖 Read original article


112. Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core ​

Author: Yogesh Kumar
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24810v1 Announce Type: new Abstract: Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches still rely on buffering clips or windows internally, lack a theoretical account of how temporal memory relates to detection latency, and b...

📖 Read original article


113. Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA ​

Author: Emanuel Kitzelmann
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24824v1 Announce Type: new Abstract: Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to correctly ground answers in the underlying graph. Current approaches to LLM-based KGQA either rely on full semantic parsing into executable queri...

📖 Read original article


114. A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments ​

Author: Jing Huang, Jihong Zhang, Hua-Hua Chang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24825v1 Announce Type: new Abstract: The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become uninten...

📖 Read original article


115. FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs ​

Author: Md Saikat Islam Khan Bappy, Oshani Seneviratne
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24846v1 Announce Type: new Abstract: Real-world data for knowledge graph question answering is often distributed across different organizations due to governance and data sovereignty constraints. While centralized systems exist, they cannot answer multi-hop questions when the required fac...

📖 Read original article


116. SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL ​

Author: Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24870v1 Announce Type: new Abstract: Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estim...

📖 Read original article


117. Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses ​

Author: Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.24876v1 Announce Type: new Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harness...

📖 Read original article


118. Progressively Learning Heterogeneous Skills in a Unified Latent Space ​

Author: Yue-Yi Zhang, Ming Gong, Linpu He, Wei-Shi Zheng, Zhilin Zhao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2608.23258v1 Announce Type: cross Abstract: We propose HetSkills, a novel framework designed to progressively learn heterogeneous skills within a unified latent space for physics-based character control. The core idea is to treat this latent space as a shared executable interface, enabling sea...

📖 Read original article


119. A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision ​

Author: Vasiliki Kondyli, Jakob Suchan, Mehul Bhatt
Published: 8/26/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CV

arXiv:2608.23572v1 Announce Type: cross Abstract: We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of complexity in embodied perception and interaction in dynamic, naturalistic settings. Grounded in theor...

📖 Read original article


120. Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers ​

Author: Mingyang Xu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23593v1 Announce Type: cross Abstract: Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, but whether these scores encode demographic attributes as objective quality is unclear. We audit four scorers (LAION-Aesthetics, PickScore, ImageReward,...

📖 Read original article


121. REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring ​

Author: Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23611v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated changes must reduce targeted quality problems without introducing new issues or altering behaviour-relevant code structures. We introduce REFINE (...

📖 Read original article


122. Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal ​

Author: Parker Fawcett
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23616v1 Announce Type: cross Abstract: An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instruction (AgentMo...

📖 Read original article


123. Identifying Latent Declarative Representations of Code for Assisting Repository Migration ​

Author: Shraddha Surana, Ashwin Srinivasan, Michael Bain
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23619v1 Announce Type: cross Abstract: Legacy software repositories embed decades of domain knowledge in undocumented code, making understanding and modernization difficult. We treat a program as the implementation of an unobserved, declarative description of its computation and investiga...

📖 Read original article


124. When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs ​

Author: Jason Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify execution traces, or enforce runtime polici es, but do not test this particular receipt-, scope-, and closed-replay design at the COMPLETE boundary ac...

📖 Read original article


125. Macro-Operator Generation and Predicate Selection for TAMP Operator Learning ​

Author: Can Emir Bora, Emre Ugur
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.23629v1 Announce Type: cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show that these operators can instead be learned directly from demonstration data. Existing methods, however, typica...

📖 Read original article


126. ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents ​

Author: YiShan Zheng, Yuan Wu, Yi Chang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to invoke external systems and complete tasks beyond text generation. However, clean end-to-end (E2E) success cannot identify where a tool-use failure ...

📖 Read original article


127. Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail ​

Author: Esmail Gumaan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information. We measure whether it is. Defining the corrective gain of a failure record as...

📖 Read original article


128. Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling ​

Author: Zizhe Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23653v1 Announce Type: cross Abstract: AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents different requirements from general-purpose code generation in software engineering, because correctness depends not only on syntax and executability...

📖 Read original article


129. Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap ​

Author: Sathishkumar Sivashanmugam
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2608.23658v1 Announce Type: cross Abstract: An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a reserve for the worst-case prefill activation. During decode-dominant phases that reserve sits idle, yet it cannot be handed to the KV pool because it ...

📖 Read original article


130. From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers ​

Author: Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu, Taoran Ji
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ME

arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tun...

📖 Read original article


131. Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model ​

Author: Shashwat Pandey, Satwik Pandey, Suresh Raghu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.23663v1 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderation, and the configuration developers can actually deploy is rarely audi...

📖 Read original article


132. The Limits of Automatic Evaluation of Creativity in Large Language Models ​

Author: Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate wheth...

📖 Read original article


133. Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it ​

Author: Irene Trigueros-Lorca, Leonardo Concepci'on, Christian Wagner, Isaac Triguero, Daniel Molina
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23752v1 Announce Type: cross Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models. Knowledge Distillation (KD) addresses this by transferring knowledge from a large network (teacher) to a small one (student), also reducing the trainin...

📖 Read original article


134. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$ ​

Author: Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.23758v1 Announce Type: cross Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typically focus on a single audio moda...

📖 Read original article


135. TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers ​

Author: Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23763v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised MCP server behaves benignly du...

📖 Read original article


136. What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development ​

Author: Christopher Brooks (School of Information, University of Michigan)
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23766v1 Announce Type: cross Abstract: Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psy...

📖 Read original article


137. Disentangled Skill Representations for Predictive Human Modeling ​

Author: Mariah Schrum, Deepak Gopinath, Srijan Srivatsa, Guy Rosman, Tiffany Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23776v1 Announce Type: cross Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounde...

📖 Read original article


138. When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk ​

Author: Liliana Santos-Deonizio, James Malamut, Ram'on Mart'inez, Dorottya Demszky
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2608.23780v1 Announce Type: cross Abstract: LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal ...

📖 Read original article


139. EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis ​

Author: Tianchi Liu, Zeyang Song, Tianrui Wang, Zhipeng Li, Chenglin Xu, Yiwen Guo
Published: 8/26/2026, 4:00:00 AM
Categories: eess.AS, cs.AI

arXiv:2608.23791v1 Announce Type: cross Abstract: Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discre...

📖 Read original article


140. Restoring Without Forgetting: Continual Learning Across Image Degradations ​

Author: Alif Ashrafee, Bartosz Krawczyk
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23799v1 Announce Type: cross Abstract: Recent progress in image restoration has converged on all-in-one architectures that jointly handle multiple degradations within a single network. These methods are effective on static benchmarks but target a closed-world setting that assumes simultan...

📖 Read original article


141. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology ​

Author: Marie-Lisa Eich, Kai Standvoss, Timo Milbich, Alexander M"ollers, Miriam H"agele, Philipp Anders, Lars Tharun, Hanna Kontradiuk, Sebastian Kons, Nader Aldoj, Recepcan Adig"uzel, Adam Narai, Lukas H"onig, Jonathan Striebel, Binru Yang, Mihnea P. Dragomir, Marvin Sextro, Philipp Keyl, Philipp Jurmeister, Rosemarie Krupar, Evelyn Ramberger, James Wells, Julika Ribbat-Idel, Andreas Kunft, Hussam Shuaib, Christian Groh'e, Reinhard B"uttner, David Horst, Klaus-Robert M"uller, Lukas Ruff, Maximilian Alber, Frederick Klauschen, Simon Schallenberg
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23803v1 Announce Type: cross Abstract: Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immunohistochemical, and molecular features. Yet pathological assessment remains largely visual and semi-quantitative...

📖 Read original article


142. Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders ​

Author: Igor Bogdanov, Changcheng Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.23809v1 Announce Type: cross Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We study this question in...

📖 Read original article


143. Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring ​

Author: Olga Manakina, Igor Bogdanov
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.23814v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal configurations. We p...

📖 Read original article


144. Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB ​

Author: Alireza Mohammadhosseini, Fatemeh Afghah
Published: 8/26/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.23824v1 Announce Type: cross Abstract: Unmanned aerial vehicle (UAV)-mounted 5G New Radio base stations (gNBs) can augment terrestrial networks with an on-demand, repositionable Frequency Range 2 (FR2) capacity layer. This flexibility, however, couples the physical network topology with r...

📖 Read original article


145. Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation ​

Author: Leila Khaertdinova, Anna Anikina, Claudia Mello-Thoms, Bulat Ibragimov
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23836v1 Announce Type: cross Abstract: Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention to diagnostically relevant regions. While eye-tracking has been widely studied in 2D medical imaging, its use for expertise assessment in CT sett...

📖 Read original article


146. Infant Care Video Dataset for Classification of Interventions Using Transformers ​

Author: Igor Bogdanov, James Green
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.23838v1 Announce Type: cross Abstract: Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges, with nurses spending approximately 25% of their time on record-keeping, while up to 60% of interventions remain undocumented. Motivated by the need...

📖 Read original article


147. Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization ​

Author: Yapeng Liu, Yuanzhao Zhai, Xudong Gong, Dawei Feng, Bo Ding, Lin Wang, Huaimin Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.23839v1 Announce Type: cross Abstract: Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliability directly dictates deployment quality and human-agent trust. However, existing evaluations rely on outcome-centric metrics as success rate or safe...

📖 Read original article


148. ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork ​

Author: Tim Beringer (Technical University of Darmstadt), Patrick Diem (Technical University of Darmstadt), Felix Wolf (Technical University of Darmstadt), Arya Mazaheri (Technical University of Darmstadt, PanocularAI)
Published: 8/26/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.23840v1 Announce Type: cross Abstract: Training large-scale AI models often outgrows a single data center, demanding sharded, multi-cluster, and decentralized training. However, the huge space of resource allocations makes exhaustive benchmarking and manual tuning impractical, while perfo...

📖 Read original article


149. Automated Synthesis of Cloud Emulators ​

Author: Archit Bhatnagar, Zhenning Yang, Sarah McClure, Yiming Qiu, Sylvia Ratnasamy, Ang Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.DC

arXiv:2608.23842v1 Announce Type: cross Abstract: DevOps programming (e.g., using CLI/API scripts or IaC frameworks) is key to cloud infrastructure management. Unlike traditional programming tasks, DevOps program testing needs provisioning and execution against actual cloud resources, which is often...

📖 Read original article


150. Coronavirus Optimization Algorithm: A Success-History Adaptive Evolutionary Framework with Archive-Assisted Search and Stagnation Recovery for Global Optimization ​

Author: Hari Mohan Pandey
Published: 8/26/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CC

arXiv:2608.23847v1 Announce Type: cross Abstract: This paper proposes the Coronavirus Optimization Algorithm (COA), a SARS-CoV-2-inspired success-history adaptive evolutionary optimizer for box-constrained continuous global optimization. COA does not model disease transmission; instead, it maps sele...

📖 Read original article


151. Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2) ​

Author: Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23858v1 Announce Type: cross Abstract: The Agent Payments Protocol (AP2), introduced by Google, enables large language model (LLM)-driven shopping agents to authorize and execute payments on behalf of users. Its signed Checkout and Payment Mandates protect the integrity of transaction dat...

📖 Read original article


152. Revelation Control ​

Author: Qinyou Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.23860v1 Announce Type: cross Abstract: Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention...

📖 Read original article


153. A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization ​

Author: Prithvi Dake, Rahul Bindlish, James B. Rawlings
Published: 8/26/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2608.23885v1 Announce Type: cross Abstract: Real-time optimization (RTO) relies on process models to locate economically optimal operating conditions. Because developing first-principles models requires significant process knowledge, data-driven alternatives are increasingly attractive. Modern...

📖 Read original article


154. A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource ​

Author: Blake Reynolds
Published: 8/26/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT

arXiv:2608.23892v1 Announce Type: cross Abstract: This article presents the abridged core of \emph{A Mathematical Theory of Interpretation} (MTI), which treats interpretation as observer-relative spectral measurement under an access structure. MTI makes interpretation a method-design problem: access...

📖 Read original article


155. Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory ​

Author: Danish Khan, Maurice D. Hanisch, Nikolai Argatoff, Evan Xie, Sandeep Sharma, Anima Anandkumar
Published: 8/26/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI

arXiv:2608.23895v1 Announce Type: cross Abstract: Kohn--Sham density functional theory (DFT) underpins electronic-structure simulations, but repeated orbital diagonalizations lead to cubic scaling, restricting quantum calculations to modest scales only. Eliminating these auxiliary orbitals while ret...

📖 Read original article


156. Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs ​

Author: Akash Raj, Sargam Sahu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23897v1 Announce Type: cross Abstract: When a code generating language model fabricates a Python package name, an adversary who has pre-registered that name on PyPI can convert that hallucination into a supply chain compromise. This event has been termed as 'slopsquatting'. We propose a t...

📖 Read original article


157. RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding ​

Author: Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23928v1 Announce Type: cross Abstract: Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the question asks about. Existing approaches face a trade-off: vision language models understand the question context b...

📖 Read original article


158. QML for Quantum Sensing under Measurement-Induced Information Loss ​

Author: Sounak Bhowmik, Himanshu Thapliyal
Published: 8/26/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2608.23934v1 Announce Type: cross Abstract: Nitrogen-vacancy (NV) centers in diamond can serve as highly sensitive solid-state quantum sensors for high-sensitivity magnetometry. However, in the noisy intermediate-scale quantum (NISQ) era, extracting reliable information from noisy, finite-shot...

📖 Read original article


159. Luce: Relightable Gaussians for 3D Asset Generation ​

Author: Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi, Muhammad Ahmed Riaz, Behrooz Shahsavari, Waleed Abdulla, David E. Jacobs
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR

arXiv:2608.23943v1 Announce Type: cross Abstract: High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) ...

📖 Read original article


160. STAIN-FL: Stealthy Targeted Attack Injection with Contextual Triggers in Federated Learning ​

Author: Ashlinder Kaur, Purnima Murali Mohan, Zengxiang Li, Tram Truong-Huu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.23952v1 Announce Type: cross Abstract: Federated video anomaly detection trains model collaboratively without sharing raw surveillance footage, but limited server-side visibility lets compromised clients to inject backdoor via malicious updates. This paper introduces STAIN-FL, a stealthy ...

📖 Read original article


161. The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses ​

Author: Dai Jiahong
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CE

arXiv:2608.23953v1 Announce Type: cross Abstract: An agent harness is what turns a language model into an autonomous agent: the surrounding code that builds the model's context, mediates its tools, runs the loop, and persists state across a long-horizon run. This layer, not the model it wraps, is in...

📖 Read original article


162. NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution ​

Author: Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.IR, cs.LG

arXiv:2608.23959v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. Jailbreak attacks bypass safety mechanisms through crafted prompts, while neuron-level attacks directly prune safety-critical neurons post-deploym...

📖 Read original article


163. Evaluating Language Models on Cross-Language Code Functional Equivalence ​

Author: Hui Sun, Anderson Uch^oa, Rohit Gheyi, Wesley K. G. Assun\c{c}~ao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2608.23961v1 Announce Type: cross Abstract: Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to believe that they can reason about program semantics. However, existing evaluations primarily focus on single-...

📖 Read original article


164. RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation ​

Author: Yueyang Quan, Anjun Gao, Yufei Xia, Minghong Fang, Zhuqing Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.IR, cs.LG

arXiv:2608.23965v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models by grounding responses in external documents, but it also exposes a critical security vulnerability: adversarial documents injected into the knowledge database can ...

📖 Read original article


165. The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem ​

Author: Elioth Sanabria
Published: 8/26/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.PF

arXiv:2608.23986v1 Announce Type: cross Abstract: Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context. The industry's accounting says this saves money. We sho...

📖 Read original article


166. Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation ​

Author: Olympia Saha, Amy Wang, Srinivasan Manoharan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.23992v1 Announce Type: cross Abstract: Large language model (LLM) agents invoke external tools to retrieve and reason over information beyond pretrained knowledge. The Model Context Protocol (MCP) standardizes how such tools are surfaced, and a proxy MCP server aggregates many backend ser...

📖 Read original article


167. SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding ​

Author: Yuchuan Wu, Xuan Luo, Yinglian Zhu, Meng Fang, Xiangyang Xue, Bin Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24011v1 Announce Type: cross Abstract: Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly g...

📖 Read original article


168. WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents ​

Author: Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.24017v1 Announce Type: cross Abstract: The emerging W3C WebMCP proposal enables LLM agents to invoke tools exposed by web pages. In multi-party web environments, however, integrating agent execution into a browser security model centered on the Same-Origin Policy (SOP) leaves insufficient...

📖 Read original article


169. IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views ​

Author: Yuchuan Wu, Ke Niu, Haiyang Yu, Zhuofan Chen, Xiangyang Xue, Bin Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24020v1 Announce Type: cross Abstract: Generating executable parametric CAD code from dimension-annotated orthographic drawings is a challenging task requiring geometric understanding, procedural reasoning, and precise numerical prediction. Existing vision-language approaches typically fo...

📖 Read original article


170. What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions ​

Author: Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Puhan Luo, Ruiqi Li, Zhiqiang Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.24022v1 Announce Type: cross Abstract: LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructi...

📖 Read original article


171. ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning ​

Author: Juntao Fang, Shifeng Xie, Ruichu Cai, Shengji Zheng, Zijian Li, Keli Zhang, Lujia Pan, Themis Palpanas, Zhifeng Hao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.24033v1 Announce Type: cross Abstract: Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models support forecasting and transferable representation learning, classification still typically requires fitting ...

📖 Read original article


172. Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings ​

Author: Muhammad Tayyab Khan, Lequn Chen, Wenhe Feng, Seung Ki Moon
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.24039v1 Announce Type: cross Abstract: Manufacturing process planning transforms heterogeneous design information into coherent manufacturing decisions. However, existing approaches focus on isolated subtasks, such as feature recognition, drawing interpretation, or tool selection, and str...

📖 Read original article


173. Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models ​

Author: Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.24042v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. Retrieval offers...

📖 Read original article


174. Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding ​

Author: Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.24048v1 Announce Type: cross Abstract: While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and p...

📖 Read original article


175. VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference ​

Author: Lyuke Wang, Zhuo Li, Guangxu Zhu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24063v1 Announce Type: cross Abstract: While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key-Value (KV) caches. Exis...

📖 Read original article


176. Mechanistic Circuit Identification for Controllable Data Generation ​

Author: Nakyung Lee, Sangwoo Hong, Jungwoo Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.24065v1 Announce Type: cross Abstract: While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. This black-box paradigm provides limited insight into how individual samples interact with a model's ...

📖 Read original article


177. ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal ​

Author: Bohan Zhang, Chenyu Xu, Yijie Mao, Yuanming Shi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV, cs.DC

arXiv:2608.24073v1 Announce Type: cross Abstract: Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and environmental surveillance. However, cloud coverage often obscures the Earth's surface, and conventional cloud-rem...

📖 Read original article


178. When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs ​

Author: Zhiyang Qi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24080v1 Announce Type: cross Abstract: In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, and encour...

📖 Read original article


179. PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding ​

Author: Md Mahadi Hasan Nahid, Davood Rafiei
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.24082v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong capabilities in table reasoning, but their effectiveness degrades as tables grow in size and complexity due to irrelevant context and difficulty localizing the evidence required for reasoning. Existing a...

📖 Read original article


180. Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents ​

Author: Nadeem Shaikh
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.24087v1 Announce Type: cross Abstract: Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own reasoning, that it is...

📖 Read original article


181. MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes ​

Author: Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy, See-Kiong Ng, Luu Anh Tuan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24107v1 Announce Type: cross Abstract: Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surroundings, and illumination. Despite its commercial relevance, no public benchmark isolates this task, and evalu...

📖 Read original article


182. Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection ​

Author: Jungwook Seo, Sangwon Son, Minjeong Kim, Seungmin Han, Seojin Yoo, Sungyong Baik
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24113v1 Announce Type: cross Abstract: Time-series anomalies can appear not only as pointwise deviations but also as changes in recurring temporal structure, such as shifted periodicity or localized oscillatory fluctuations. However, existing LLM-based time-series anomaly detection method...

📖 Read original article


183. PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control ​

Author: Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.24115v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without usi...

📖 Read original article


184. TransPhy: Visual In-Context Learning for Physically Grounded Image Editing ​

Author: Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang, Zukai Chen, Lei Yang, Quan Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24119v1 Announce Type: cross Abstract: Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation tra...

📖 Read original article


185. Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking ​

Author: Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc Pham-Nam Ho, Trinh Le Ba Khanh, Chi Dai Tran, Duong Khac Vu, Son Hong Phan, Hyung-Min Jeon, Jae Wook Jeon
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24130v1 Announce Type: cross Abstract: Multi-camera 3D perception systems for warehouse scenes are trained largely on synthetic data and evaluated on physically captured environments. The resulting synthetic-to-real gap, which corrupts ground-plane localization and cross-camera identity a...

📖 Read original article


186. From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender ​

Author: Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24132v1 Announce Type: cross Abstract: Product catalogs in fast-moving service businesses are shifting from static, independently priced SKUs toward dynamically bundled, discount-coupled offerings--a shift that strains the tree-based classifiers traditionally preferred for sparse and high...

📖 Read original article


187. PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment ​

Author: Ziqi Cui, Shangyu Lou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR

arXiv:2608.24133v1 Announce Type: cross Abstract: People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perceived. Existing geospatial retrieval remains largely POIcentric and metadata-driven, making it difficult to sat...

📖 Read original article


188. Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection ​

Author: Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen, Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Hyung-Min Jeon, Chi Dai Tran, Son Hong Phan, Duong Khac Vu, Trinh Le Ba Khanh, Jae Wook Jeon
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24154v1 Announce Type: cross Abstract: Real-world deployment of traffic surveillance systems is bottlenecked by geographic domain shift, in which models trained in one city underperform when applied to an unseen target city. Conventional domain adaptation relies on hyperparameter-sensitiv...

📖 Read original article


189. LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes ​

Author: Youcheng Zong, Runda Jia, Dakuo He
Published: 8/26/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.ET, cs.SY

arXiv:2608.24156v1 Announce Type: cross Abstract: Industrial actor--critic methods usually represent continuous actions as anonymous numerical coordinates. They must therefore learn from limited interactions which process variables each action affects, in which direction, and after what delay. Fixed...

📖 Read original article


190. Preference Optimization for Non-Verbal Vocalization Synthesis ​

Author: Haoyang Li, Chenglin Xu, Junchuan Zhao, Yuang Cao, Liumeng Xue, Yiwen Guo, Eng Siong Chng
Published: 8/26/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.LG

arXiv:2608.24163v1 Announce Type: cross Abstract: Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for N...

📖 Read original article


191. Tlow: Flow-based Item Tokenizer for Recommendation ​

Author: Nian Li, Chonggang Song, Jingtao Ding, Lingling Yi, Yong Li, Qingmin Liao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.24176v1 Announce Type: cross Abstract: Item tokenizer encodes semantic embeddings into token IDs to replace the randomly assigned item IDs used in traditional recommendation models, fundamentally addressing the problems of excessive parameters and cold starts. However, the most common tok...

📖 Read original article


192. 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection ​

Author: Fawzia Zehra (Fuzzy), Kara-Isitt, Sonal Khosla, Stephen Swift
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24191v1 Announce Type: cross Abstract: Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for conte...

📖 Read original article


193. Contrastive Branch Policy Optimization ​

Author: Ying Wang, Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun, Jingli Yang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24300v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for succe...

📖 Read original article


194. SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling ​

Author: Shahed Masoudian, Markus Frohmann, Emmanouil Karystinaios, Navid Rekabsaz, Markus Schedl
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24304v1 Announce Type: cross Abstract: Recent controllable text generation (CTG) for sentiment control has largely focused on decoder-based large language models, making causal attention the dominant paradigm. While effective for fluent generation, these models still struggle to satisfy c...

📖 Read original article


195. Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning ​

Author: Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.HC, cs.LG

arXiv:2608.24340v1 Announce Type: cross Abstract: The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together dif...

📖 Read original article


196. Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis ​

Author: Marc Rodr'iguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24342v1 Announce Type: cross Abstract: Synthetic image generation is a promising strategy to address data scarcity and the underrepresentation of clinically important phenotypes in medical imaging, yet generating images that faithfully reflect meaningful patient characteristics remains ch...

📖 Read original article


197. FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision ​

Author: Qiming Xie, Wenjie Zheng, Xiangqing Shen, Rui Xia
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24350v1 Announce Type: cross Abstract: To reduce the hallucination risk caused by outcome-driven rewards in large language models trained through reinforcement learning with verifiable rewards, existing mitigation approaches introduce process-level factual supervision. However, due to coa...

📖 Read original article


198. Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs ​

Author: Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang, Jun Sun, Guoheng Sun, Xiaoning Ren, Haijun Wang, Ting Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2608.24354v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for...

📖 Read original article


199. Markerless Pose Estimation for Resistance Training Technique Assessment ​

Author: Joseph Turner, Jeff Clark, Nawid Keshtmand
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24384v1 Announce Type: cross Abstract: Resistance training can be a high risk activity, and safe form is essential to avoiding injury. Laboratory-based movement analysis provides quantitive technique assessment, yet is not easily accessible. Markerless pose estimation infers body landmark...

📖 Read original article


200. Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning ​

Author: Ruihan Liu, Yu Ji, Jianbo Yu, Shifu Yan, Qingchao Jiang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24386v1 Announce Type: cross Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open challenge. While E(3)-equivariant neural networks excel at point estimates, they lack rigorous confidence measure...

📖 Read original article


201. Multilevel Fair Allocation under Additive Preferences ​

Author: Maxime Lucet, Nawal Benabbou, Aur'elie Beynier, Nicolas Maudet
Published: 8/26/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA

arXiv:2608.24400v1 Announce Type: cross Abstract: We study multilevel fair resource allocation with tree-structured hierarchical relations among agents. At each level, the problem can be viewed locally as allocating an agent's bundle to its children, the overall allocation being a trace of this proc...

📖 Read original article


202. Evaluating Deep Multivariate Imputation Models on Wearable Device Data ​

Author: Skye Goodman, Roussel Desmond Nzoyem, Leandro Junges, Peter Kissack, Yasser Qureshi, Amberly Brigden, Jeff Clark, Nawid Keshtmand
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.QM

arXiv:2608.24436v1 Announce Type: cross Abstract: Wearable device data enables continuous health monitoring, but suffers from structured missingness: features sharing a physical sensor drop out together. Deep imputation methods such as BRITS and SAITS have seen limited evaluation on multimodal physi...

📖 Read original article


203. Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning ​

Author: Hang Chen, Jiaying Zhu, Wenya Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.24482v1 Announce Type: cross Abstract: Mechanistic Localization bridges mechanistic interpretability and post-training optimization by isolating critical parameters via interpretative approaches and then guiding parameter-efficient Supervised Fine-Tuning (SFT) in a ``locating-then-tuning'...

📖 Read original article


204. When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study ​

Author: Mohit Singh Chauhan, Vipin Gyanchandani, Dylan Bouchard
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.24492v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time. Prior work has proposed combining UQ signals via...

📖 Read original article


205. Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR ​

Author: Janis Aiad, Aghiles Drali, Aymen El Ouadrhiri, Anass Ettahiri, Yasser Oufqir, Simon Patry, David Cortes, Marianne Clausel, Emilie Devijver
Published: 8/26/2026, 4:00:00 AM
Categories: stat.ML, cs.AI

arXiv:2608.24500v1 Announce Type: cross Abstract: The STAR (Student-Teacher Achievement Ratio) experiment (1985, Tennessee, USA) is a landmark hierarchical dataset designed to assess the impact of class size on student outcomes, with observations nested within classes. To encode class-level interven...

📖 Read original article


206. LumiXAI: A Modular Full-Stack Framework for Feature Attribution ​

Author: Alfio Ferrara, Lorenzo Gatta, Sergio Picascia, Elisabetta Rocchetti
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.24524v1 Announce Type: cross Abstract: Feature attribution is a central tool of model interpretability, yet the software through which it is applied remains fragmented: individual tools specialize along narrow axes, such as a single modality, a code API or a GUI, or a fixed rather than ex...

📖 Read original article


207. FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment ​

Author: Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24551v1 Announce Type: cross Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymm...

📖 Read original article


208. StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment ​

Author: Wentao Yang, Zhenye Xu, Ruoyi Li, Musen Zhang, Yao Guo
Published: 8/26/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.MA

arXiv:2608.24555v1 Announce Type: cross Abstract: Prehospital stroke assessment aims to accurately identify stroke symptoms and make rapid decisions through standardized procedures within an extremely narrow time window, thereby saving valuable time for subsequent treatment. In clinical practice, FA...

📖 Read original article


209. COCI: Conference Organisers and Content Identifier ​

Author: Angelo Salatino, Francesco Osborne, Alexis Vizcaino, Aliaksandr Birukou, Enrico Motta
Published: 8/26/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2608.24559v1 Announce Type: cross Abstract: Despite the critical role of grey literature in scholarly communication, artefacts such as Calls for Papers (CfPs) remain largely isolated from modern Scholarly Knowledge Graphs. The unstructured and highly heterogeneous nature of these documents has...

📖 Read original article


210. Across the Loss Landscape with Progressive Growth ​

Author: Paul Caillon, Christophe Cerisara, Alexandre Allauzen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24568v1 Announce Type: cross Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geometry of the minima found by stochastic optimization. We study how incremental grow-and-optimize strateg...

📖 Read original article


211. $\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions ​

Author: Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook
Published: 8/26/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, q-fin.RM

arXiv:2608.24582v1 Announce Type: cross Abstract: Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear str...

📖 Read original article


212. Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks ​

Author: Yulong Dou, Han Wu, Guo Chen, Fangmao Ju, Zhiming Cui, Dinggang Shen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24597v1 Announce Type: cross Abstract: Electroencephalography (EEG) is a widely used window into human brain function, but most EEG models remain tied to a one-dataset-one-model supervised paradigm. Recent EEG foundation models offer a route toward reusable representations, but most remai...

📖 Read original article


213. A Literate Programming Environment for Human and Machine Agents ​

Author: Adam T. Burke
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.PL

arXiv:2608.24644v1 Announce Type: cross Abstract: This paper introduces an environment for constructing literate programs in concert with language-aware machine agents. This environment includes a grammar for executable program essays, a parser that treats names as first-class objects, an internal n...

📖 Read original article


214. Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems ​

Author: Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2608.24650v1 Announce Type: cross Abstract: System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain costly and often infeasible. However, modern LLM serving now evolves faster than human-driven simulato...

📖 Read original article


215. Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration ​

Author: Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.DC, cs.ET, cs.LG

arXiv:2608.24664v1 Announce Type: cross Abstract: We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures...

📖 Read original article


216. On-policy Distillation with Verifiable Reward ​

Author: Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24696v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense t...

📖 Read original article


217. Constrained Hyperparameter Optimization for Streaming Data ​

Author: Bruno Veloso, Jo~ao Gama
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24712v1 Announce Type: cross Abstract: Optimization of hyperparameters is a critical factor to obtain optimal model performance. While existing research has predominantly concentrated on batch-learning scenarios, addressing the complexities inherent in data streams presents a challenge. T...

📖 Read original article


218. Deep Learning Super Resolution for Satellite Cloud Mask Downscaling ​

Author: Angelos Georgakis, Valentina Kanaki, Giorgos Giannopoulos, Stella Girtsou, Ioannis Kontogiorgakis, Charalampos Kontoes, Kostas Philippopoulos
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24715v1 Announce Type: cross Abstract: A vast amount of optical satellite data is being transmitted to Earth-based servers every day, and more than half of this data is affected by haze or clouds. Additionally, this data suffers from the fundamental trade-off between spatial and temporal ...

📖 Read original article


219. Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity ​

Author: Heng Zhang, Haotian Xiang, Qin Lu, Konstantinos D. Polyzos, Tara Javidi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24721v1 Announce Type: cross Abstract: Hyperparameter selection remains a key challenge in Bayesian optimization (BO) and Bayesian active learning (AL), as model misspecification can lead to suboptimal performance, while more accurate fully Bayesian treatments typically rely on computatio...

📖 Read original article


220. Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets ​

Author: Meghal Dani, Stefanie Liebe
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24727v1 Announce Type: cross Abstract: EEG foundation models pretrained via self-supervised learning promise transferable representations, but their generalization remains limited, especially across diverse clinical datasets. Full fine-tuning is impractical for resource-constrained clinic...

📖 Read original article


221. Method, Mind, and Morality: How People Make Sense of Artificial Intelligence ​

Author: Jacy Reese Anthis, Erik Brynjolfsson, James Evans
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG, stat.ML

arXiv:2608.24748v1 Announce Type: cross Abstract: How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dynamics of AI through an open-ended, mixed-methods study with computational text analysis of millions of AI-related newspaper articles and soc...

📖 Read original article


222. The RAT: A Unified Bayesian Model for RAG Evaluation ​

Author: Pius von D"{a}niken, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24753v1 Announce Type: cross Abstract: Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jo...

📖 Read original article


223. Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations ​

Author: Sohini Gupta, Bahareh Tolooshams
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.24762v1 Announce Type: cross Abstract: Many disentanglement methods represent generative factors using Euclidean product coordinates, although the underlying factor spaces may wrap, collapse, or have position-dependent geometry. We introduce factor-space structure, combining factor domain...

📖 Read original article


224. Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks ​

Author: Weimin Zhou
Published: 8/26/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG, stat.CO

arXiv:2608.24768v1 Announce Type: cross Abstract: The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tasks. However, analytical computation of the IO test statistic is generally intractable. Numerical approaches based on Markov-chain Mon...

📖 Read original article


225. Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy ​

Author: Md Shahriar Sajid
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.24771v1 Announce Type: cross Abstract: Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid intervention for survival. Early diagnosis and preventive measures can greatly reduce life loss and disabilities. Recent advancements in...

📖 Read original article


226. Automatic Model Card Generation Using an LLM ​

Author: Tajkia Rahman Toma, Balreet Grewal, Cor-Paul Bezemer
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.24807v1 Announce Type: cross Abstract: Model cards are structured documents that summarize key information about machine learning models to improve transparency, usability, and accountability. However, they often lack a consistent structure, and many models provide no model cards, making ...

📖 Read original article


227. Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows ​

Author: Miao Liu, Zhizhe Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.24842v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information a...

📖 Read original article


228. LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training ​

Author: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd"aus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Sch"olkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.24845v1 Announce Type: cross Abstract: We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is ...

📖 Read original article


229. Fuzzy Segmentations of a String ​

Author: Armen Kostanyan, Arevik Harmandayan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2201.13427v2 Announce Type: replace Abstract: This article discusses a particular case of the data clustering problem, where it is necessary to find groups of adjacent text segments of the appropriate length that match a fuzzy pattern represented as a sequence of fuzzy properties. To solve thi...

📖 Read original article


230. Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives ​

Author: Lening Li, Zhentian Qian, Jianan Xia, Qiren Geng, Huasheng Zhang, Liang Hu, Qishuang Li, Junqiang Lou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, math.OC

arXiv:2304.10041v2 Announce Type: replace Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications expressed in linear temporal logic. To learn an optimal policy that maximizes the satisfaction probability, we compo...

📖 Read original article


231. Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs ​

Author: Shaojie Zhu, Zhaobin Wang, Chengxiang Zhuo, Hui Lu, Bo Hu, Zang Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC

arXiv:2312.17535v2 Announce Type: replace Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led to large language models (LLMs) attracting widespread attention. However, restricted by expensive costs, many studies have to focus on the abilit...

📖 Read original article


232. LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis ​

Author: Lecheng Zheng, Zhengzhang Chen, Dongjie Wang, Chengyuan Deng, Reon Matsuoka, Haifeng Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2406.05375v4 Announce Type: replace Abstract: Root cause analysis (RCA) is crucial for enhancing the reliability and performance of complex systems. However, progress in this field has been hindered by the lack of large-scale, open-source datasets tailored for RCA. To bridge this gap, we intro...

📖 Read original article


233. Efficient LLM Collaboration via Planning ​

Author: Byeongchan Lee, Jonghoon Lee, Dongyoung Kim, Jaehyung Kim, Kyungjoon Park, Dongjun Lee, Jinwoo Shin
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achieve remarkable results across diverse tasks, they often incur substantial monetary inference cost, ma...

📖 Read original article


234. Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light ​

Author: Mani Hamidi, Terrence W. Deacon
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2507.11482v5 Announce Type: replace Abstract: Artificial learning systems are graduating from passive learners to increasingly autonomous agents, lending pragmatic urgency to the question of what constitutes agency. Reinforcement learning (RL) offers arguably the most explicit formulation of a...

📖 Read original article


235. Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning ​

Author: Tiberiu-Andrei Georgescu, Alexander W. Goodall, Dalal Alrajeh, Francesco Belardinelli, Sebastian Uchitel
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.02605v3 Announce Type: replace Abstract: Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent's actions remain compliant with formal specifications. Classical shielding approaches, however, are often static, in the sense that they assume fixed ...

📖 Read original article


236. UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models ​

Author: Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Kun Kuang, Zhongxiang Dai
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.08873v3 Announce Type: replace Abstract: Large language models (LLMs) are shifting from answer providers to intelligent tutors in educational settings, yet current supervised fine-tuning methods only learn surface teaching patterns without dynamic adaptation capabilities. Recent reinforce...

📖 Read original article


237. ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering ​

Author: Ge Yan, Chung-En Sun, Linbo Liu, Tsui-Wei Weng
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.13979v2 Announce Type: replace Abstract: Large reasoning models achieve strong performance on diverse tasks by producing extended chains of thought. Self-reflection, the ability to review and revise prior reasoning steps, is widely regarded as a key contributor to this performance. Howeve...

📖 Read original article


238. Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge ​

Author: Runhao Zhao, Weixin Zeng, Wentao Zhang, Chong Chen, Zhengpin Li, Xiang Zhao, Lei Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.10485v5 Announce Type: replace Abstract: Domain-specific knowledge graphs (DKGs) are critical yet often suffer from limited coverage compared to General Knowledge Graphs (GKGs). Existing tasks to enrich DKGs rely primarily on extracting knowledge from external unstructured data or complet...

📖 Read original article


239. Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models ​

Author: Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini, Luca Pappalardo, Fosca Giannotti, Francesco Giannini
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2602.02304v3 Announce Type: replace Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context learning. Current explainability methods are structurally ill-suited to ...

📖 Read original article


240. CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support ​

Author: Yichen Wu, Kailong Fan, Sangjoon Park, Yuhan Liu, Zhiyi Shi, Sekeun Kim, Dania Daye, Hana Farzaneh, Xiang Li, Raul Uppot, Yujin Oh, Quanzheng Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2602.09159v2 Announce Type: replace Abstract: Recent multi-agent frameworks have shown promise for oncology decision support, yet most assume centralized data access and rely on prompt-based assignment, limiting their applicability in privacy-sensitive clinical settings. We propose Contributio...

📖 Read original article


241. PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools ​

Author: Yusheng Li, Tianjun Feng, Yunfeng Chen, Chun-Yi Tsai, Yihan Sun, Ayan Das, Kaoutar El Maghraoui, Shuxin Lin, Dhaval Patel
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.01532v3 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critical \emph{Prognostics and Health Management (PHM)} is unanswered. Pri...

📖 Read original article


242. Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints ​

Author: Minh-Khoi Pham, Thang-Long Nguyen Ho, Thao Thi Phuong Dao, Tai Tan Mai, Minh-Triet Tran, Marie E. Ward, Una Geary, Rob Brennan, Nick McDonald, Martin Crane, Marija Bezbradica
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.01841v4 Announce Type: replace Abstract: Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbalance, and distribution shift. While tabular in-context learning (TICL) and retrieval-augmented methods perform...

📖 Read original article


243. ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams ​

Author: Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.15994v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However, when faced with complex topological structures involving branching paths, converging flows, and cyclic depe...

📖 Read original article


244. Housing Potential Common Data Model and City Digital Twin ​

Author: Megan Katsumi, Mark Fox, Anderson Wong, Divnoor Chatha
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.05535v2 Announce Type: replace Abstract: The evaluation of housing potential requires consideration of a location from multiple perspectives, ranging from zoning and land use to population characteristics and access to services. This research introduces the Housing Potential Common Data M...

📖 Read original article


245. Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust ​

Author: Shijun Lei, Quang Nguyen, Swapneel S Mehta, Zeping Li, Huichuan Fu, Xiaolong Zheng, Siki Chen, Yunji Liang, Philip Torr, Zhenfei Yin
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.10059v3 Announce Type: replace Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social and economic simulation. While prior work has discovered strategic deception by LLM agents in f...

📖 Read original article


246. EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design ​

Author: Gioele Molinari, Florian Felten, Soheyl Massoudi, Mark Fuge
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2605.19743v3 Announce Type: replace Abstract: Engineering-agent systems are proliferating, but differences in tasks, tools, and success criteria make demonstrations difficult to compare and failures difficult to diagnose. We introduce a capability-based evaluation framework for tool-connected ...

📖 Read original article


247. Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control ​

Author: Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, physics.flu-dyn

arXiv:2606.08405v3 Announce Type: replace Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific control design in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control a...

📖 Read original article


248. Atomic Units of X: The Compression Layer of Intelligence ​

Author: Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.12634v2 Announce Type: replace Abstract: This paper proposes a theoretical and empirical framework for understanding intelligence as a process of atomic compression and compositional reuse. It argues that scalable cognitive, biological, computational, and organisational systems reduce com...

📖 Read original article


249. SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents ​

Author: Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.08253v2 Announce Type: replace Abstract: We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents, unifying multi-channel retrieval under reciprocal-rank fusion, bi-temporal recall, multi-scope isolation, role-based access, verified erasure, and a has...

📖 Read original article


250. Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models ​

Author: Kevin Murphy
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09696v4 Announce Type: replace Abstract: A primary goal of science is to learn mechanistic or causal world models from data. These models can be used to explain some phenomenon of interest. They also provide the ability to answer interventional ``what if'' questions (i.e., to predict the ...

📖 Read original article


251. The Dynamics of Intelligence Explosions ​

Author: Toby Ord
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, econ.TH

arXiv:2608.14426v2 Announce Type: replace Abstract: AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be able to produce an intelligence explosion, with rapidly escalating AI capabilities. I explore the mathematics of the most explosive possibilitie...

📖 Read original article


252. Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument ​

Author: Miko{\l}aj Sienicki, Krzysztof Sienicki
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.GL, quant-ph

arXiv:2608.14673v2 Announce Type: replace Abstract: We present an independent human assessment of the proof developed in Chapter 6 of OpenAI's Ten Advances in Mathematics and Theoretical Computer Science. An initial audit appeared to identify a polarity error in a greedy conditioning lemma. Subseque...

📖 Read original article


253. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies ​

Author: Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Navan Preet Singh, Qingqing Mao, Ritankar Das
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2608.16645v3 Announce Type: replace Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous o...

📖 Read original article


254. ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting ​

Author: Rui Wang, Yeteng Wu, Xianlin Zhang, Mengshi Qi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.20009v2 Announce Type: replace Abstract: Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion. However, existing benchmarks rarely expose object-level physical properties as ex...

📖 Read original article


255. What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies ​

Author: Narcis Marincat
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2608.20054v3 Announce Type: replace Abstract: Multi-module neural systems often expose every module to the full input. We test whether a slot-selective evidence-masking regime -- restricting each module to its own evidence span -- changes which solutions gradient-based training discovers. Four...

📖 Read original article


256. SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing ​

Author: Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22018v2 Announce Type: replace Abstract: Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, supporting arguments, and target-level labels. Documents with multiple targets, conflicting local readings, or cul...

📖 Read original article


257. GenCoord: Skill-Path Commitments under Private Information ​

Author: Peng He, Junning Zhu, Haohan Yuan, Jianpeng Liang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22055v2 Announce Type: replace Abstract: Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither local view determines who should act, what should be handed off, or how the joint task should continue. We in...

📖 Read original article


258. Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models ​

Author: Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.MM

arXiv:2608.22232v2 Announce Type: replace Abstract: Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and ...

📖 Read original article


259. CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents ​

Author: Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22577v2 Announce Type: replace Abstract: Long-horizon GUI agents can retain complete action histories as compact text, but only a few historical screenshots fit in active context. We formulate this as budgeted fidelity restoration: every event remains summarized, while a fixed budget $B$ ...

📖 Read original article


260. SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems ​

Author: Xiang Wang, Shigang Quan, Tingzhen Chang, Kang Yang, Sitong Chen, Yabo Fan, Xingxing Wang, Zhaodian He
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.22979v2 Announce Type: replace Abstract: Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quantization is compact but introduces boundary distortion, whereas dense soft quantization couples representatio...

📖 Read original article


261. MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks ​

Author: Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao, Long Chen, Weigao Sun, Feida Zhu, Yiran Zhong, Steven Hoi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23035v2 Announce Type: replace Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind ...

📖 Read original article


262. Apodex 1.1: Scaling Agentic Intelligence for Complex Work ​

Author: B. An, B. Li, B. Wang, B. Zhang, B. L. Wang, C. Feng, C. Wei, C. Xue, C. Zhang, D. Ng, D. Ye, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, G. Sun, H. Ji, H. Xu, H. Yang, H. Ye, H. Zhang, H. Zhao, J. Li, J. Lin, J. Xia, K. Jin, K. Wang, K. Yang, L. Bing, L. Lei, L. Su, Le. Wang, Lu. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du, S. Li, S. Lin, S. Nie, S. Wang, S. Zhang, S. Z. Wang, T. Ge, Ta. Q. Fang, Ti. Q. Fang, W. Fang, W. Li, W. Zhang, X. Chen, X. Li, X. Tang, X. Wang, X. Xu, X. Zhang, X. Q. Wang, X. Y. Wang, Y. Deng, Y. Gao, Y. Hu, Y. Li, Y. Sui, Y. Wang, Y. Xiao, Y. Zhang, Y. Zhou, Z. Chen, Z. Cheng, Z. Feng, Z. Liang, Z. Liu, Z. Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.23283v2 Announce Type: replace Abstract: General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delive...

📖 Read original article


263. MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction ​

Author: Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.23397v2 Announce Type: replace Abstract: Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existing agents struggle to convert experience into reusable process knowl...

📖 Read original article


264. Screening Autism Spectrum Disorder in children using Deep Learning Approach : Evaluating the classification model of YOLOv26s by comparing with other models ​

Author: Subash Gautam, Sagar Pathak, Prabin Sharma, Bidhya Shrestha, Kisan Thapa, Shubham Joshi, Mala Deep Upadhaya, Dikshya Thapa, Chandiprasad Chintalapati, Sagar Duwal, Angela Upreti, Salik Ram Khanal
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2306.14300v2 Announce Type: replace-cross Abstract: Autism spectrum disorder (ASD) is a developmental condition that presents significant challenges in social interac- tion, communication, and behavior. Early intervention plays a pivotal role in enhancing cognitive abilities and reducing autis...

📖 Read original article


265. HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA ​

Author: Xinyue Chen, Pengyu Gao, Jiangjiang Song, Xinjian Chen, Xiaoyang Tan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2402.01767v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by integrating external documents during generation. However, retrieval accuracy can degrade when the knowledge base contains many semantically and ...

📖 Read original article


266. Intrinsic PAPR: Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering ​

Author: Alireza Moazeni, Shichong Peng, Yanshu Zhang, Chirag Vashist, Ke Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR, cs.LG

arXiv:2407.00500v2 Announce Type: replace-cross Abstract: Recent point-based intrinsic decomposition and inverse rendering methods have advanced the modelling of the shading and albedo of 3D scenes. However, we identify a fundamental limitation: these methods suffer from a misattribution issue, wher...

📖 Read original article


267. Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control ​

Author: Yaron Veksler, Sharon Hornstein, Han Wang, Maria Laura Delle Monache, Daniel Urieli
Published: 8/26/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2412.02520v4 Announce Type: replace-cross Abstract: Connected automated vehicles (CAVs) equipped with adaptive cruise control (ACC) create new opportunities for highway congestion mitigation. Traditional practice relies on Eulerian variable speed limits (VSL) which regulate traffic through roa...

📖 Read original article


268. Generative AI for Validating Physics Laws ​

Author: Maria Nareklishvili, Nicholas Polson, Vadim Sokolov
Published: 8/26/2026, 4:00:00 AM
Categories: astro-ph.SR, astro-ph.GA, cs.AI

arXiv:2503.17894v3 Announce Type: replace-cross Abstract: We propose generative learner for estimating heterogeneous treatment effects and characterizing the full distribution of causal effects. The learner takes the form of a multi-head feed-forward neural network with three jointly estimated subne...

📖 Read original article


269. Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review ​

Author: Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Leon Witt, Muhammad Asif Ali, Yukai Miao, Dan Li, Qingsong Wei
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2504.18346v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect information, remains one of the leading challenges for LLMs. This raises the question of how to accurate...

📖 Read original article


270. Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric ​

Author: Jatin Kumar Arora, Soutrik Bandyopadhyay, Sunil Sulania, Shubhendu Bhasin
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2505.23197v5 Announce Type: replace-cross Abstract: Path planning for autonomous robots faces a fundamental trade-off between path length and obstacle clearance. While existing algorithms typically prioritize a single objective, we introduce the Unified Path Planner (UPP), a graph-search algor...

📖 Read original article


271. Quasar: A Programming Language Specialized for LLM Code Actions ​

Author: Stephen Mell, Botong Zhang, David Mell, Shuo Li, Ramya Ramalingam, Nathan Yu, Stephan Zdancewic, Osbert Bastani
Published: 8/26/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CR, cs.LG

arXiv:2506.12202v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often call external tools to solve tasks. One effective strategy is for LLMs to write code, enabling them to use complex control flow such as conditionals and loops. Such code actions are typically represented as ...

📖 Read original article


272. From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs ​

Author: Anh Ho, Thanh Le-Cong, Bach Le, Christine Rizkallah
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2506.13182v3 Announce Type: replace-cross Abstract: [...] Since then, various APR approaches, especially those leveraging the power of large language models (LLMs), have been rapidly developed to fix general software bugs. Unfortunately, the effectiveness of these advanced techniques in the co...

📖 Read original article


273. A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs ​

Author: Kethmi Hirushini Hettige, Jiahao Ji, Cheng Long, Shili Xiang, Gao Cong, Jingyuan Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2506.20073v3 Announce Type: replace-cross Abstract: Spatio-temporal data mining plays a pivotal role in informed decision making across diverse domains. However, existing models are often restricted to narrow tasks, lacking the capacity for multi-task inference and complex long-form reasoning ...

📖 Read original article


274. Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities ​

Author: Georges Sfeir, Gabriel Nova, Stephane Hess, Sander van Cranenburgh
Published: 8/26/2026, 4:00:00 AM
Categories: econ.EM, cs.AI

arXiv:2507.21790v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potential of LLMs as assis...

📖 Read original article


275. NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs ​

Author: Birong Pan, Jianhao Chen, Mayi Xu, Qiankun Pi, Yuanyuan Zhu, Ming Zhong, Tieyun Qian
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2508.09473v2 Announce Type: replace-cross Abstract: Ensuring robust safety alignment while preserving utility is critical for the reliable deployment of Large Language Models (LLMs). However, current techniques fundamentally suffer from intertwined deficiencies: insufficient robustness against...

📖 Read original article


276. An Information-Flow Perspective on Explainability Requirements: Specification and Verification ​

Author: Bernd Finkbeiner, Hadar Frenkel, Julian Siber
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2509.01479v4 Announce Type: replace-cross Abstract: Explainable systems expose information about why certain observed effects are happening to the agents interacting with them. We argue that this constitutes a positive flow of information that needs to be specified, verified, and balanced agai...

📖 Read original article


277. STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification ​

Author: Zongsen Qiu, Jianjun Wang, Yue Zhou, Zibo Zhou, Rui Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.03754v2 Announce Type: replace-cross Abstract: Responding to rising global food security needs, precision agriculture and deep learning-based plant disease diagnosis have become crucial. Yet, deploying high-precision models on edge devices is challenging. Most lightweight networks use att...

📖 Read original article


278. Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships ​

Author: Zhuoyue Zhang, Haitong Xu, Carlos Guedes Soares
Published: 8/26/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2509.15959v2 Announce Type: replace-cross Abstract: Autonomous navigation in maritime domains is accelerating alongside advances in artificial intelligence, sensing, and connectivity. Opaque decision-making and poorly calibrated human-automation interaction remain key barriers to safe adoption...

📖 Read original article


279. VGGT-DP: Generalizable Robot Control via Vision Foundation Models ​

Author: Shijia Ge, Yijun Liu, Yinxin Zhang, Shuzhao Xie, Weixiang Zhang, Mingcai Zhou, Zhi Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2509.18778v2 Announce Type: replace-cross Abstract: Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly focus on policy design, they often neglect the structure and capacity of visual encoders, limiting spa...

📖 Read original article


280. Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics? ​

Author: Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2510.14249v2 Announce Type: replace-cross Abstract: Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval, text-guided music generation, and audio captioning. Central to these tasks are joint language-audio emb...

📖 Read original article


281. Monotone and Separable Set Functions: Characterizations and Neural Models ​

Author: Soutrik Sarangi, Yonatan Sverdlov, Nadav Dym, Abir De
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.23634v5 Announce Type: replace-cross Abstract: Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector functions so that the natural partial order on sets is preserved, namely $S\subseteq T \text{ if and only if } ...

📖 Read original article


282. CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution ​

Author: Christian Schiffer, Zeynep Boztoprak, Jan-Oliver Kropp, Julia Th"onni{\ss}en, Katia Berr, Hannah Spitzer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid
Published: 8/26/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG

arXiv:2511.01870v3 Announce Type: replace-cross Abstract: Studying the cellular architecture of the human cerebral cortex is essential for understanding how the brain is organized from the micro to the macro level, and how it functions. However, investigating complex texture patterns in histological...

📖 Read original article


283. Robust Motion Generation using Part-level Reliable Data from Videos ​

Author: Boyuan Li, Sipeng Zheng, Bin Cao, Ruihua Song, Zongqing Lu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2512.12703v2 Announce Type: replace-cross Abstract: Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However, some human parts in many video frames cannot be seen due to off-screen captures or occlusions. It bring...

📖 Read original article


284. Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library ​

Author: Oliver Stritzel, Nick H"uhnerbein, Simon Rauch, Itzel Zarate, Lukas Fleischmann, Moike Buck, Attila Lischka, Christian Frey
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2512.16715v3 Announce Type: replace-cross Abstract: In recent years, Predictive Process Mining (PPM) techniques based on artificial neural networks have evolved as a method for monitoring the future behavior of unfolding business processes and predicting Key Performance Indicators (KPIs). Howe...

📖 Read original article


285. Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations ​

Author: Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M"oller, Vera Schmitt
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2601.00282v2 Announce Type: replace-cross Abstract: Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs, generated by LLMs to justify their own outputs, require reas...

📖 Read original article


286. Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes ​

Author: Chen Ling, Tongwei Zhang, Hanqian Li, Nai Ding
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.07737v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability to process action scenes that contradict everyday common sense remains undertested. To address this ...

📖 Read original article


287. Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model ​

Author: Song-Ju Kim
Published: 8/26/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, q-bio.NC

arXiv:2601.10034v3 Announce Type: replace-cross Abstract: Decision making often exhibits context dependence that is difficult to accommodate within a single non-invasive classical probability model. This paper develops a quantum-like extension of the Tug-of-War (QTOW) decision-making model to ask wh...

📖 Read original article


288. TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning ​

Author: Daixian Liu, Jiayi Kuang, Yinghui Li, Yangning Li, Di Yin, Haoyu Cao, Xing Sun, Ying Shen, Hai-Tao Zheng, Liang Lin, Philip S. Yu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2601.16520v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding, yet precise compositional spatial reasoning under geometric constraints remains underexplored. Existing benchmarks ma...

📖 Read original article


289. Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility ​

Author: Honglin Lin, Zheng Liu, Chonghan Qin, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.17027v2 Announce Type: replace-cross Abstract: While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models of...

📖 Read original article


290. Ad Insertion in LLM-Generated Responses ​

Author: Shengwei Xu, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang, Grant Schoenebeck
Published: 8/26/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CL

arXiv:2601.19435v2 Announce Type: replace-cross Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fleeting, context-dependent user intent---the specific informati...

📖 Read original article


291. Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging ​

Author: Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML

arXiv:2602.03702v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained in continual or open-ended settings, where the total training horizon is not known in advance. Despite this, most existing pretraining recipes are not anytime: they rely on horizon-dependent lear...

📖 Read original article


292. ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents ​

Author: Cong Pang, Xuyu Feng, Yujie Yi, Jiaqi Su, Zixuan Chen, Jiawei Hong, Tiankuo Yao, Nang Yuan, Jiapeng Luo, Lewei Lu, Xin Lou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.10863v2 Announce Type: replace-cross Abstract: Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer is correct, but not which acquired information enabled it. This difficulty is amplified by text-deri...

📖 Read original article


293. PatientHub: A Unified Framework for Patient Simulation ​

Author: Sahand Sabour, TszYam NG, Minlie Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2602.11684v2 Announce Type: replace-cross Abstract: As Large Language Models increasingly power role-playing applications, simulating patients has become a valuable tool for training counselors and scaling therapeutic assessment. However, prior work remains fragmented: existing approaches rely...

📖 Read original article


294. You Can Learn Tokenization End-to-End with Reinforcement Learning ​

Author: Sam Dauncey, Roger Wattenhofer
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.13940v3 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend towards architectures becoming increasingly end-to-end. Prior work has shown promising results at sca...

📖 Read original article


295. VLANeXt: Recipes for Building Strong VLA Models ​

Author: Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang, Runze Yang, Yihang Luo, Zhonghua Wu, Wei-Shi Zheng, Chen Change Loy
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2602.18532v3 Announce Type: replace-cross Abstract: Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape r...

📖 Read original article


296. ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents ​

Author: Bowen Zhou, Zhou Xu, Wanli Li, Jingyu Xiao, Pingan Gan, Haoqian Wang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2603.00188v2 Announce Type: replace-cross Abstract: Training-free KV cache compression is essential for deploying vision-language GUI agents under memory and latency constraints, yet existing methods are designed for generic language workloads and ignore the distinctive structure of GUI intera...

📖 Read original article


297. EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training ​

Author: Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fi\v{s}el, Tanel Alum"ae, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.02041v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are predominantly trained on English-centric data, resulting in uneven performance for smaller languages. We study whether continued pretraining (CPT) can improve Estonian capabilities in multilingual LLMs while p...

📖 Read original article


298. ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models ​

Author: Harry Owiredu-Ashley
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2603.10068v2 Announce Type: replace-cross Abstract: Most adversarial evaluations of large language model (LLM) safety assess single prompts and report binary pass/fail outcomes, which fails to capture how safety properties evolve under sustained adversarial interaction. We present ADVERSA, an ...

📖 Read original article


299. msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models ​

Author: Subina Khanal, Seshu Tirupathi, Merim Dzaferagic, Marco Ruffini, Torben Bach Pedersen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.16497v3 Announce Type: replace-cross Abstract: Time series foundation models (TSFMs) require diverse, real-world datasets to adapt across varying domains and temporal frequencies. However, current large-scale datasets predominantly focus on low-frequency time series with sampling interval...

📖 Read original article


300. Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models ​

Author: Xiaojie Gu, Sherry T. Tong, Aosong Feng, Sophia Simeng Han, Jinghui Lu, Yingjian Chen, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Rex Ying, Irene Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce O...

📖 Read original article


301. Beyond OAuth: Task-Scoped Authorization for AI Agents via Natural Language Slices ​

Author: Reshabh K Sharma, Linxi Jiang, Shuo Chen, Zhiqiang Lin
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.PL

arXiv:2603.17170v2 Announce Type: replace-cross Abstract: AI agents increasingly execute users' natural-language (NL) tasks by calling Web services, yet today's Web authorizes these calls through OAuth, which grants permissions over operators (e.g., TRANSFER), not operations (operator plus operands,...

📖 Read original article


302. Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification ​

Author: Giampaolo Bovenzi, Domenico Ciuonzo, Jonatan Krolikowski, Antonio Montieri, Alfredo Nascita, Antonio Pescap`e, Dario Rossi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.LG

arXiv:2603.25507v2 Announce Type: replace-cross Abstract: Network Traffic Classification (NTC) increasingly relies on data-driven models, yet its practical deployment is often constrained by limited labeled data, strict privacy requirements, and the cost of collecting representative traffic traces. ...

📖 Read original article


303. Ollivier-Ricci Curvature of Riemannian Manifolds and Directed Graphs with Applications to Graph Neural Networks ​

Author: Eleanor P Wiesler
Published: 8/26/2026, 4:00:00 AM
Categories: math.DG, cs.AI, cs.SI, math.CO

arXiv:2604.14211v2 Announce Type: replace-cross Abstract: This thesis is an exposition of Ollivier-Ricci Curvature of metric spaces as introduced by Yann Ollivier, which is based upon the 1-Wasserstein Distance and optimal transport theory. We present some of the major results and proofs that connec...

📖 Read original article


304. Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts ​

Author: Gabriel Jason Lee, Jathurshan Pradeepkumar, Jimeng Sun
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2604.16926v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models have shown strong potential for learning generalizable representations from large-scale neural data, yet their clinical deployment is hindered by distribution shifts across clinical settings, dev...

📖 Read original article


305. RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction ​

Author: Md Shifatul Ahsan Apurba, Md Selim, Jin Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.00901v2 Announce Type: replace-cross Abstract: The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differences in imaging protocols and scanner models, CT images acquired by different means may show large diff...

📖 Read original article


306. Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters ​

Author: Lingxiao Kong, Cong Yang, Oya Deniz Beyan, Zeyd Boukhers
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2605.02867v3 Announce Type: replace-cross Abstract: Despite significant advances in Reinforcement Learning (RL), model performance remains highly sensitive to algorithm and hyperparameter configurations, while generalization gaps across environments complicate real-world deployment. Although p...

📖 Read original article


307. Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval ​

Author: Zeyu Yang, Xu Han, Qi Ma, Jason Chen, Anshumali Shrivastava
Published: 8/26/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2605.06647v3 Announce Type: replace-cross Abstract: Retrieval-augmented agents are increasingly the interface to large knowledge bases, yet most treat retrieval as a black box: they issue exploratory queries, inspect snippets, and reformulate until evidence emerges. This resembles how a newcom...

📖 Read original article


308. Outlier-Robust Diffusion Solvers for Inverse Problems ​

Author: Yang Zheng, Jiahua Liu, Tongyao Pang, Wen Li, Zhaoqiang Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.09477v2 Announce Type: replace-cross Abstract: Methods based on diffusion models (DMs) for solving inverse problems (IPs) have recently achieved remarkable performance. However, DM-based methods typically struggle against outliers, which are common in real-world measurements. In this work...

📖 Read original article


309. CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving ​

Author: Minqing Huang, Yujiao Xiang, Zihan Liang, Jiajie Huang, Jingqi Wang, Yuheng Zhou, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, Gong Che
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.10426v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thou...

📖 Read original article


310. ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching ​

Author: Shuoheng Zhang, Yifu Yuan, Hongyao Tang, Yan Zheng, Qiaojun Yu, Pengyi Li, Guowei Huang, Helong Huang, Xingyue Quan, Jianye Hao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2605.11048v2 Announce Type: replace-cross Abstract: Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remain a significant challenge due to complex contact dynamics that demand high-precision force...

📖 Read original article


311. Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation ​

Author: Zixuan Yang, Yiqun Chen, Wei Yang, Erhan Zhang, Zihan Shen, Xiaochi Wei, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.26958v2 Announce Type: replace-cross Abstract: Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailable. Existing rubric-based methods typically rely on pointwise LLM-as-a-judge scoring, but abs...

📖 Read original article


312. Skill-Conditioned Gated Self-Distillation for LLM Reasoning ​

Author: Jiazhen Huang, Xiao Chen, Xiao Luo, Yong Dai, Senkang Hu, Yuzhi Zhao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.28791v2 Announce Type: replace-cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision. Existing methods usually assume trusted PI, such as reference answe...

📖 Read original article


313. A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink ​

Author: Yuhang Jiang, Bowen Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.00930v2 Announce Type: replace-cross Abstract: Mechanistic interpretability routinely reads a probe and labels its top-activating units as the circuit executing the computation. We test the move in Mamba, on the state sink: the selective state-space analogue of the Transformer attention s...

📖 Read original article


314. SaliMory: Orchestrating Cognitive Memory for Conversational Agents ​

Author: Kai Zhang, Xinyuan Zhang, Hongda Jiang, Shiun-Zu Kuo, Hyokun Yun, Ejaz Ahmed, Shereen Oraby, Ziyun Li, Sanat Sharma, Ann Lee, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Xin Luna Dong
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.04120v2 Announce Type: replace-cross Abstract: Conversational agents that serve as lifelong companions must maintain persistent memory across all interactions. However, simply expanding context windows with raw retrieval degrades reasoning quality, while training memory agents via standar...

📖 Read original article


315. When Can One Neuron Fix Repetition Loops in LLMs? ​

Author: Aristotelis Lazaridis, Aman Sharma, Dylan Bates, Brian King, Vincent Lu, Jack FitzGerald
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.13705v2 Announce Type: replace-cross Abstract: The Gemma 4 instruction-tuned models share a reproducible failure: on long factual enumeration prompts, such as TV episodes, the 88 IAU constellations, or the 151 original Pokemon, they collapse into repetition, either a tight verbatim loop o...

📖 Read original article


Author: Fei-Yueh Chen, Chun Huang Lin, Chan Wei Hsu, Kuan Hsuan Yeh, Zih-Ching Chen, Kuan-Ming Chen, Patrick Chung-Chia Huang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2606.18699v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal reasoning remains underexplored. We present TW-LegalBench that utilizes Taiwanese legal system's rich o...

📖 Read original article


317. RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation ​

Author: Pengzhi Yang, Xinyu Wang, Pengyu Jing, Kehan Wen, Yiduo Qu, Zhenhao Huang, Minghao Fu, Xin Liu, Yaheng Shen, Fan Shi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2606.22027v4 Announce Type: replace-cross Abstract: Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poo...

📖 Read original article


318. Co-occurring Associated REtained concepts in Diffusion Unlearning ​

Author: Miso Kim, Georu Lee, Yunji Kim, Hoki Kim, Jinseong Park, Woojin Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2606.24192v2 Announce Type: replace-cross Abstract: Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. As illustrated in Fig.1, unlearni...

📖 Read original article


319. Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop ​

Author: Chenmu Zhang, Boris I. Yakobson
Published: 8/26/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.LG

arXiv:2606.29717v3 Announce Type: replace-cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has produced standard public benchmarks and many published machine-learning models for the task (D...

📖 Read original article


320. A Unified Algebraic Framework for Classification Performance Evaluation ​

Author: Ronaldo C. Prati
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.04028v2 Announce Type: replace-cross Abstract: We propose a unified algebraic framework for classification performance evaluation covering binary, multiclass, multilabel, ordinal, hierarchical, cost-sensitive, and soft-label settings. Actual and predicted labels are represented as binary ...

📖 Read original article


321. Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution ​

Author: Ning Liu, P Aditya Sreekar, Kalle Kujanp"a"a, Zhaoxuan Zhu, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, Tianyu Yang, Shahnawaz Alam, Rose Yu, Baoyuan Liu, Kristina Klinkner, Shervin Malmasi
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08960v2 Announce Type: replace-cross Abstract: Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural co...

📖 Read original article


322. The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle ​

Author: Augusto Camargo
Published: 8/26/2026, 4:00:00 AM
Categories: cs.PF, cs.AI

arXiv:2607.13393v3 Announce Type: replace-cross Abstract: A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself unable to when the incumbent's own output has quietly become the specification. This note offer...

📖 Read original article


323. Discrete Diffusion Models: A Unified Framework from Tokenization to Generation ​

Author: Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Dawn Song, Philip S. Yu, Xue Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.13431v2 Announce Type: replace-cross Abstract: Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffu...

📖 Read original article


324. EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology ​

Author: Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.19261v4 Announce Type: replace-cross Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most pathology benchmarks evaluate models on pre-cropped patches or p...

📖 Read original article


325. GraphVid: Interactive Graph-Controllable Video Generation ​

Author: Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjiao Yu, Adheesh Juvekar, Muntasir Wahed, Ismini Lourentzou
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.21580v2 Announce Type: replace-cross Abstract: Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control...

📖 Read original article


326. LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ​

Author: Junsung Hwang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24555v2 Announce Type: replace-cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: a fixed low-rank sketch shared across pages ...

📖 Read original article


327. MOSAIC: Masked Outsourcing of Secure AI Computations ​

Author: James Hsin-yu Chiang, Sheila Zingg, Kari Kostiainen, Srdjan Capkun
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.29221v2 Announce Type: replace-cross Abstract: We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the s...

📖 Read original article


328. TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction ​

Author: Rasa Hosseinzadeh, Alex Labach, Zexin Xue, Shuyi Han, Valentin Thomas, Anthony L. Caterini
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.01400v2 Announce Type: replace-cross Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. However, recent approaches with either cell-based architectures or retrieval have sacrificed efficiency for raw performance, restricting t...

📖 Read original article


329. Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset ​

Author: Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoru Mo, Yaolong Ju
Published: 8/26/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM

arXiv:2608.06165v3 Announce Type: replace-cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with **kern score ...

📖 Read original article


330. AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization ​

Author: Peng Xu, Chengcheng Wang, Shaohua Wan
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.07557v2 Announce Type: replace-cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containi...

📖 Read original article


331. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol ​

Author: Christoph Trattner
Published: 8/26/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.08882v4 Announce Type: replace-cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own? I call this epistemic transfer. It refers t...

📖 Read original article


332. ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning ​

Author: Harshil Lodhiya
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.14773v2 Announce Type: replace-cross Abstract: The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---has been benchmarked almost entirely on clean data. We show that this choice conceals a large capabilit...

📖 Read original article


333. MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation ​

Author: Lie Li, Wen Li, Junxiao Shen, Guosheng Hu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.15299v2 Announce Type: replace-cross Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented heterogeneity in layer-wise redundancy. We demonstrate that this unifo...

📖 Read original article


334. From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation ​

Author: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Zhengrui Chen, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Yuan Wang, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.18076v2 Announce Type: replace-cross Abstract: Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate e...

📖 Read original article


335. Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas ​

Author: Dmitry V. Alexandrov
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CC, cs.PL

arXiv:2608.18445v3 Announce Type: replace-cross Abstract: We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a combinatorial framework originally motivated by Boolean satisfiability, based on triplet structures and a filter that we call...

📖 Read original article


336. Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay ​

Author: Haiyue Zhang
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.19760v2 Announce Type: replace-cross Abstract: Audited against policy-conditional ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals we audit -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own co...

📖 Read original article


337. ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations ​

Author: Naveen Venkat, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.20539v2 Announce Type: replace-cross Abstract: Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digi...

📖 Read original article


338. Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation ​

Author: Yanglei Gan, Peng He, Run Lin, Peiyuan Jiang, Yifan Wang, Qiao Liu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.20804v2 Announce Type: replace-cross Abstract: Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on...

📖 Read original article


339. Scaling Muon for Diffusion Transformers ​

Author: Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.20818v2 Announce Type: replace-cross Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establish Muon's...

📖 Read original article


340. CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents ​

Author: Jiancheng Wang, Mingli Zhu, Tong Zhang, Jiaqi Ruan, Wei Wang, Siyuan Liang, Dacheng Tao
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21114v2 Announce Type: replace-cross Abstract: Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their perturbations vary sharply over time under a strict per-frame pert...

📖 Read original article


341. Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds ​

Author: Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Daniel Kurzawe, Jan Philip Wahle, Bela Gipp
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.21170v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual p...

📖 Read original article


342. Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores ​

Author: Yu Pan, Hongfeng Yu
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.21829v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation treats the document store as a frozen input, and the offline pipelines that do build structure over it build it unsupervised -- a whole corpus indexed at uniform effort, with no signal about which structure a qu...

📖 Read original article


343. Inferring Action from Future Latent State for Robotic Manipulation ​

Author: Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren
Published: 8/26/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG

arXiv:2608.22067v2 Announce Type: replace-cross Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action m...

📖 Read original article


344. SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models ​

Author: Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao, Ruichu Cai
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22354v2 Announce Type: replace-cross Abstract: Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory but potentially becoming unstable under extreme-context extrapolation. By tracking RWKV-7 over sequences of up to 100M tokens, we empirically identify a...

📖 Read original article


345. Functional compatibility as a determinant of persistent neural learning ​

Author: Hossein Javidnia
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22462v2 Announce Type: replace-cross Abstract: Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behavio...

📖 Read original article


346. The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models ​

Author: Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.22876v2 Announce Type: replace-cross Abstract: Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training o...

📖 Read original article


347. Molecular LLM Agents: From Architectural Design to Scientific Autonomy ​

Author: Jiatong Li, Wengyu Zhang, Weida Wang, Yuxuan Ren, Wei Liu, Chenyang Mao, Yuqiang Li, Yatao Bian, Changmeng Zheng, Xiaoyong Wei, Qing Li
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23104v2 Announce Type: replace-cross Abstract: Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, code, or web environments, molecular LLM agents must perceive, reason about, and act upon chemical objec...

📖 Read original article


348. How Much Regularization Survives Averaging? Update Masking in Federated Learning ​

Author: Wenhao Yan, Fu Kuroda, Yucheng Jin, Zhenke Chen
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.23286v2 Announce Type: replace-cross Abstract: Federated learning on non-IID data seeks flat minima to generalize across clients, and existing methods borrow sharpness-aware minimization from centralized training. There is a second way to reach flat minima, in which the regularization com...

📖 Read original article


349. Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents ​

Author: Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Zeyu Wang, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Hongyi Fu, Jianxiong Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.23329v2 Announce Type: replace-cross Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and D...

📖 Read original article


350. Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data ​

Author: Yifei Song, Kun Efimov-Zhang, Claire Gardent
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.23391v2 Announce Type: replace-cross Abstract: Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different generation tasks. However, most prior work on data-to-text (D2T) generation has focused on specific ta...

📖 Read original article


351. What's the Catch? Evaluating Temporal Consistency in Vision-Language Models ​

Author: Marek Hradil, Danae S'anchez Villegas
Published: 8/26/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2608.23474v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection p...

📖 Read original article


352. Best Practice Critic Optimization ​

Author: Penghui Qi, Xiangxin Zhou, Wee Sun Lee
Published: 8/26/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.23566v2 Announce Type: replace-cross Abstract: Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but s...

📖 Read original article