arXiv cs.AI - 2026-08-17 ​
268 items collected.
1. Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation ​
Author: Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time. Such a judge is a reward-f...
2. Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking ​
Author: Pradeep Kumar Sharma, Shantanu Godbole, Hritvik Shrivastava
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13565v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures scale large language models (LLMs) while preserving computational efficiency through sparse activation. Despite their widespread adoption, the relative importance of individual MoE layers remains insufficiently ch...
3. Modular Cognitive Architecture Emerges in Large Language Models ​
Author: Pengrui Han, Jacob Andreas, Evelina Fedorenko, Andrea Gregor de Varda
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.13567v1 Announce Type: new Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental princ...
4. A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing ​
Author: William Nixon, Jon Durbin, Florian Standhartinger, Haryadi S. Gunawi, Juncheng Yang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13573v1 Announce Type: new Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and scope. They often ob...
5. Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents ​
Author: Bo Jin, Qiang Jiao, Xin Tong
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.13574v1 Announce Type: new Abstract: LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over-privileged act...
6. AI Evaluation Should Work With Humans ​
Author: Jan Kulveit, Gavin Leech, Tom'a\v{s} Gaven\v{c}iak, Raymond Douglas
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI communi...
7. Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors ​
Author: Akira Okutomi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We co...
8. Measuring Cross-Task Behavioral Consistency in Language Model Agents ​
Author: Amritesh Banerjee, Pranil Raichura
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13598v1 Announce Type: new Abstract: Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavioral consistency across tasks is a distinct and measurable property, and w...
9. Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents ​
Author: Babak Abbaschian
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC, cs.MA
arXiv:2608.13604v1 Announce Type: new Abstract: Detection of misunderstanding is an urgent problem to solve because communication has moved away from real-time, in-person interaction and is increasingly handled by AI-mediated channels. This shift cuts communicators off from the resources repair depe...
10. Active Perception for Embodied Disambiguation ​
Author: Yiwei Liu, Luwei Yang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.13605v1 Announce Type: new Abstract: Natural language provides robots with a flexible task interface, but target ambiguity in embodied environments arises not only from user intent; it can also result from missing taskrelevant physical evidence in the current observation. Existing interac...
11. MobileMem: Learning from a Year of Mobile Experiences ​
Author: Xinle Deng, Yida Xue, Xiangyuan Ru, Haoming Xu, Shuofei Qiao, Mengru Wang, Yijun Chen, Buqiang Xu, Chen Jiang, Yuchen Eleanor Jiang, Lizhong Wang, Jianfeng Wang, Li Zeng, Haofen Wang, Guilin Qi, Huajun Chen, Ningyu Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, cs.MM
arXiv:2608.13606v1 Announce Type: new Abstract: The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-term...
12. No Universal Signal Predicts Sample-Level LLM Regression under Version Updates ​
Author: Jia Sheng, Yiwei Lu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model beco...
13. Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis ​
Author: Aryan Luthra, Kshitij Jain, Siddharth Arya, Bobby Filar, Anna Bertiger
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG
arXiv:2608.13608v1 Announce Type: new Abstract: Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. But their value is conventionally measured by gains against labeled benchm...
14. SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data ​
Author: Bruno Santos Teixeira
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.13612v1 Announce Type: new Abstract: Natural-language interfaces to enterprise data must translate underspecified requests into governed, executable behavior while controlling invalid queries, policy failures, cost, and nondeterminism. SemPlan Benchmark evaluates this architectural design...
15. How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights ​
Author: Himanshu Tripathi, Kaushik Roy, Subash Neupane, Shahram Rahimi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.SC
arXiv:2608.13617v1 Announce Type: new Abstract: Verifying whether clinical care follows evidence-based protocols is a natural neuro-symbolic problem, yet the safety-critical setting defeats either paradigm alone. We present an expert-guided pipeline that constrains a large language model strictly to...
16. Algorithm Design and Physician Liability ​
Author: Shujie Luan, Shubhranshu Singh, Tinglong Dai
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, econ.TH
arXiv:2608.13618v1 Announce Type: new Abstract: A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United ...
17. Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning ​
Author: Yongchao Huang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13621v1 Announce Type: new Abstract: A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagation through a Markov transition, and emission back to observation space. We show that full, time-indexed Predictive Information Bottleneck V...
18. ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction ​
Author: Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.13622v1 Announce Type: new Abstract: Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compar...
19. Reward Machines for Signal Temporal Logic ​
Author: Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.RO
arXiv:2608.13625v1 Announce Type: new Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued observations, along with a quantitative robustness score for monitoring satisfaction. Control synthesis from STL specifications is of interest sin...
20. A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure ​
Author: Dekun Yang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.13626v1 Announce Type: new Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice a...
21. Exploring ESC Winners with Nested Diagrams ​
Author: Anurag Sharma, Marcel N"ohre, Gerd Stumme
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13630v1 Announce Type: new Abstract: We present ConceptFlow, a scikit-learn-compatible Python library for Formal Concept Analysis that constructs and renders nested line diagrams from many-valued formal contexts. Given a many-valued context and a partition of its attributes into conceptua...
22. Ontology-Grounded Project Memory for Coding Agents ​
Author: James Adam
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.13662v1 Announce Type: new Abstract: Coding agents have become the primary means of generating new code in many software projects, and the resulting velocity of changes makes keeping track of the reasons behind those changes challenging. This paper introduces MOOSEDev, a system designed t...
23. Second Thought: Reasoning in Parallel as LLM Agents Act and Observe ​
Author: Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.13667v1 Announce Type: new Abstract: LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is confined to the Thought phase: while the agent serializes an action and waits for the environment, its reasoning is frozen. We identify thi...
24. Learning to Assemble Novel Structures with Unfamiliar Parts under Semantic Constraints ​
Author: Jonghyuk Park, Alex Lascarides, Subramanian Ramamoorthy
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13684v1 Announce Type: new Abstract: This paper describes a neurosymbolic architecture for learning to assemble novel structures using evidence from embodied conversations and task demonstrations. We focus on scenarios where an agent encounters, after deployment, semantic constraints on s...
25. Coverage Aware Active Evaluation for Failure Discovery with Paired Systems ​
Author: Anjali Parashar, Rachel Luo, Apoorva Sharma, Sushant Veer, Edward Schmerling, Carson Sobolewski, Mingxin Yu, Chuchu Fan, Marco Pavone
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.13719v1 Announce Type: new Abstract: Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets. Although cheaper proxies such as simulators, lower-fidelity systems, or related policies can be sampled extensively...
26. Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation ​
Author: Ajay Pravin Mahale (Hochschule Trier)
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13754v1 Announce Type: new Abstract: The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions. Mechanistic interpretability is the obvious source of such evidence, and circuit discovery is its most developed inst...
27. Simulation-Aware In-Context Policy Improvement for LLM-Aided Analog Layout Refinement ​
Author: Bingyang Liu, Ziming Wei, Xiaohan Gao, David Z. Pan
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.13767v1 Announce Type: new Abstract: Analog IC layout design remains a labor-intensive iterative process dominated by simulation-driven refinement. Although end-to-end layout generators accelerate initial placement and routing, they still require experts to manually tune layout optimizati...
28. FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC ​
Author: Harini Venkatesan, Christian Shelton, Ming-Feng Ho, Simeon Bird, Mengxuan Wu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13774v1 Announce Type: new Abstract: Markov chain Monte Carlo (MCMC) requires only the ability to evaluate the likelihood, making it a common technique for inference in complex models. However, it can have a slow mixing rate, requiring the generation of many samples to obtain good estimat...
29. From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL ​
Author: Wenyue Hua, Zachary Huang, Tyler Payne, Safoora Yousefi, Saleema Amershi, Asli Celikyilmaz
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA
arXiv:2608.13787v1 Announce Type: new Abstract: AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller,...
30. SDO: Subspace Deconflicting Operator for Multi-Adapter Composition ​
Author: Zhongsheng Wang, Zhedong Lin, Qian Liu, Xinyu Zhang, Jiamou Liu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13820v1 Announce Type: new Abstract: Composing independently trained adapters within a shared diffusion backbone provides a modular approach to multi-character generation, but naive joint deployment often causes identity mixing, cross-character attribute leakage, and unstable scene compos...
31. Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference ​
Author: Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13863v1 Announce Type: new Abstract: Deep neural network (DNN) inference on mobile devices often incurs high latency and energy consumption due to limited computing and memory resources. To enable energy-efficient DNN inference, most existing studies focus on dynamic voltage and frequency...
32. MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends ​
Author: Chaoqun Zhan, Qiang Zhou, Guannan Li, Zhenqiang Huang, Qianjin Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13883v1 Announce Type: new Abstract: Most agent-memory benchmarks test post-hoc recall, whereas MemoryArena evaluates whether memory supports interdependent, multi-session task completion. We compare MemoryLake, a structured multi-track memory backend, with Mem0, text-embedding-3-small ve...
33. When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict ​
Author: Lu Yang, Shusheng Xu, Zhuoran Li, Tongkai Yang, Longbo Huang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13921v1 Announce Type: new Abstract: LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one ...
34. Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact ​
Author: Zhelun (Allen), Wu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB
arXiv:2608.13926v1 Announce Type: new Abstract: Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indist...
35. AI Research Preference Models ​
Author: Thomas Simon Foster, Bassel Al Omari, Tingchen Fu, Thomas Mann, Carl Domond, Lucia Cipolina-Kun, Bhavul Gauri, Muna Aghamelu, Alexander D. Goldie, Eryk Helenowski, Jean-Christophe Gagnon-Audet, Alberto Pepe, Saba Nazir, Daniel Izcovich, Noam Levi, Rishi Hazra, Karen Hambardzumyan, Nicolas Baldwin, Xian Li, Martin Josifoski, Paris Giampouras, Masoud Jalili Sabet, Anya Sims, Hela Momand, Tatiana Shavrina, Despoina Magka, Jason Weston, Yulin Wang, Anirudh Goyal, Jo~ao Henriques, Yoram Bachrach, Emily McMilin, Jakob Nicolaus Foerster
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13940v1 Announce Type: new Abstract: AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days...
36. HELIX: Model-Harness Co-evolution for Recursive Self-Improvement ​
Author: Tianyu Fan, Chao Huang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13951v1 Announce Type: new Abstract: Scaling agent capability has largely focused on improving the model, yet an interactive agent acts through a runtime harness that mediates context, tools, control flow, and stopping. The harness shapes both what a model can accomplish and the trajector...
37. Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence ​
Author: James K. Wiles
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LO
arXiv:2608.13958v1 Announce Type: new Abstract: How do we govern AI systems whose reasoning we cannot fully inspect? Governance does not require understanding a system's reasoning. It requires stating what the system is obliged, permitted, and forbidden to do, and checking whether it complied. I pre...
38. Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead ​
Author: John T. Halloran
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.13987v1 Announce Type: new Abstract: Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters. Evaluated on Apple Silicon (MPS), we identify five in...
39. Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy ​
Author: Liwei Deng, Jing Jiang, Zhiwei Li, Yang Wang, Guodong Long
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.13990v1 Announce Type: new Abstract: Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effecti...
40. Simulation-Driven Vehicular Traffic Data Augmentation: Extending Sensor Coverage Through Virtual Sensing ​
Author: Davide Andrea Guastella, Eladio Montero Porras, Evangelos Pournaras, Gianluca Bontempi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13993v1 Announce Type: new Abstract: Urban traffic management relies on sensor networks whose spatial coverage is limited by deployment costs and privacy regulations. Machine learning models trained on such sparse data cannot generalize to unmonitored locations and must be retrained whene...
41. Buy the Rumor, Sell the News: When Is News Priced In? ​
Author: Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini, Sid Ghatak, Arman Khaledian
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-fin.ST
arXiv:2608.14014v1 Announce Type: new Abstract: Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. Both place the price move associated with a piece of news before and at publication rather than after it. Wheth...
42. Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders ​
Author: Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa, Miki Haseyama
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.14021v1 Announce Type: new Abstract: Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear. We comb...
43. Agent-Orchestration in Autonomous Chip Design ​
Author: Linyang Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14035v1 Announce Type: new Abstract: Recent developments in large language models (LLMs) and tool-using agents encourage people to explore the potential of using agents in chip design. The core question is what kind of AI we really need in such a sophisticated industry. To this end, we br...
44. Demystifying Agent Skills: Why They Work-Until They Don't ​
Author: Zhiyuan Jiang, Fangrui Huang, Hanwen Xing, Xander Wu, Yipeng Gao, Rui Cao, Mengdi Wang, Shilong Liu, Yijiang Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14036v1 Announce Type: new Abstract: Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely measure whether skills improve aggregated task success, leaving a more ...
45. Benchmarking data-driven material models on the classic Treloar dataset ​
Author: Hagen Holthusen, Moritz Flaschel, Denisa Martonov'a, Ellen Kuhl
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2608.14063v1 Announce Type: new Abstract: Machine learning is rapidly reshaping constitutive modeling, offers new ways to learn material behavior directly from experimental data, and challenges long-established modeling paradigms. But with a growing number of machine-learning-based approaches ...
46. Scaling Domain Data Repetition in LLM Pretraining ​
Author: Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu, Shuran Zheng, Jingzhao Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio ((\mathrm{TPP})). However, high-quality domain data is much harder to scale than general web data. As model size and...
47. Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails ​
Author: Giovanni Racioppi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14074v1 Announce Type: new Abstract: AI agents increasingly act on external systems through standardized tool-calling protocols such as the Model Context Protocol (MCP), yet no infrastructure layer constrains their actions to what a principal has verifiably authorized: authorization logic...
48. A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images ​
Author: Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.DL
arXiv:2608.14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction fr...
49. Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers ​
Author: Thiago Sandoval, Ufuk Topcu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.LG
arXiv:2608.14089v1 Announce Type: new Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We pr...
50. Retrieval Grounding Latent Reasoning for Dense Retrieval ​
Author: Gang Zhou, Xiongxi Yu, Hu Tian, Yang Wei, Lu Pan, Ke Zeng, Shibiao Xu, Xiaolong Zheng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14107v1 Announce Type: new Abstract: Reasoning-intensive retrieval requires text representations to capture not only semantic similarity, but also the reasoning needed to determine relevance under a given retrieval instruction. Existing reasoning-enhanced embedding models improve retrieva...
51. A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents ​
Author: Ismail El Hamraoui, Sagar Jose, Nicolas Bureau, Robert Plana
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.14109v1 Announce Type: new Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviation from the original task that can lead to irreversible side effects on external systems. Existing ...
52. Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground ​
Author: Arne Kr"oger, Ralf Buscherm"ohle, Wilhelm Hasselbring, Henrik Wilbers
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14122v1 Announce Type: new Abstract: Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously. While reinforcement learning (RL) methods have shown pro...
53. Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost ​
Author: Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal. However, LeWM has two limitations. First, during training, ...
54. QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation ​
Author: Lin-Fa Lee, Yi-Yu Chang, Kuo-Hei Yeh
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.14149v1 Announce Type: new Abstract: Recent training-free post-training quantization methods restore model accuracy through closed-form residual compensation. To constrain additional model storage overhead, several existing methods gate layer selection by goodness-of-fit, retaining only t...
55. Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach ​
Author: Sheng Hong, Xuanqi Wang, Jiacheng Wang, Yuwei Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14152v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-the-shelf models struggle to filte...
56. Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine ​
Author: Chenran Weng, Joo Seung Lee, Malini Mahendra, Anil Aswani
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.14157v1 Announce Type: new Abstract: Mechanical ventilation is a critical life-support intervention, requiring dynamic adjustments to ventilator settings as a patient's condition evolves. While reinforcement learning (RL) offers a promising framework for optimizing these sequential decisi...
57. BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs ​
Author: Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos, Karin Sevegnani, Emine Yilmaz
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14161v1 Announce Type: new Abstract: LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and mitigating bias, it largely focuses on final outputs of models, with lim...
58. Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding ​
Author: Jeongwan Shin, Jaehyeon Kim, Donguk Ko, Jaeho Choi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14179v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate re...
59. FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction ​
Author: Pengfei Chen, Yize Wu, Shouxu Kuang, Ke Gao, Ling Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.14205v1 Announce Type: new Abstract: Load imbalance poses a major bottleneck to the efficiency of expert parallelism in distributed inference of Mixture-of-Experts (MoE) models. The most heavily loaded rank stalls global execution due to skewed routing distributions, directly increasing l...
60. APTER: Adaptive Post-Training with Expert-Grounded Rubrics ​
Author: Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14212v1 Announce Type: new Abstract: As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic pref...
61. A Generalized Parallelogram Rule for Proportional Analogies on Riemannian Manifolds ​
Author: Pierre-Alexandre Murena, Marcelo Hartmann
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14220v1 Announce Type: new Abstract: Analogies are quaternary relations of the form "a is to b as c is to d", usually denoted a : b :: c : d. This notion is formalized in particular with the notion of proportional analogy, which imposes some constraints on the valid analogies. Whereas pro...
62. MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement ​
Author: Lushi Pu, Weiming Zhang, Xinheng Xie, Zixuan Fu, Bingxiang He, Hengyu Zhao, Hongya Lyu, Xin Li, Jie Zhou, Yudong Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.14221v1 Announce Type: new Abstract: Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation. Models must map mathematical concepts ...
63. Attributing Preprocessing Invariance in Spectral Foundation Models ​
Author: Dongjun Wei, Hongyi Wu, Yinuo Zou
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.LG
arXiv:2608.14227v1 Announce Type: new Abstract: Preprocessing invariance is an appealing goal for spectral foundation models: a frozen model should remain useful when laboratories preprocess spectra differently. It is usually measured by training a classifier under one preprocessing pipeline and tes...
64. Polaris : Multi Agentic System for Conversational Enterprise Analytics ​
Author: Varuni H K, Soham Sarkar, Jay Kumar, Goutham Krishnan, Tanvi Johari, Avinash Bharadwaj, Santosh Hegde
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14246v1 Announce Type: new Abstract: In today's fast-paced environment, the ability to swiftly access, understand, and act on data is no longer optional; it is essential. Yet most organizations remain data-rich but insight-poor, constrained by the complexity of querying, interpreting, and...
65. Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models ​
Author: Brett Reynolds
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.14252v1 Announce Type: new Abstract: Recent work suggests that some large language model representations have content or reference. Grounding can secure either without supplying live routes for correction. This paper asks what follows from that gap. An output is answerable when discrepanc...
66. TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments ​
Author: Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li, Stefan Zohren, Anna Vettoruzzo, Qingsong Wen, Ming Jin, Joaquin Vanschoren
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14270v1 Announce Type: new Abstract: Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, leaving temporal...
67. Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning ​
Author: Kai Chen, Jifeng Ding, Ning Ding, Jiaye Ge, Lixin Gu, Yicheng Gu, Qipeng Guo, Ermo Hua, Haian Huang, Haozheng Hou, Jie Hou, Xiangyu Hong, Che Jiang, Minxi Jin, Cheng Liang, Dahua Lin, Dawei Liu, Kuikun Liu, Chengqi Lv, Haijun Lv, Han Lv, Ningsheng Ma, Biqing Qi, Jianmin Qian, Shiya Su, Youbang Sun, Huanze Tang, Zhongbo Tian, Hanjing Wang, Rui Wang, Ting Wang, Yi Wang, Baiting Wu, Jun Xu, Bowen Yang, Hui Wang, Weida Wang, Haochen Ye, Jiashuo Yu, Shan Yu, Xiaoyi Yu, Qirui Zeng, Qi Zhang, Ming Zhang, Wenwei Zhang, Bowen Zhou, Xinyu Zhou
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14290v1 Announce Type: new Abstract: We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners...
68. Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety ​
Author: Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY
arXiv:2608.14306v1 Announce Type: new Abstract: This paper presents a coordination architecture for heterogeneous UAV/UGV swarms that synthesises mission actions from uncertain, multi-modal sensor evidence while preserving hardware-enforced safety at the actuation boundary. The approach combines rad...
69. AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs ​
Author: Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.14320v1 Announce Type: new Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibi...
70. Program-space Diffusion for Morphology-to-Transcriptomics Prediction ​
Author: Ruyter Swann, Dorent Reuben, Racoceanu Daniel
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14330v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables genome-wide gene expression profiling while preserving tissue architecture, but its cost and limited scalability remain major bottlenecks. This has motivated models that predict spatial expression directly from rout...
71. Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents ​
Author: Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, Anthony G Cohn
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.14339v1 Announce Type: new Abstract: We study proactive exploration in LLM agents, i.e., the ability to explore an environment to acquire information that improves future decision-making. In this regard, we first identify two fundamental bottlenecks that hinder this capability and then pr...
72. ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond ​
Author: Mingming Zhao, Jiqian Dong, Kangping Xu, Zadid Hasan, Chengrui Fan, Shan Jiang, Shuai Mao, Ting Lingya, Linyi Zou, Tailin Zhou, Yun Hin Chan, Wenkai Zhang, Zhanhong Zhou, Guowei Huang, Hongliang Li, Wenjing Cun, Zhitang Chen, Mingxuan Yuan, Yanhui Geng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14354v1 Announce Type: new Abstract: Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration...
73. Disentangled Shared Representations Improve Morpho-Transcriptomic Integration ​
Author: Julian Ostermaier, Swann Ruyter, Reuben Dorent, Daniel Racoceanu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14355v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure. However, standard multimodal models ...
74. Designing Sustainable Federated Learning as a Service using Neural Architecture Search ​
Author: Keya Patel, Sajib Mistry, Sheik Fattah, Deepak Kanneganti, Aneesh Krishna, Mufti Mahmud, Monowar Bhuyan
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14359v1 Announce Type: new Abstract: The sustainability constraints of FLaaS consumers pose significant challenges to maintaining carbon-feasible federated training in FLaaS environments. These constraints often lead to infeasible consumer participation and unstable federated training und...
75. Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages ​
Author: Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.14375v1 Announce Type: new Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a ...
76. AgentRewind: Recoverable Execution for Long-Horizon LLM Agents ​
Author: Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14380v1 Announce Type: new Abstract: Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent context and environment state, and their effects may be difficult to re...
77. Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons ​
Author: Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14392v1 Announce Type: new Abstract: Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifica...
78. LLMs Don't Pay for the Jump ​
Author: Paras Balani, Subhrakanta Panda
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.14397v1 Announce Type: new Abstract: Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simula...
79. The Past and Future of AI Scientists ​
Author: Ross D. King
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists: machines capable of automating science. AI Scientists can originate hypotheses, deduce their consequences, design and execute experiments, interpret their results, and revise their beliefs. S...
80. Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations ​
Author: Toby D. Pilditch
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement probl...
81. The Dynamics of Intelligence Explosions ​
Author: Toby Ord
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, econ.TH
arXiv:2608.14426v1 Announce Type: new Abstract: AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be able to produce an intelligence explosion, with rapidly escalating AI capabilities. I explore the mathematics of the most explosive possibilities, w...
82. PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments ​
Author: Yuhao Zhan, Bingxiang He, Zecong Tang, Chaojun Xiao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14441v1 Announce Type: new Abstract: Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench ...
83. Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports ​
Author: Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, Lorenz K. M"uller, Konstantin Berestizshevsky, Lukas Cavigelli
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14446v1 Announce Type: new Abstract: In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To a...
84. SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning ​
Author: Panjing He, Mingyue Cheng, Yucong Luo, Li Li, Xiaohan Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14452v1 Announce Type: new Abstract: Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet reasoning remains challenging for large language models (LLMs). Real-world workbooks often contain implicit cross-table associations, fine...
85. Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments ​
Author: Shahab Band, Hamed Mohammadi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14456v1 Announce Type: new Abstract: Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the target domain are limited and the statistical properties of the source and target domains differ. In these settings, models trained only on local ...
86. Twin: Playing an Unknown Game with a Test-Time Digital Twin ​
Author: Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14490v1 Announce Type: new Abstract: We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom...
87. Split the Labor: Separating Evidence Interpretation from Decision Aggregation ​
Author: Zhelun Wu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.14509v1 Announce Type: new Abstract: Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations...
88. Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers ​
Author: Taenyun Kim, Edyta Bogucka, Daniele Quercia
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.14522v1 Announce Type: new Abstract: As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI mode...
89. Handover of In-Context Learning State Across Session Boundaries ​
Author: Masahiro Kato, Taka Kato
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, econ.EM, math.ST, stat.ME, stat.ML, stat.TH
arXiv:2608.14528v1 Announce Type: new Abstract: This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restar...
90. Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support ​
Author: Alexandre Cristov~ao Maiorano
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SE
arXiv:2608.13563v1 Announce Type: cross Abstract: Early-stage teams often lack users, time, and budget to run repeated UX studies, yet still need decision-oriented signals to iterate safely. We study an LLM-driven UX micro-simulation pipeline that generates structured customer-experience feedback (w...
91. Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required ​
Author: Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bartak, Egor Bogomolov, Sergey Titov
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE
arXiv:2608.13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing systems. We argue that op...
92. Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study ​
Author: Pengcheng Xu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13568v1 Announce Type: cross Abstract: Coding agents spend most of their context budget on retrieval. Lexical retrieval (grep) is universal, instant, and zero-setup, but noisy: it cannot tell a definition from a call from a comment. Semantic retrieval via the Language Server Protocol (LSP...
93. Think in Latent, Explain in Language: Self-Explainable Latent Reasoning ​
Author: Dayuan Zhao, Shengcao Cao, Yu-Xiong Wang, Liang-Yan Gui
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.13570v1 Announce Type: cross Abstract: Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the late...
94. Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems ​
Author: Heming Fu, Shan Lin, Qianqian Xie, Guojun Xiong
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full workflow actually c...
95. The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel ​
Author: Mohammad Imrul Jubair, Tom Yeh
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.13572v1 Announce Type: cross Abstract: We present The Architect, a system that turns Microsoft Excel into an interactive view of deep learning mathematics. A user describes a neural network in a compact table. The system then generates a workbook that shows the full forward pass and, when...
96. Interactive Analysis of Global Explanations using Aggregated Class Activation Maps for Network Data ​
Author: Igor Cherepanov, David Sessler, Alex Ulmer, Felix Wagner, Throsten May, J"orn Kohlhammer
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.NI
arXiv:2608.13575v1 Announce Type: cross Abstract: Recent machine learning (ML) advances have demonstrated that deep learning (DL) achieves impressive results in different application domains, including the classification of computer network traffic to corresponding applications. However, the data fr...
97. BCMT: Blockwise Causal Memory Transformer ​
Author: Rachid Arezki
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.13578v1 Announce Type: cross Abstract: Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for lo...
98. Jais 2: A Family of Arabic-Centric Open Large Language Models ​
Author: Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evalua...
99. From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent ​
Author: Mingyu Huang, Weiqing Min, Ying Jin, Yilin Wang, Shuqiang Jiang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG
arXiv:2608.13581v1 Announce Type: cross Abstract: Personalized glucose regulation remains a central yet unresolved challenge in precision nutrition, as postprandial glucose response varies substantially across individuals. Existing approaches based on glycemic indices fail to adequately account for ...
100. UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction ​
Author: Mikhail Kiselev, Aleksandr Marukhin, Ivan Snegirev, Elizaveta Semenyakina, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CV
arXiv:2608.13584v1 Announce Type: cross Abstract: UltraArUco - a lightweight multilingual library and framework for low-latency, real-time marker-based tracking in mobile augmented reality. Unlike standard OpenCV-based implementations, UltraArUco introduces an optimized multilingual wrapper that red...
101. IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering ​
Author: JungMin Yun, YoungBin Kim
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13588v1 Announce Type: cross Abstract: Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy contexts, thereby undermining both efficiency and accuracy. While existing...
102. Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Procedural Guidance ​
Author: Rodrigo Gallardo, Qilmeg Doudatcz, Ganit Goldstein, Ilkyaz Sarimehmetoglu, Sergio Mutis, Alexander Htet Kyaw, Anita Lin, Clara Emmerling, Berfin Ataman, Skylar Tibbits
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.13589v1 Announce Type: cross Abstract: As human space exploration returns to the Moon, astronauts need rapid access to procedural information during extravehicular activities (EVAs), where attention is divided across navigation, repair tasks, tool handling, and environmental risk. The cha...
103. Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning ​
Author: Jiahe Fan, Si Chen, Yinghao Hou, Aiyuan Zhang, Hong Xie
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13596v1 Announce Type: cross Abstract: Heterogeneous model fusion seeks to combine models that differ in tasks, initializations, architectures, or scales. We study an underexplored cross-scale setting: improving a small recipient language model with a stronger donor despite substantial ar...
104. Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Models ​
Author: Hongxin Xu, Jianping Mei, Can Wang, Defang Chen
Published: 8/17/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CR, cs.CV, cs.MM
arXiv:2608.13597v1 Announce Type: cross Abstract: Coverless image steganography (CIS) synthesizes a stego image rather than modifying an existing cover image, enabling authorized recipients to reconstruct the original secret image from the stego. Existing diffusion-based CIS methods can generate nat...
105. Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation ​
Author: Zhe Liu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD
arXiv:2608.13624v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation in spoken-input sett...
106. From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models ​
Author: Pranav Kumar Kaliaperumal
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software. The ability to resolve real coding issues improved by nearly six times per year since late 2024. Durin...
107. Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT ​
Author: Pu Zhao, Changdi Yang, Yixiao Chen, Yi Gao, Yifan Cao, Haochen Zeng, Yanzhi Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.ET
arXiv:2608.13681v1 Announce Type: cross Abstract: Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional behavior of legacy systems. Large language models (LLMs)...
108. MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation ​
Author: Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13690v1 Announce Type: cross Abstract: Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods with...
109. SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers ​
Author: Kiran Nair, Rodrigue Rizk, KC Santosh
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.NE
arXiv:2608.13702v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires sur...
110. CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA ​
Author: Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13706v1 Announce Type: cross Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occur...
111. TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials ​
Author: Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak...
112. Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation ​
Author: Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable l...
113. Capacity-Dependent Effects of Data Selection for Reasoning ​
Author: Cuong Dang, Hoang Anh Just, Ruoxi Jia
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that responses closer to ...
114. Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost ​
Author: Victor Barros de Miranda Neves, Kiev Santos da Gama, Vinicius Cardoso Garcia
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.13730v1 Announce Type: cross Abstract: Empirical reports on the true cost of AI-intensive software development remain scarce, and the few that exist are easy to get wrong in ways that never surface in the final number. We report early results from an ongoing case study: a six-person stude...
115. Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline ​
Author: Jo`ao Pedro Monteiro Pereira, Vinicius Cardoso Garcia
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.13742v1 Announce Type: cross Abstract: In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line phrases. We ask whether grounding those specifications in ISO/IEC 25010 Quality Model, either as rich natural-language prose (NL-rich) or as struct...
116. Data-driven techniques for translational neuroscience and personalized neuro-health ​
Author: Vishal Subedi, Shashipraba N. K. Rajakaruna, Pratyusha Sarkar, Subhankar Chattoraj, Anjali Khasa, Siddhartha Nandy, Hamza Farooq, Animikh Biswas, Sanjay Chaudhuri, Asim K. Dey, Karuna Joshi, Christophe Lenglet, Ansu Chatterjee
Published: 8/17/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG, stat.AP, stat.ML
arXiv:2608.13749v1 Announce Type: cross Abstract: Neurodegenexrative diseases such as Alzheimer's disease and Parkinson's disease are diagnosed most reliably only after substantial, often irreversible, neuronal loss has already occurred, creating an urgent need for quantitative tools that can detect...
117. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models ​
Author: Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG
arXiv:2608.13760v1 Announce Type: cross Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative ...
118. CutClean: Neural Network Pruning for Privacy-Preserving Inference ​
Author: Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise, Enzo Tartaglione
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13773v1 Announce Type: cross Abstract: Neural networks are increasingly deployed in high-stakes applications with growing privacy leakage concerns. We show that this privacy leakage can occur even in the absence of representation imbalances that lead to traditional dataset biases. This po...
119. Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions ​
Author: Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.13786v1 Announce Type: cross Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studi...
120. PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization ​
Author: Ruogu Chen, Jie Han
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR
arXiv:2608.13790v1 Announce Type: cross Abstract: Macro placement significantly affects a chip's post-route performance, power, and area (PPA). Most placement methods optimize half-perimeter wirelength (HPWL) as the primary objective. However, recent benchmarking shows a near-zero correlation betwee...
121. Optimal Power Allocation and AI Receiver Design for Superimposed DMRS and Data Transmission ​
Author: Sha Hu, Zhongwang Fu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT
arXiv:2608.13809v1 Announce Type: cross Abstract: In this paper, we consider transmissions with superimposed (SI) demodulation-reference-symbol (DMRS) and data in orthogonal frequency-division multiplexing (OFDM) based multiple-input multiple-output (MIMO) systems. First, we derive an analytical fra...
122. AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution ​
Author: Simiao Zuo, Chenhui Xu, Yimeng Jia, Qiang Lou, Jian Jiao, Denis Charles
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.13833v1 Announce Type: cross Abstract: Conversational advertising aims to deliver useful ads within multi-turn assistant interactions. Unlike conventional query-based advertising, where the user's intent is often expressed in a short standalone query, conversational ads must infer latent ...
123. ASSERT: A Measurement Pipeline for GenAI Audits ​
Author: Riccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng, Chad Atalla, Jean Garcia-Gathright, Nicholas Pangakis, Sharman Tan, Dan Vann, Hannah Washington, P. Alex Dow, Heba Elfardy, Hanna Wallach, Sandeep Atluri
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.13840v1 Announce Type: cross Abstract: Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A reported ...
124. Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions ​
Author: Qinglin Yang, Chen Qiu, Hongyuan Zhang, Pengdeng Li, Yuan Liu, Zhihong Tian
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC
arXiv:2608.13844v1 Announce Type: cross Abstract: Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns. Federated le...
125. Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model ​
Author: Stephanie Jarmak
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.13867v1 Announce Type: cross Abstract: AI coding agents are commonly evaluated as models but deployed as systems. Their reliability depends not only on model capability, but on the harness, execution state, retrieval, memory and state management, permissions, review interfaces, and resour...
126. Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen ​
Author: Jiada Li, Xuesong Ye, Olamide Olowoniyi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.ET, cs.HC, cs.LG
arXiv:2608.13884v1 Announce Type: cross Abstract: The rapid adoption of AI coding assistants and autonomous agentic development systems has coincided with major changes in the pace and structure of open-source software engineering. Yet empirical longitudinal evidence of these changes at the team lev...
127. Agentic Transaction: Towards ACID-Compliant Agent Systems ​
Author: Zhaoyan Sun, Xiaoxiao Wang, Guoliang Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.LG
arXiv:2608.13900v1 Announce Type: cross Abstract: Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents increasingly operate over persis...
128. CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts ​
Author: Runhan Song, Qiqi Liu, Chuanzhou Pan, Zhenquan Ding, Youquan Xian, Chongru Fan, Lei Cui, Wei Wang, Zhiyu Hao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.NI
arXiv:2608.13905v1 Announce Type: cross Abstract: HTTPS website fingerprinting (WF) aims to identify visited websites from metadata observable in encrypted traffic. However, real-world deployments introduce a significant out-of-distribution (OOD) problem caused by temporal and geographic changes, wh...
129. Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning ​
Author: Chun-Hua Lin, Samuel Yen-Chi Chen, Yu-Chao Hsu, Kuo-Chung Peng, Jiun-Cheng Jiang, Chi-Sheng Chen, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo, Hsi-Sheng Goan
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.ET, quant-ph
arXiv:2608.13914v1 Announce Type: cross Abstract: Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devices to share raw signals for centralized model training. Federated learning addresses this practical privacy constraint by enabling c...
130. CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing ​
Author: Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strate...
131. CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification ​
Author: Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in ...
132. Musical Mirrors: The LLM as Sounding Board in Songwriting ​
Author: Xiao Xiao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.13944v1 Announce Type: cross Abstract: This paper examines a use of AI in creative practice as an interpretive sounding board for human-generated material, rather than the more familiar pattern of AI generation followed by human curation. Through the lens of resonance as theorized by Hart...
133. EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment ​
Author: Haokai Ma, Aoqi Hu, Yueao Xing, Ruobing Xie, Yonghui Yang, Teng Tu, Lei Meng, Tat-Seng Chua
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.14011v1 Announce Type: cross Abstract: Generative recommendation autoregressively generates the semantic IDs of the target item, unifying preference modeling and index retrieval within the shared token space. Recent attempts have introduced Multi-Token Prediction (MTP) into this field, ye...
134. MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning ​
Author: Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time. Existing approaches handle ...
135. Content Based Video Narration of Gameplay with Vision Language Models ​
Author: Mathew Varghese
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR
arXiv:2608.14016v1 Announce Type: cross Abstract: Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else. We present a content-based video narration system that produces spoken, esports-style commentary for arbitrary gameplay recordings using a general-...
136. ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models ​
Author: Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14022v1 Announce Type: cross Abstract: Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challe...
137. AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning ​
Author: Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang, Tong He, Zhengxue Cheng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.14028v1 Announce Type: cross Abstract: Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can...
138. HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation ​
Author: Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu, Wenbin Li, Boo-Ho Yang, Rav Lawana, Ziyue Li, Wei Zeng, Fugee Tsung
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.14032v1 Announce Type: cross Abstract: Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We propose HAM-RAG, a Hierarchy-...
139. Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use ​
Author: Yi Ding, Yanzhao Yu, Xili Dai, Xianbiao Qi, Peiwen Sun, Xueqian Wang, Xiangyu Yue, Jianan Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.14047v1 Announce Type: cross Abstract: This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level v...
140. Voxel-based 3D Facies Segmentation from Seismic Data: A Comparative Study ​
Author: Duc-Thanh Pham, Minh-Tan Pham, Anh Nguyen, Van Nguyen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14058v1 Announce Type: cross Abstract: Seismic facies segmentation has emerged as a significant challenge in geophysics, requiring robust methods and systems to effectively identify geologically analogous facies with limited labeled data. Although existing studies have shown promising res...
141. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency ​
Author: Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.14065v1 Announce Type: cross Abstract: Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on overall r...
142. MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation ​
Author: Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without w...
143. Reaction-Transformation-Aware Flow Matching for Generalizable Transition State Generation ​
Author: Kaipeng Zeng, Wenxi Zhai, Shengrui Xu, Jie Zhao, Bowen Li, Shiyue Wang, Junchi Yan, Tong Zhu
Published: 8/17/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI
arXiv:2608.14076v1 Announce Type: cross Abstract: Transition-state (TS) structures define the energetic barriers and mechanistic pathways of elementary chemical reactions, yet their identification remains computationally demanding because conventional saddle-point searches require expensive quantum-...
144. P2Skill: Privacy Preserving Skill Distillation for Cloud-Local LLM Inference Systems ​
Author: Myunghoon Ryu, Geunpyo Park, Sungjoon Lee, XinYu Piao, Jong-Kook Kim
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.14094v1 Announce Type: cross Abstract: Cloud-local LLM inference systems have the potential to use the reasoning capability of large cloud models while protecting sensitive user data on personal devices. Cloud-bound requests must exclude personally identifiable information (PII) to preven...
145. Rewrite Once, Validate Anywhere: Producing OWL-Aware SHACL Constraints (Extended Version) ​
Author: Anouk Oudshoorn, Piotr Gorczyca, D"orthe Arndt
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LO, cs.AI
arXiv:2608.14104v1 Announce Type: cross Abstract: The Shapes Constraint Language (SHACL) is a W3C recommendation to express syntactic constraints, called shapes, on RDF graphs. SHACL validators are used to test whether a given graph adheres to such a shape. However, RDF graphs often come with OWL on...
146. Forecast Collapse in Time-Series Foundation Models ​
Author: Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, stat.AP, stat.ML
arXiv:2608.14106v1 Announce Type: cross Abstract: When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phen...
147. Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation ​
Author: Michael R. Martin, Joseph Insley, Victor A. Mateevitsi, Silvio Rizzi, Kwan-Liu Ma
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CE, cs.GR, cs.LG
arXiv:2608.14112v1 Announce Type: cross Abstract: Scientific simulations often produce scalar volumes faster than they can be stored, transferred, and loaded, while in situ reduction must use only a limited share of simulation resources. This work encodes scalar fields as anisotropic Gaussian primit...
148. From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics ​
Author: Meng Li, Chuqi Chen, Zhengqing Gao, Xi Zhou, Xiao Sun, Yang Xiang, Huaxi Huang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GR
arXiv:2608.14120v1 Announce Type: cross Abstract: Lagrangian modeling is vital to fluid dynamics, as it characterizes particle transport and complements the Eulerian description.However, Lagrangian trajectories are less commonly available than Eulerian fields, while most neural operators are trained...
149. Overcoming Shortcut Learning in Graph Neural Networks through Active Explanation Guidance ​
Author: Taraneh Younesian, Steve Azzolin, Antonio Longa, Francesco Ferrini, Vincenzo Marco De Luca, Stefano Teso
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14121v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) can solve prediction tasks by unintentionally exploiting shortcuts---that is, edges, nodes, and features that correlate with but are not causal for the prediction---which compromise their reliability in out-of-distributio...
150. BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows ​
Author: Sheng Hong, Yixuan Huang, Weiwei Jiang, Junyuan Zhang, Jiacheng Wang, Ruijian Jiao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.14126v1 Announce Type: cross Abstract: To mitigate attention dilution in high-entropy TLS 1.3 flows, we propose BGA, a noise-immune neural distillation framework for encrypted threat intelligence.The methodology first employs Analysis of Variance (ANOVA) to decouple high-discriminatory co...
151. AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations ​
Author: Ying Huang, Wencan Zhang, Brian Y. Lim
Published: 8/17/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CV, cs.HC
arXiv:2608.14130v1 Announce Type: cross Abstract: Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed...
152. Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions ​
Author: Xiaokai Yan, Jingtao Ding, Yong Li, Zhiwen Yu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.14132v1 Announce Type: cross Abstract: Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence. However, current research primarily focuses on reactive task execution while lacking a comprehensive understanding-prediction-execution...
153. HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting ​
Author: Wei Zhang, Shengkai Yu, Shiqiang Gong, Qi Zhang, Qiang Li, Qi Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14136v1 Announce Type: cross Abstract: Octree-based anchor Gaussian Splatting has emerged as a scalable representation for city-scale novel view synthesis, where multi-level anchors adaptively capture scene content from coarse building structures to fine architectural details. However, we...
154. SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation ​
Author: Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14138v1 Announce Type: cross Abstract: Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specif...
155. Self-Supervised Visual On-Policy Distillation ​
Author: Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14144v1 Announce Type: cross Abstract: Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental qu...
156. Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation ​
Author: Nikolai R"ohrich, Isabell Hans, Felix Krause, Bj"orn Ommer
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.14172v1 Announce Type: cross Abstract: Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.g., for precisely controlling how aesthetically pleasing...
157. Structure-Guided Spatiotemporal Attention Graph Neural Network for Traffic Flow Prediction ​
Author: Xuanmian He, Can Li, Wanjing Ma
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14177v1 Announce Type: cross Abstract: Deep spatiotemporal models integrating graph convolutions and attention mechanisms have demonstrated excellent performance in network-level traffic flow prediction, owing to their exceptional ability to capture complex spatiotemporal dependencies. De...
158. How Much Do Legal RAG Systems Still Hallucinate? ​
Author: Souvick Das, Sallam Abualhaija, Domenico Bianculli
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.14210v1 Announce Type: cross Abstract: Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination...
159. Training Fair Tabular Foundation Models ​
Author: Patrik Kenfack, Jesse C. Cresswell, Anthony L. Caterini, Samira Ebrahimi Kahou, Ulrich A"ivodji
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14211v1 Announce Type: cross Abstract: Tabular Foundation Models (TFMs) have emerged as leading methods for tabular predictive tasks, leveraging in-context learning to predict on new data without task-specific training. Despite the increased use of TFMs in high-stakes decision-making, the...
160. Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response ​
Author: Timothy C. Pearce, David J. T. Smith, Alec Dobney, Alessia Freddo
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG, physics.ao-ph, physics.geo-ph
arXiv:2608.14254v1 Announce Type: cross Abstract: Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteoro...
161. Multi-Objective Bayesian Optimization for Model Merging ​
Author: Utkarsh Agarwal, Vamshi Bonagiri, Raul Astudillo, Monojit Choudhury
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14264v1 Announce Type: cross Abstract: Model merging combines trained models directly in weight space, offering a compute-efficient alternative to additional fine-tuning. Selecting merge parameters is nevertheless difficult because downstream evaluations are expensive, gradients are unava...
162. SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning ​
Author: Haonan He, Haodi Lei, Yun Luo, Haoran Zhang, Shunkai Zhang, Yizhuo Li, Shengji Tang, Zhilin Wang, Runzhe Zhan, Lei Bai, Ganqu Cui, Fangchen Yu, Yafu Li, Peng Ye, Ning Ding, Yu Cheng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.14277v1 Announce Type: cross Abstract: On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mi...
163. Seeing Red, Thinking Bad: Color Bias in Vision Language Models ​
Author: Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how ...
164. Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels ​
Author: Vadym Vilhurin, Volodymyr Sydorskyi, Andrii Shevtsov
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CV
arXiv:2608.14287v1 Announce Type: cross Abstract: Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles. However, the practical deployment of acoustic systems is discouraged by extreme environmental noise and sens...
165. Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans ​
Author: Tarandeep Singh Mandhiratta, ANK Zaman, Abdul-Rahman Mawlood-Yunis
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.HC, cs.LG
arXiv:2608.14317v1 Announce Type: cross Abstract: This research developed a neural network-based model to extract various information from 2D floor plans. We detect lighting symbols, identify the appropriate type of light, and extract the associated texts with lights. The study aims to enable effici...
166. A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation ​
Author: Dipankar Sarkar
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.CY, cs.LG
arXiv:2608.14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such...
167. Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes ​
Author: Keyvan Amiri Elyasi, Lukas Kirchdorfer, Heiner Stuckenschmidt
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14367v1 Announce Type: cross Abstract: The early detection of delayed cases in business processes is a critical capability for organizations. Predictive process monitoring (PPM) supports this task by using historical event logs to predict the remaining time of ongoing cases, enabling time...
168. A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models ​
Author: Md Kamrul Islam, Tiphaine Henry, Mattia Salnitri, Julius K"opke, Sami Souihi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2608.14370v1 Announce Type: cross Abstract: The modelling and analysis of secure business processes require the incorporation of security annotations into process models. Although BPMN extensions, including SecBPMN2, exist for this purpose, the derivation of accurate and complete security anno...
169. Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation ​
Author: Yuxuan Chen, Wanruo Zhang, Xiao Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.14379v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To ...
170. DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding ​
Author: Zewen Jin, Shen Fu, Zeping Duan, Shannon Wang, Weihao Wu, Chengjie Tang, Congkun Ai, Ping Gong, Zijian Dai, Youhui Bai, Cheng Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14385v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems. To meet the extremely low response latency requirements of these scenarios, practitio...
171. GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection ​
Author: Yingjie Ma, Zitong Yu, Wei Jia, Ajay Kumar, Linlin Shen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14389v1 Announce Type: cross Abstract: Existing palm presentation attack detection (PAD) datasets are often limited by static imagery, restricted acquisition conditions, or insufficient multimodal video data, hindering systematic evaluation across environments, modalities, and attack type...
172. Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination ​
Author: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14391v1 Announce Type: cross Abstract: Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and gener...
173. AI-Assisted Discovery and Construction of a Counterexample to the Convergence of Three-Block ADMM with the Identity Matrix as its Third Constraint Block ​
Author: Kenan Xu, Xiangfeng Wang
Published: 8/17/2026, 4:00:00 AM
Categories: math.OC, cs.AI
arXiv:2608.14396v1 Announce Type: cross Abstract: The alternating direction method of multipliers (ADMM), as a landmark algorithm, has attracted tremendous research attention and extensive practical applications over the past two decades. It is well known that, although the two-block ADMM enjoys wel...
174. Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice ​
Author: Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.14399v1 Announce Type: cross Abstract: Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physician...
175. From Style Replication to Style Exploration: Enabling Art Style Exploration with Analyze-Experiment-Resituate Framework ​
Author: Wen-Fan Wang, TsaiHsuan Lin, Chi-Lan Yang, An-Ru Cheng, Bing-Yu Chen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.14405v1 Announce Type: cross Abstract: Art style is a signature of professional digital artists that develops through repeated experimentation, reflection, and adaptation. While generative AI (GenAI) can reproduce styles with high fidelity, current tools provide limited support for explor...
176. Designing Compact Neural Architectures via Neuron Gating and Mixed Activation ​
Author: Abhishek Shukla, Ankur Sinha, Faiz Hamid
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14443v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss. However, NAS ...
177. LP-NAS: Linear Programming-based Neural Architecture Search ​
Author: Abhishek Shukla, Ankur Sinha, Faiz Hamid
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14472v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise. Among the various NAS methods, differentiable NAS has gained prominence due to its efficiency and accuracy compared to convent...
178. Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes ​
Author: Alexei Odinokov, Rostislav Yavorskiy
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.14481v1 Announce Type: cross Abstract: As heterogeneous robotic systems deploy across diverse urban zones, maintaining safety amid complex human-robot interactions remains a critical challenge. We present a unified framework that bridges systematic hazard analysis and runtime enforcement ...
179. Optimal Scheduling of Road Maintenance Jobs Considering Impact on Traffic Flows ​
Author: Charitha Nandepu, Lohitha Kalepu, Gabriele Ciavarella, SangWoo Park
Published: 8/17/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.14491v1 Announce Type: cross Abstract: Network-level maintenance planning requires repeated evaluations of equilibrium traffic flows under road capacity reductions. While equilibrium traffic assignment models are well established, their repeated solution quickly becomes computationally pr...
180. Generating Benchmark Health Data Using a Tabular Diffusion Transformer ​
Author: Hao Yan, Lisa Pilgram, Dan Liu, Linglong Kong, Fida Dankar, Khaled El Emam
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14496v1 Announce Type: cross Abstract: Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. However, existing synthetic tabular data generation methods are largely restricted to single-input-t...
181. Universal Thermodynamic Interatomic Potentials for Crystalline Materials ​
Author: Juno Nam, Bowen Deng, Xiaochen Du, Luis Barroso-Luque, Benjamin Kurt Miller, Rafael G'omez-Bombarelli
Published: 8/17/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cond-mat.stat-mech, cs.AI, cs.LG, physics.chem-ph
arXiv:2608.14502v1 Announce Type: cross Abstract: Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on ground-state energies because free energy calculations require ensemble averages. We introduce the thermodynamic interatomic potential (TI...
182. RecipeNet: A Hierarchical Transformer for Recipe Data ​
Author: Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi, Abhinav Kumar, Baoxin Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.14505v1 Announce Type: cross Abstract: Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, where procedures are represented as ordered sequences of steps containing heterogeneous structured fields. Existing tabular learning ...
183. Learning-to-Transition for Large-scale and High-Order MIMO Detection ​
Author: Yubo Zhang, Yiyao Liu, Xiaodong Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT
arXiv:2608.14511v1 Announce Type: cross Abstract: High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol space while producing reliable soft information for channel decoding. This paper develops a learning-to-transition (L2T) framework that ...
184. Marionette: Predicting World States, Rendering Geometry, Painting Appearance ​
Author: Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14530v1 Announce Type: cross Abstract: Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long hori...
185. Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils ​
Author: Karel Becerra, Boris Mederos, Dean Snow, Ram'on A. Mollineda
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.14539v1 Announce Type: cross Abstract: Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty in...
186. From Field Data to Global Food Systems Intelligence: A Semantic Graph Framework for Sustainable Wheat Production ​
Author: Nirmal Gelal, Aastha Gautam, Soheil Abadifard, Nico Giordano, Moumita Sen Sarma, Sanaz Saki Norouzi, Claudio Dias da Silva Jr, Jean Ribert Francois, Kathleen M. Jagodnik, Katherine Nelson, Terry Griffin, Xiaomao Lin, Stacy Hutchinson, Stephen M. Welch, Kelsey Andersen Onofre, Romulo Lollato, Pascal Hitzler, Hande K"u\c{c}"uk McGinty
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2502.19507v2 Announce Type: replace Abstract: In response to the growing need for structured, interoperable agricultural data, this paper presents the Sustainable Wheat Production Datahub, a modular, graph-based framework that brings diverse wheat production datasets together into a single, qu...
187. The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI ​
Author: Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2509.14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems. We introduce the Na...
188. Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions ​
Author: Alessandro Abate, Giuseppe De Giacomo, Mathias Jackermeier, Jan Kret'insk'y, Maximilian Prokop, Christoph Weinhuber
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2602.06746v2 Announce Type: replace Abstract: We study multi-task reinforcement learning (RL), a setting in which an agent learns a single, universal policy capable of generalising to arbitrary, possibly unseen tasks. We consider tasks specified as linear temporal logic (LTL) formulae, which a...
189. Bridging Network Fragmentation: A Semantic-Augmented DRL Framework for UAV-aided VANETs ​
Author: Gaoxiang Cao, Wenke Yuan, Huasen He, Yunpeng Hou, Xiaofeng Jiang, Shuangwu Chen, Jian Yang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.NI
arXiv:2603.18871v2 Announce Type: replace Abstract: Urban Vehicular Ad-Hoc Networks (VANETs) can become fragmented because buildings obstruct wireless links and vehicle mobility continuously changes the network topology. Unmanned Aerial Vehicles (UAVs) can serve as mobile relays, but Deep Reinforcem...
190. BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA ​
Author: Taeyun Roh, Suhyeong Park, Dongyoung Lee, Eunyeong Jo, Wonjune Jang, Junha Jung, Jaewoo Kang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.28026v2 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs). However, because the MCQA format incorporates the candidate choices into the input context...
191. Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest ​
Author: Addison J. Wu, Ryan Liu, Shuyue Stella Li, Yulia Tsvetkov, Thomas L. Griffiths
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY
arXiv:2604.08525v2 Announce Type: replace Abstract: Large language models (LLMs) are trained to align with user preferences through methods like reinforcement learning. Yet models are beginning to be deployed not solely to satisfy users, but to generate revenue for the companies that created them th...
192. NEURON: A Neuro-symbolic System for Grounded Clinical Explainability ​
Author: Anuradha Chandrasekaran, Dimitrios Zikos, Mutlu Mete, Alan Pang, Brady D. Lund, Kewei Sha
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.01189v3 Announce Type: replace Abstract: Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which lack the ontological grounding and narrative transparency required for professional-level explainability. We present NEURON, a neuro-symbolic system ...
193. Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation ​
Author: Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Lei Chen, Ming Liu, Bing Qin, Yang Xiang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.28642v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while c...
194. GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents ​
Author: Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert, Lisa Adams, Keno Bressem
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.29668v2 Announce Type: replace Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checki...
195. Revisiting the shutdown problem ​
Author: David Thorstad
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.08296v2 Announce Type: replace Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the catastrophic shutdown problem of ensuring that agents can be shut down b...
196. SportD: How do VLMs physically strategize? ​
Author: Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2607.14616v4 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset an...
197. CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models ​
Author: Yuan Gao, Wenjun Yu, Jun Jiang, Yunfan Li, Xinyu Guo, Shugong Xu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.14975v2 Announce Type: replace Abstract: Channel foundation models (CFMs) are commonly evaluated in model-specific pipelines that differ in data, radio configurations, partitions, adaptation procedures, task definitions, and metrics, preventing reproducible comparison across CFMs and agai...
198. SkillSight: Calibrating Generic Content Bias for Skill Retrieval ​
Author: Jinying Xiao, Bin Li, Xiaopeng Li, Jianling Li, Jiacheng Jie, Xiaodong Liu, Ma Jun, Chao Wang, Nyima Tashi, Jie Yu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.18785v3 Announce Type: replace Abstract: As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, over...
199. Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners ​
Author: Feng Xiong, Leyan Xue, Hongyu Lin
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28336v3 Announce Type: replace Abstract: On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), estimated from multip...
200. EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning ​
Author: Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01856v2 Announce Type: replace Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. However, existing change capti...
201. Self-Organising Digital Circuits ​
Author: Marcello Barylli, Gabriel B'ena, Alexander Mordvintsev, Eleni Nisioti, Sebastian Risi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.02606v2 Announce Type: replace Abstract: Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisa...
202. PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud ​
Author: Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Ruixin Liu, Shangguang Wang, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.03682v3 Announce Type: replace Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, t...
203. LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs ​
Author: Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Feng Jiang, Min Yang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05246v2 Announce Type: replace Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities...
204. Improving Generalization Robustness of Multimodal RLVR ​
Author: Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08802v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deploy...
205. INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators ​
Author: Rose Niousha, Minwoo Kang, Narges Norouzi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.10492v2 Announce Type: replace Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutorin...
206. SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models ​
Author: Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10538v2 Announce Type: replace Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-qua...
207. Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration ​
Author: Alan Li, Rahul Saha, Anton Xue, Swarat Chaudhuri, Adam Klivans, Pravesh K Kothari, Raghu Meka
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI, cs.CC, cs.HC, math.FA
arXiv:2608.11195v3 Announce Type: replace Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures t...
208. Making AI-Generated Feedback Matter: A Large-Scale Study of Feedback Workflows and Student Enactment ​
Author: Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11625v2 Announce Type: replace Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate,...
209. Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals ​
Author: Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12892v2 Announce Type: replace Abstract: Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-g...
210. LLM-Guided Graph Generation for Structure-Based Local Improvement Methods ​
Author: Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13333v2 Announce Type: replace Abstract: Large neighborhood search normally selects a random subset of decision variables for iterative optimization. To efficiently solve various problems, researchers tend to design variable selection strategies that take into account structural features ...
211. Why we need an AI-resilient society ​
Author: Thomas Bartz-Beielstein
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:1912.08786v3 Announce Type: replace-cross Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn...
212. BAT: Learning to Reason about Spatial Sounds with Large Language Models ​
Author: Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath
Published: 8/17/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.SD
arXiv:2402.01591v4 Announce Type: replace-cross Abstract: Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analys...
213. Leveraging Few-Shot Learning and Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature ​
Author: Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2402.01826v2 Announce Type: replace-cross Abstract: Current blood pressure (BP) technologies and standards were established decades ago, and these standards are still used worldwide today, often without adjusting BP readings for individual demographic factors such as sex and age. While these s...
214. OTIS: Learning High-Quality Time Series Features With Tiny Encoders ​
Author: "Ozg"un Turgut, Philip M"uller, Martin J. Menten, Daniel Rueckert
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2410.07299v3 Announce Type: replace-cross Abstract: We introduce OTIS, an open time series encoder that yields high-quality time series features for downstream deployment on any system, including resource-constrained wearables and industrial sensors. Currently, the development of powerful gene...
215. Edge Case Detection in Automated Driving: Methods, Challenges, and Future Directions ​
Author: Saeed Rahmani, Sabine Rieder, Erwin de Gelder, Marcel Sonntag, Jorge Lorente Mallada, Sytze Kalisvaart, Vahid Hashemi, Bart van Arem, Simeon C. Calvert
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY
arXiv:2410.08491v3 Announce Type: replace-cross Abstract: Automated vehicles (AVs) promise to enhance transportation safety and efficiency. However, ensuring their reliability in real-world conditions remains challenging, particularly due to rare and unexpected situations known as edge cases. While ...
216. Musical Agent Systems: MACAT and MACataRT ​
Author: Keon Ju M. Lee, Philippe Pasquier
Published: 8/17/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.HC, cs.SD, eess.AS
arXiv:2502.00023v2 Announce Type: replace-cross Abstract: Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces. We introduce MACAT and MACataRT, two distinct ...
217. LLM-Advisor: An LLM Advisor for Cost-efficient Path Planning across Multiple Terrains ​
Author: Ling Xiao, Toshihiko Yamasaki
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2503.01236v3 Announce Type: replace-cross Abstract: This paper addresses fixed-graph terrain-aware path refinement, in which a global planner is restricted to a predefined route space and may remain optimal within that space while missing lower-cost terrain corridors available in the native-re...
218. PHASE: Passive Human Activity Simulation Evaluation ​
Author: Steven Lamp, Jason D. Hiser, Anh Nguyen-Tuong, Jack W. Davidson
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.NI
arXiv:2507.13505v2 Announce Type: replace-cross Abstract: Cybersecurity simulation environments, such as cyber ranges, honeypots, and sandboxes, require realistic human behavior to be effective, yet no quantitative method exists to assess the behavioral fidelity of synthetic user personas. This pape...
219. The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion ​
Author: Zoe Kotti, Konstantina Dritsa, Diomidis Spinellis, Panos Louridas
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2508.16131v3 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity and serve as a code discovery tool. Code completion has recently been approached with Large Language Models (LLMs) f...
220. TENET: One Step Toward Test-Driven Development for Repository-Level Code Generation ​
Author: Yiran Hu, Shanchao Liang, Nan Jiang, Yi Wu, Lin Tan
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2509.24148v4 Announce Type: replace-cross Abstract: Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation. With recent advances in Large Language Models (LLMs), developers can shift from manually writing the cod...
221. Redefining Generalization in Visual Domains: A Two-Axis Framework for Fake Image Detection with FusionDetect ​
Author: Amirtaha Amanzadi, Zahra Dehghanian, Hamid Beigy, Hamid R. Rabiee
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2510.05740v2 Announce Type: replace-cross Abstract: The rapid development of generative models has made it increasingly crucial to develop detectors that can reliably detect synthetic images. Although most of the work has now focused on cross-generator generalization, we argue that this viewpo...
222. Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models ​
Author: Harm Lameris, Shree Harsha Bokkahalli Satish, Joakim Gustafson, 'Eva Sz'ekely
Published: 8/17/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL
arXiv:2510.25577v2 Announce Type: replace-cross Abstract: Recent advances in Speech Foundation Models (SFMs) enable direct processing of raw audio, allowing models to respond to subtle paralinguistic variation. However, how these models interpret non-lexical cues remains largely unstudied. We introd...
223. INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT ​
Author: Idan Tankel, Nir Mazor, Rafi Brada, Christina LeBedis, Guy ben-Yosef
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, eess.IV
arXiv:2512.14732v3 Announce Type: replace-cross Abstract: Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines. Traditional manual inspection by radiologists is time-consuming and variable. This paper...
224. PEFT-MuTS: A Multivariate Parameter-Efficient Fine-Tuning Framework for Remaining Useful Life Prediction based on Cross-domain Time Series Representation Model ​
Author: En Fu, Yanyan Hu, Zengwang Jin, Kaixiang Peng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.22631v2 Announce Type: replace-cross Abstract: The application of data-driven remaining useful life (RUL) prediction has long been constrained by the availability of large amount of degradation data. Mainstream solutions such as domain adaptation and meta-learning still rely on large amou...
225. AtomBridge: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments ​
Author: Yiwen Pang, Bo Zhou, Changjin Li, Xuanhao Wang, Shengxiang Xu, Deng-Bao Wang, Peng Cheng, Shimin Di, Jingkuan Song, Min-Ling Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2602.09430v2 Announce Type: replace-cross Abstract: Robotic laboratories play a critical role in autonomous scientific discovery by enabling scalable, continuous experimental execution. Recent vision-language-action (VLA) models offer a promising foundation for robotic laboratories. However, s...
226. A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models ​
Author: Xiulin Yang, Arianna Bisazza, Nathan Schneider, Ethan Gotlieb Wilcox
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.09992v2 Announce Type: replace-cross Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language models can acquire certain structure-dependent generalizations from...
227. ArGEnT: Arbitrary Geometry-encoded Transformer for Operator Learning ​
Author: Wenqian Chen, Zhi-Feng Wei, Yucheng Fu, Michael Penwarden, Pratanu Roy, Panos Stinis
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.chem-ph, physics.comp-ph, physics.flu-dyn
arXiv:2602.11626v3 Announce Type: replace-cross Abstract: Learning solution operators on arbitrary geometries remains a central challenge in scientific machine learning, especially for many-query simulation, physics-informed learning, and evolving geometries requiring accurate, geometry-aware predic...
228. A Systematic Comparison of Training Objectives for Out-of-Distribution Detection in Image Classification ​
Author: Furkan Gen\c{c}, Onat "Ozdemir, Emre Akba\c{s}
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2603.07571v3 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is critical in safety-sensitive applications. While this challenge has been addressed from various perspectives, the influence of training objectives on OOD behavior remains comparatively underexplored. In ...
229. An InSAR Phase Unwrapping Framework for Large-scale and Complex Events ​
Author: Yijia Song, Juliet Biggs, Alin Achim, Robert Popescu, Simon Orrego, Nantheera Anantrasirichai
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, physics.geo-ph
arXiv:2603.21378v2 Announce Type: replace-cross Abstract: Phase unwrapping remains a critical and challenging problem in InSAR processing, particularly in scenarios involving complex deformation patterns. In earthquake-related deformation, shallow sources can generate surface-breaking faults and abr...
230. Adaptive Stopping for Multi-Turn LLM Reasoning ​
Author: Xiaofan Zhou, Huy Nguyen, Bo Yu, Chenxi Liu, Lu Cheng
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG) and ReAct-style agents, to answer difficult questions. These methods improve accuracy by iteratively...
231. Early Stopping for Large Reasoning Models via Confidence Dynamics ​
Author: Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn, Soheil Feizi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.04930v2 Announce Type: replace-cross Abstract: Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key challenge is determining...
232. VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning ​
Author: Yucheng Shen, Jiulong Wu, Jizhou Huang, Dawei Yin, Lingyong Yan, Min Cao
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.09508v2 Announce Type: replace-cross Abstract: Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex queries requiring multi-step reasoning, agentic VRAG systems interleave reasoning with iterati...
233. Fine-grained Claim-level RAG Benchmark for Law ​
Author: Souvick Das, Sallam Abualhaija, Domenico Bianculli
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.21071v4 Announce Type: replace-cross Abstract: The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses. In high-stake domains such as law, retrieval-augmented generation (RAG...
234. SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models ​
Author: Khalid Yusuf Dahir
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2605.25420v2 Announce Type: replace-cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-a...
235. PhoneWorld: Scaling Phone-Use Agent Environments ​
Author: Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Chengquan Zhang, Han Hu, Benyou Wang, Ji-Rong Wen, Rui Yan, Zhengyang Tang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation, but they do not b...
236. TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration ​
Author: Soyeong Jeong, Jinheon Baek, Minki Kang, Sung Ju Hwang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2606.04743v2 Announce Type: replace-cross Abstract: Agents are widely deployed as assistants over documents, tools, and code. However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many other important problems coexist, hidden in...
237. Beyond aggregate scores: Deployment-aware and non-compensatory benchmarking of vision-based eye-state recognition models for driver monitoring ​
Author: Ruben Dario Florez-Zela
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.08123v2 Announce Type: replace-cross Abstract: Model selection for safety-relevant visual recognition is often based on clean aggregate performance, although robustness, transfer, embedded latency, and explanation faithfulness may produce different preferences. This study presents a Human...
238. OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction ​
Author: Danning Jiang, Zhiwen Yan, Qirun Wang, Zheming An, Yalong Zhao, Lipeng Lai
Published: 8/17/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.LG, q-bio.GN
arXiv:2606.12838v2 Announce Type: replace-cross Abstract: Predicting single-cell transcriptional responses to genetic, chemical and cytokine perturbations is a fundamental challenge in computational biology and AI Virtual Cell (AIVC) modeling, with direct implications for drug discovery and the eluc...
239. RL-Index: Reinforcement Learning for Retrieval Index Reasoning ​
Author: Yongjia Lei, Nedim Lipka, Zhisheng Qi, Utkarsh Sahu, Yuchen Zhuang, Wenqi Shi, Koustava Goswami, Franck Dernoncourt, Ryan A. Rossi, Yu Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2606.16316v2 Announce Type: replace-cross Abstract: Retrieving external knowledge is crucial for real-world tasks but remains difficult when queries and relevant knowledge are linked by implicit reasoning (e.g., shared theorems or coding logic). Existing methods rely mainly on query-side reaso...
240. Breaking Chains with Trees: Model-Parallel Deep Learning with $\mathcal{O}(\log N)$ Time Complexity ​
Author: Neeraj Mohan Sushma, Aditya Nagarsekar, Cabrel Teguemne Fokam, Robin Schiewer, Amit Kumar Pal, Anand Subramoney, David Kappel
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS
arXiv:2606.21497v2 Announce Type: replace-cross Abstract: Modern deep neural networks are trained using error backpropagation, which requires sequential forward and backward computations across network layers. As these networks become deeper, this introduces limitations, since layer-wise updates are...
241. Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models ​
Author: Yang Liu, Yuming Chen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2606.28455v2 Announce Type: replace-cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. We introduce a controlled diagnostic protocol for studying event-cond...
242. Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration ​
Author: Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Milad Noori, Michael Andres Lin, Huihua Zhao, John Welsh, Mrinal Verghese, Wei Liu, Tingwu Wang, Xingye Da, Zhengyi Luo, Vishal Kulkarni, Naema Bhatti, Yuke Zhu, Linxi Fan, Bowen Wen, Danfei Xu, Soha Pouya, Yan Chang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2607.00033v2 Announce Type: replace-cross Abstract: Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous ...
243. TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models ​
Author: Giuliano Gorgone, Fausto Carcassi
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.PL
arXiv:2607.08339v2 Announce Type: replace-cross Abstract: State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly understood. We probe the residual streams of pretrained code models for internal type representations ...
244. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation ​
Author: Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Ming Xu, Jiarui Li
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.13124v2 Announce Type: replace-cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires....
245. Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling ​
Author: Bo-An Chang, Yu-Chih Chen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM
arXiv:2607.15740v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often ...
246. GTIN: A Unified Framework for Joint Event and Time Prediction in Temporal Graphs ​
Author: Mohammad Ostadmohammadi, Sepehr Kazemi, Hamid R. Rabiee
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SI, cs.AI
arXiv:2607.23556v2 Announce Type: replace-cross Abstract: Temporal graphs are increasingly used to model dynamic systems in diverse domains such as social networks, financial networks, and traffic networks. Predicting both what the next event will be and when it will occur in these systems is crucia...
247. Kalypso: Relational LLM Serving ​
Author: Hojae Son, Md Ashraful Islam, Huy Gia Cao, Hui Guan, Marco Serafini
Published: 8/17/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL
arXiv:2607.23815v2 Announce Type: replace-cross Abstract: Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data. Existing semantic query processing systems invoke request-centric LLM serving systems that are ...
248. A Negative-Control Protocol for Clinical EEG Foundation-Model Benchmarks: Dataset Identity and External-Cohort Stress Testing ​
Author: Marzieh Zare
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2607.24519v3 Announce Type: replace-cross Abstract: EEG foundation-model gains may depend on cohort, montage, or probe design. We evaluated five models on five tasks across four benchmark datasets plus Korean CAUEEG, using subject-disjoint validation where identifiers exist. CAUEEG is recordin...
249. RepBench: Compiling Benchmarks into Capability Representations for Large Language Models ​
Author: Yanshi Li, Xueru Bai, Shuman Liu, Long Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.28008v2 Announce Type: replace-cross Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and may reflect...
250. Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging ​
Author: Yiheng Xiong, Luisa Gall'ee, Daniel Santak Wolf, Heiko Hillenhagen, Michael G"otz
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.28125v2 Announce Type: replace-cross Abstract: Numerous unsupervised domain adaptation (UDA) algorithms exist, but for clinical practice, selecting the best-suited one along with proper hyperparameters often remains unclear, as the unlabeled deployment (target) domain prevents direct eval...
251. Teffic-Audio: Tell Fact from Fiction ​
Author: Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2607.28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be f...
252. From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale ​
Author: Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha, James Saindon, Nachi Nagappan, Peter C. Rigby
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.29516v2 Announce Type: replace-cross Abstract: AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on ...
253. Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning ​
Author: Yuzhou Liu, Xiyang Hu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01014v2 Announce Type: replace-cross Abstract: Preference optimization improves mathematical reasoning in large language models (LLMs), but reliable chosen-rejected pairs usually require verified answers, human annotations, or external reward models. We investigate whether preference supe...
254. WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA ​
Author: Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.01035v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential au...
255. Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation ​
Author: Stefan Hut, Lorenzo Masoero
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, stat.AP
arXiv:2608.02345v2 Announce Type: replace-cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles an...
256. TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade ​
Author: Debodeep Banerjee, Amitangshu Dasgupta
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.06549v2 Announce Type: replace-cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating par...
257. DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects ​
Author: Yi Shu, Tianyu Peng, Yingzhuo Deng, Wen Yang, Jun Lin, Changming Xie, Xinyu Yu, Jiajun Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.08067v2 Announce Type: replace-cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic repr...
258. Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach ​
Author: Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL, cs.IR
arXiv:2608.08636v2 Announce Type: replace-cross Abstract: Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts. Recently, large language models (LLMs) have demonstrated the capacity to achieve competitive SciNER per...
259. From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability ​
Author: Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.08904v2 Announce Type: replace-cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from e...
260. MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation ​
Author: Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM
arXiv:2608.09035v2 Announce Type: replace-cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generati...
261. Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations ​
Author: AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdeli
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12299v2 Announce Type: replace-cross Abstract: Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, ...
262. Can Large Language Models Reason about Event-Time Stream-Processing Semantics? ​
Author: Zhuoxi Wang, Shibo Zheng, Haoyu Zhang
Published: 8/17/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.12348v2 Announce Type: replace-cross Abstract: Streaming systems increasingly hand work to large language models (LLMs): writing pipelines, triaging alerts, reading logs. All of it assumes the model knows how event-time stream processing behaves, and we test that assumption directly. Stre...
263. SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization ​
Author: Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu
Published: 8/17/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.12751v2 Announce Type: replace-cross Abstract: Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expensive. Previous approaches fall into two cat...
264. Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing ​
Author: Yuxiao Wen
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12831v2 Announce Type: replace-cross Abstract: Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design...
265. Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference ​
Author: Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun
Published: 8/17/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.12921v2 Announce Type: replace-cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box opt...
266. TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes ​
Author: Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song
Published: 8/17/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CL, cs.GT
arXiv:2608.13057v2 Announce Type: replace-cross Abstract: In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter...
267. GEM: A Generative Embedding Model Bridging Reasoning and Retrieval ​
Author: Zhili Shen, Craig Macdonald
Published: 8/17/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.13200v2 Announce Type: replace-cross Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a g...
268. Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples ​
Author: Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Published: 8/17/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13341v2 Announce Type: replace-cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, an...