arXiv cs.AI - 2026-08-15 ​
299 items collected.
1. Position: Reasoning is a Learnable Rule-Based Process ​
Author: Rachel Lawrence, Jacqueline Maasch
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and r...
2. Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists ​
Author: Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action re...
3. Position: The Alignment Community is Unintentionally Building a Censor's Toolkit ​
Author: Sarah Ball, Phil Hackemann
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.12346v1 Announce Type: new Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniq...
4. Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments ​
Author: Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same ju...
5. Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing ​
Author: Sabeur Lajili, Zaki Brahmi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12371v1 Announce Type: new Abstract: Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper ...
6. Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning ​
Author: Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that r...
7. Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese ​
Author: Rian Touchent (ALMAnaCH)
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12373v1 Announce Type: new Abstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's deci...
8. Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation ​
Author: Liming Liu, Mingze Wang, Tuo Zhao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12385v1 Announce Type: new Abstract: As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound,...
9. Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization ​
Author: Xuefei Wang, Jun Han, Zixuan Wang, Qingkai Zeng, Xiao Wang, Ruijie Wang, Jianxin Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse...
10. Research Assistant: AstraZeneca's Agentic System for R&D ​
Author: Piotr Grabowski, Mohamed Alameen, Jorge Bretones, Sabina Cardell, Miguel Carmona, Gavin Edwards, Ben Grainger, Sameh Hassan, Erik Jansson, Artur Kuziakhmetov, Albert Maristany, Hebatallah Mohamed, Andriy Nikolov, Sebastian Nilsson, Mark O'Donoghue, James Pacileo, Ashiq Sultan, Alex Voegele, Michael Ughetto
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12395v1 Announce Type: new Abstract: We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together e...
11. Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction ​
Author: Mariya I. Vasileva
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the composit...
12. MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents ​
Author: Kaichao Liang, Yuqi Cui, Hao Kong, Xinyuan Huang, Guohaotian Hou, Qingcan Kang, Liang Chen, Yiyang Yin, Ke Ye, Jiaquan Guo, Da Chen, Lingan Zeng, Yixing Peng, Rong Yao, Shixiong Kai, Mingxuan Yuan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.IT, math.IT
arXiv:2608.12428v1 Announce Type: new Abstract: Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt the...
13. Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents ​
Author: Guodong Xu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12476v1 Announce Type: new Abstract: Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an...
14. $\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution ​
Author: Aofan Liu, Shiyuan Song, Yiyan Qi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12522v1 Announce Type: new Abstract: LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEv...
15. CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence ​
Author: Michael Georgiades, Charalambia Varnava
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12555v1 Announce Type: new Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS s...
16. Trie Automata for Constrained Decoding over Large Finite Sets ​
Author: Xingzi Xu, Karim Bouyarmane
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.FL
arXiv:2608.12574v1 Announce Type: new Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-p...
17. Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces ​
Author: Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model ...
18. Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting ​
Author: Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a cli...
19. DiG-bench: Discovery in Games ​
Author: Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, J"urgen Schmidhuber, Michal Valko, Joshua Tenenbaum, Thomas L. Griffiths, Zeb Kurth-Nelson, James C. R. Whittington
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge w...
20. Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues ​
Author: Haoyuan Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behav...
21. @skills: Attention is all you have ​
Author: Li Yin (Atlas), Zhi Li (Atlas), Zhan Shi (Atlas), Haoran Zhang (Atlas), Haebin Seong (Atlas), Zhangyang (Atlas), Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This ...
22. Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence ​
Author: Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, o...
23. SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries ​
Author: Oguz Serdar, Cuneyt Mertayak
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBen...
24. General Probabilities of Causation with Causal Knowledge ​
Author: Xin Shu, Zhen Lei, Ang Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, stat.ML
arXiv:2608.12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the p...
25. Designing AI Pipelines for Decision-Ready ITSM Intelligence ​
Author: Archan Dutta, Yash Dharmadhikari, Marat Valiullin, Rahul Guha, Alexander Liss
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12670v1 Announce Type: new Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and e...
26. On the Expressive Power of Transformers ​
Author: Phokion Kolaitis, Rik Sengupta
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CC
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expr...
27. Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy ​
Author: Ravi Teja Chunduri, Srikaran Reddy Boya, Deep Narayan Mishra, Ajay Kumar B, Karthik Kumaran, Pranay Kona
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across ...
28. Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs ​
Author: Saleh Almohaimeed, Saad Almohaimeed, Mousa Jari, Fahad Alotaibi, Khalid A. Alobaid
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG has focused on preventing unauthorized users from accessing sensitive data. Howev...
29. The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis ​
Author: Danial Sharifrazi, Saadat Behzadi, Julakha Jahan Jui, Mojtaba Mohammadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make ...
30. Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies ​
Author: Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer, Roberto Dailey, Babak Hodjat, Risto Miikkulainen, Xin Qiu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can ...
31. Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence ​
Author: Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, su...
32. Correct Is Not Governed: Provenance Integrity in Agentic Workflows ​
Author: Jesus Salas
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.12761v1 Announce Type: new Abstract: Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later c...
33. PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs ​
Author: Sadat Shahriyar, Shareef Ahmed, Abdullah Al Arafat
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12762v1 Announce Type: new Abstract: Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alte...
34. ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs ​
Author: Jiale Cui, Yueyao Yuan, Kaixi Zhong, Xiaogang Xu, Jiafei Wu, Zhe Liu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research...
35. CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation ​
Author: Yuchen Liu, Zongzhen Yang, Binhang Qi, Hailong Sun, Xiang Gao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12842v1 Announce Type: new Abstract: Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade me...
36. Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories ​
Author: Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a dist...
37. Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents ​
Author: Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operatio...
38. AI and Consumer Rights in India Working Paper ​
Author: Omir Kumar, Sriya Sridhar, Vibhav Mithal, Balaraman Ravindran
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI pro...
39. ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification ​
Author: Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12877v1 Announce Type: new Abstract: Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to dec...
40. Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals ​
Author: Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12892v1 Announce Type: new Abstract: Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid ...
41. Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence ​
Author: Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.12895v1 Announce Type: new Abstract: Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent hando...
42. Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence ​
Author: Jakub Pokrywka, {\L}ukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12928v1 Announce Type: new Abstract: We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containi...
43. FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving ​
Author: Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12932v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a c...
44. Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses ​
Author: Lei You
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12935v1 Announce Type: new Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactu...
45. Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$ ​
Author: Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12961v1 Announce Type: new Abstract: The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has no...
46. OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways ​
Author: Mao Jiayang, Wang Lanfeng, Peng Zhao-Han
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.12995v1 Announce Type: new Abstract: Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that ...
47. Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic ​
Author: Costin B\u{a}dic\u{a}, Amelia B\u{a}dic\u{a}
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13018v1 Announce Type: new Abstract: Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions....
48. From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion ​
Author: Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we...
49. BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs ​
Author: Sanjeev Manivannan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.ET
arXiv:2608.13046v1 Announce Type: new Abstract: Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the syste...
50. DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition ​
Author: Amogh Joshi, Animesh Mukherjee, Sergey Utyuzhnikov
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13048v1 Announce Type: new Abstract: In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM in...
51. VALG: An Agentic System for ML Theory Research ​
Author: Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.OC, stat.ML
arXiv:2608.13060v1 Announce Type: new Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem th...
52. Uniform Herding: Exemplar Replay with Representation Refresh ​
Author: Krishna Subedi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13061v1 Announce Type: new Abstract: As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded ...
53. Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) ​
Author: Sam Mao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.13063v1 Announce Type: new Abstract: Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self...
54. Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds ​
Author: Lucia Mal'i\v{c}kov'a
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavio...
55. EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding ​
Author: Shuailei Zhang, Muyun Jiang, Wei Zhang, Jinbo Chen, Zhiwei Guo, Yong Li, Yi Ding, Cuntai Guan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13072v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset mul...
56. SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference ​
Author: Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with...
57. Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment ​
Author: Gupta Lovi Raj, Kaur Kamalpreet, Dama Sri Ram, Parani Prajithaa
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC
arXiv:2608.13100v1 Announce Type: new Abstract: Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character ...
58. Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting ​
Author: Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13108v1 Announce Type: new Abstract: Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion ...
59. SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback ​
Author: Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13120v1 Announce Type: new Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but d...
60. Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement ​
Author: Aoxin Ni
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13129v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey exam...
61. Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing ​
Author: Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13156v1 Announce Type: new Abstract: Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, e...
62. SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents ​
Author: Chang Liu, Yuqi Zhang, Yiman Zhong, Boyi Liu, Hengjun Wang, Shuyue Wei
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13173v1 Announce Type: new Abstract: Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with...
63. Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents ​
Author: Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, Leilei Gan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13179v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signa...
64. TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems ​
Author: Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior proc...
65. Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test ​
Author: Saveliy Batruin
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restr...
66. vToken: Token-Level Virtualization for Reclaimable KV Caches ​
Author: Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.OS
arXiv:2608.13263v1 Announce Type: new Abstract: Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at ...
67. Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability ​
Author: Alan Woodward, Andrew Rogoyski
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.13272v1 Announce Type: new Abstract: A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both the legal power and the political will to decide which other countries may use these systems. In June 2026 the Unit...
68. Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision ​
Author: Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart Vanrumste, Benjamin Filtjens
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13283v1 Announce Type: new Abstract: Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vi...
69. NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space ​
Author: Eleftherios Mylonas, Angelos Kouprizas, Michael Birbas, Alexios Birbas
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13293v1 Announce Type: new Abstract: Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a challenge addressed by hardware-aware Neural Architecture Search (NAS). While recent works incorporate quantization...
70. StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems ​
Author: Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13317v1 Announce Type: new Abstract: Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that tok...
71. LLM-Guided Graph Generation for Structure-Based Local Improvement Methods ​
Author: Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13333v1 Announce Type: new Abstract: Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving different problems, researchers tend to design variable selection strategies by taking into account structural features...
72. LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning ​
Author: Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13344v1 Announce Type: new Abstract: Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language...
73. Rules or Character? Scaling Laws for AI Safety Design ​
Author: Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13345v1 Announce Type: new Abstract: Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, s...
74. TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies ​
Author: Xiaokang Qu, Jianliang Ma, Zao Fan, Tianshu Chu, Tianlong Fan, Linyuan L"u
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.NI
arXiv:2608.13389v1 Announce Type: new Abstract: Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. Existing NetOps automation tools mainly operate after t...
75. Jointly Predicting Courses and Grades Using a Transformer-Based Model ​
Author: Paul Savala
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13409v1 Announce Type: new Abstract: Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, partic...
76. Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings ​
Author: Mirko Tritella, Riccardo Pozzi, Matteo Palmonari
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13410v1 Announce Type: new Abstract: Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamen...
77. Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development ​
Author: Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13417v1 Announce Type: new Abstract: Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neit...
78. Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes ​
Author: Aimilios Hadjiliasi, Louis Nisiotis
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13420v1 Announce Type: new Abstract: Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and techn...
79. RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level ​
Author: Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13428v1 Announce Type: new Abstract: Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptati...
80. Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension ​
Author: Alison R. Panisson, Maria Eduarda W. M. Vianna, Italo Firmino da Silva, Heitor Henrique da Silva, Rafaela Fernandes Savaris, Bernardo Pandolfi Costa, Martin Augusto Gagliotti Vigil, Jim Lau, Agenor Hentz, Andr'ea Sabedra Bordin, Alexandre Leopoldo Gon\c{c}alves, Roberto Rodrigues-Filho
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13447v1 Announce Type: new Abstract: Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artificial ...
81. A Unifying Perspective on Causal World Models: From Observations to Representations to Structure ​
Author: Avinash Kori, Fabrizio Russo
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.13456v1 Announce Type: new Abstract: World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from...
82. MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination ​
Author: Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.13476v1 Announce Type: new Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reas...
83. AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) ​
Author: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.13492v1 Announce Type: new Abstract: This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are rep...
84. QuoteBench: How Matched Scores Can Hide Command-Path Failures ​
Author: Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.13547v1 Announce Type: new Abstract: LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this...
85. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist ​
Author: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.13558v1 Announce Type: new Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the fu...
86. The AI Accountability Ecosystem in the Era of Language Models ​
Author: Chris Percy, Artur d'Avila Garcez
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.12320v1 Announce Type: cross Abstract: This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose th...
87. LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning ​
Author: Yubo Li, Ramayya Krishnan, Rema Padman
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12321v1 Announce Type: cross Abstract: When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activat...
88. What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting ​
Author: Poli Nemkova, Haeshitha Indukuri
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12322v1 Announce Type: cross Abstract: Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic ...
89. Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance ​
Author: Mika Okamoto, Ansel Kaplan Erol, Kutluhan Erol
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations tes...
90. When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models ​
Author: Alex Chao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.12324v1 Announce Type: cross Abstract: People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among...
91. Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition ​
Author: Suman Paudel, Sarbin Sayami
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and...
92. AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement ​
Author: Guilherme C. Oliveira, Stephanie Fong, Zimu Wang, Clarice Lee, Xiangyu Zhao, Duy Khoa Pham, Duong Nhu, Yiwen Jiang, Jiahe Liu, Zhongxing Xu, Dwarikanath Mahapatra, Dominic Dwyer, Zongyuan Ge
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.12329v1 Announce Type: cross Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured...
93. Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching ​
Author: Yang Liu, Bin Chong, Chongyang Zhang, Hao Zheng, Jiayu Liang, Xu Kefu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12331v1 Announce Type: cross Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences an...
94. Vision-Language Models are Fragile Multilingual Associators ​
Author: Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language...
95. Steering the Language Axis: From Linear Decodability to Causal Control ​
Author: Arnav Srivastav
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from hidden states, or i...
96. StorySpark: Module-wise Evolutionary Search for Story Premise Generation ​
Author: Yang Yang, Zining Zhong, Qian Cao, Jindong Li, Boyun Xu, Kaishen Yuan, Menglin Yang, Yutao Yue
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12336v1 Announce Type: cross Abstract: A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stage planning, controllability, coherence, and prose expansion, while premise-level ideation remains comparatively...
97. Mimicry without understanding: the origins of decision bias in large language models ​
Author: Eldad Yechiam, Adi Tarabeih
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or...
98. StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics? ​
Author: Zhuoxi Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.12348v1 Announce Type: cross Abstract: Streaming systems increasingly hand work to large language models (LLMs) -- writing pipelines, triaging alerts, reading logs -- and all of it assumes the model knows how event-time stream processing behaves. We test that assumption head-on. StreamRea...
99. From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks ​
Author: Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He, Angel Hsu, Michael P. Vandenbergh
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.12350v1 Announce Type: cross Abstract: The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side manageme...
100. Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection ​
Author: Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom)
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort. This paper reports lessons from designing and imple...
101. Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF ​
Author: Joseph R. Simons, David A. Broniatowski
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.12352v1 Announce Type: cross Abstract: AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in ...
102. Humans are Missing from AI Coding Agent Research ​
Author: Zora Z. Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei, Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SE
arXiv:2608.12355v1 Announce Type: cross Abstract: Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows. As these systems ma...
103. Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio ​
Author: Sherzod Turaev, Saja Aldabet, Mary John, Namya Musthafa, Mamoun Awad, Nazar Zaki, Khaled Shuaib
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.12356v1 Announce Type: cross Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is dif...
104. Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles ​
Author: Sudhir Alladi Venkatesh
Published: 8/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.12358v1 Announce Type: cross Abstract: Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Read...
105. EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector ​
Author: Javier Gonzalez-Ruiz, Carlos Rodriguez-Pardo, Alice Di Bella, Paolo Mastropietro, Jose Pablo Chavez-Avila, Massimo Tavoni
Published: 8/15/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.CY, cs.LG, cs.MA, cs.SY, eess.SY, q-fin.EC
arXiv:2608.12363v1 Announce Type: cross Abstract: European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. A notable example is Italy's 2026 Decreto Bollette package, which proposes to...
106. FluctlightDB: A Memory Model of Data for AI Agents ​
Author: Ganesh S
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.12365v1 Announce Type: cross Abstract: For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model asked which vectors lie nearest a query. Neither was built for cue-driven, provenance-weighted recall across long ...
107. Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities ​
Author: Hanna Abi Akl, Fabien Gandon, Catherine Faron, Pierre Monnin
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12374v1 Announce Type: cross Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation motivates our study...
108. From Observation to Intervention: Memory in Brains and Large Language Models ​
Author: Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer, Jia Wu
Published: 8/15/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CL
arXiv:2608.12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader associations, how new ...
109. Query Timing Produces Opposite Positional Biases Between LLMs and Humans ​
Author: Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12387v1 Announce Type: cross Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been...
110. Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models ​
Author: Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings. However, ex...
111. A Hierarchical Energy-Based Model for Multimodal Cognition ​
Author: Subir Varma
Published: 8/15/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2608.12398v1 Announce Type: cross Abstract: We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language. Followi...
112. SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents ​
Author: Ruitao Wang, Yuwen Hao, Menglin Yang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.12429v1 Announce Type: cross Abstract: Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data synthesis methods reduce manual annotation, but they still face two key limitations: they often fail to cover the...
113. Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review ​
Author: Joel Abenhaim
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.12440v1 Announce Type: cross Abstract: This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the ta...
114. Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection ​
Author: Iyad Assaad Nekka, Hamida Seba, Khaled Walid Hidouci, Karima Amrouche
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12441v1 Announce Type: cross Abstract: Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information sys...
115. SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization ​
Author: Yuanyu Li, Jintao Xu, Zijiang Liu, Yongzhi Qi, Ningxuan Kang, Jianshen Zhang, Wei Qi, Chen Xie, Zuo-Jun Max Shen
Published: 8/15/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.OC
arXiv:2608.12443v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single bes...
116. Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts ​
Author: Seyyed Ali Hoseini, Javad Baseri, Hamid Saadatfar, Edris Hoseini Gol, AmirHossein Eshghi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12446v1 Announce Type: cross Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. This study investigates wheth...
117. SchemaLink: An Intelligent Web Editor for LinkML Schema Curation ​
Author: Emanuele Cavalleri, Paolo Perlasca, J. Harry Caufield, Justin Reese, Christopher J. Mungall, Marco Mesiti
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.HC
arXiv:2608.12529v1 Announce Type: cross Abstract: Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical data. Even if it is a quite recent proposal, it has been applied in several biomedical contexts. Developing an...
118. Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling ​
Author: Nadia Mehjabin, Henry Kautz, Subigya Nepal
Published: 8/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.12582v1 Announce Type: cross Abstract: AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intentio...
119. What Makes a Peer? Valuation-Anchored Similarity in Private Markets ​
Author: Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta
Published: 8/15/2026, 4:00:00 AM
Categories: q-fin.ST, cs.AI, cs.LG
arXiv:2608.12594v1 Announce Type: cross Abstract: As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due dili...
120. Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks ​
Author: Andrew Cheng, Ali Eslamian, Jie Cheng, Mehdi Zargham, Qiang Cheng
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12597v1 Announce Type: cross Abstract: Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the late...
121. PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping ​
Author: Chikao Tsuchiya, Dhaval Bhanderi, David Ilstrup, Hsinmin Cheng, Christopher Ostafew
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12600v1 Announce Type: cross Abstract: A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, which limits model generalization in diverse environments. To address this limitation, we propose a teacher-student...
122. LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning ​
Author: Yi Wu, Zhimin Hu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2608.12626v1 Announce Type: cross Abstract: Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources prevent the model from maintaining strategic coherence over thousands of ste...
123. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory ​
Author: Le Zhang, Ke Sun
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.HC
arXiv:2608.12627v1 Announce Type: cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic se...
124. Novels generated by language models show compressed formal variation ​
Author: Mehdy Sedaghat Payam, Justin Quinn
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12630v1 Announce Type: cross Abstract: While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this stud...
125. Interpretable Causal Discovery via Causal-Effect Constraints ​
Author: Cixuan Zhang, Guy Van den Broeck, Benjie Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.12640v1 Announce Type: cross Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized...
126. Demand Transfer Estimation at Scale via Restricted Logit Modeling ​
Author: Lakshya Garg, Deep Narayan Mishra, Swapnil Yadav, Haoan Wang, Sujal Alugubelli, Karthik Kumaran, Anupriya Sharma
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12680v1 Announce Type: cross Abstract: Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e. expected demand) wi...
127. Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging ​
Author: Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12689v1 Announce Type: cross Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information...
128. Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks ​
Author: Xiaoyan Feng, Yanjun Zhang, He Zhang, Leo Yu Zhang, Shirui Pan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.12713v1 Announce Type: cross Abstract: Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerab...
129. HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement ​
Author: Zhengyi Lu, Aswini Sivakumar, Jie Hu, Yao Qiang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.12715v1 Announce Type: cross Abstract: Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schr"odinger Bridges (SB) shorten transport from noise to clean speech but lea...
130. Error-Aware Reverse Auction Mechanism for Large Language Model Routing ​
Author: Haolong Chen, Zhengyuan Xin, Liang Zhang, Lei Xue, Guangxu Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2608.12719v1 Announce Type: cross Abstract: Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bot...
131. ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval ​
Author: Haolong Chen, Liang Zhang, Zhuo Li, Lei Xue, Guanrxu Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12720v1 Announce Type: cross Abstract: While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogen...
132. PatientAct: Theory-Grounded Mental Health Client Simulation ​
Author: Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.12750v1 Announce Type: cross Abstract: LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes ...
133. SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization ​
Author: Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.12751v1 Announce Type: cross Abstract: Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expensive. Previous approaches fall into two categories:...
134. Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents ​
Author: Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12764v1 Announce Type: cross Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPS...
135. Memorization Diagnostics for Code LLMs Should be Scale-Aware ​
Author: Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend'e F. Bissyand'e
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.12771v1 Announce Type: cross Abstract: The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense ...
136. CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives ​
Author: Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.12779v1 Announce Type: cross Abstract: Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to...
137. PIPES: Securing Agent Perception with Provenance and Priors ​
Author: Sanjay Kariyappa, Severin Klingler, G. Edward Suh
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.12789v1 Announce Type: cross Abstract: Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or what it should convey. We show that this gap enables state-corruption attacks, in which attacker-co...
138. Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors ​
Author: Qiao Li, Xiaomeng Fu, Wangjia Yu, Runze He, Baisen Wang, Jiao Dai, Jizhong Han
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12806v1 Announce Type: cross Abstract: The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure:...
139. Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing ​
Author: Yuxiao Wen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12831v1 Announce Type: cross Abstract: Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design gives e...
140. From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options ​
Author: Obed Junias, Maria Leonor Pacheco
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12836v1 Announce Type: cross Abstract: Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducin...
141. AQuA: Recursively Self-Improving Quantitative Trading Research Agents ​
Author: Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12841v1 Announce Type: cross Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which...
142. Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval ​
Author: Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12843v1 Announce Type: cross Abstract: Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained re...
143. FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation ​
Author: Yuchen Zheng, Sihan Xu, Jingwen Yang, Xiangrui Cai, Haiwei Zhang, Xiaojie Yuan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2608.12845v1 Announce Type: cross Abstract: Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are...
144. Falsehood and Impossibility Are Different Directions in an AI's Representation of Language ​
Author: Yoon Pyo Lee
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12852v1 Announce Type: cross Abstract: Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal op...
145. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving ​
Author: Bing Zhan, Shuyao Shang, Jiahao Gu, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Lue Fan, Zhaoxiang Zhang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.12854v1 Announce Type: cross Abstract: Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM prio...
146. A Compositional Theory of Curvature in Probabilistic Circuits ​
Author: Hrithik Suresh, Sahil Sidheekh, Shelar Parth Vijay, Yasir Z, Sriraam Natarajan, Narayanan Chatapuram Krishnan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12869v1 Announce Type: cross Abstract: Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes ...
147. SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data ​
Author: Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12876v1 Announce Type: cross Abstract: Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenan...
148. Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation ​
Author: Rana Muhammad Ahmed (Department of Computer Science, Bahria University, Islamabad, Pakistan), Sabahat Abbas (Department of Computer Science, Bahria University, Islamabad, Pakistan)
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.12880v1 Announce Type: cross Abstract: Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatm...
149. NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents ​
Author: Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12898v1 Announce Type: cross Abstract: Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two...
150. EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction ​
Author: Danyu Li, Ling Zhou, Rubing Huang, Xian Zhong, Bin Zou, Kui Jiang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12906v1 Announce Type: cross Abstract: RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI P...
151. InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers ​
Author: Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to...
152. Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference ​
Author: Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun
Published: 8/15/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.12921v1 Announce Type: cross Abstract: The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimizatio...
153. H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities ​
Author: Julius Broermann, Oliver M"uller, Michael D"oring, Jochen Baumeister
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12926v1 Announce Type: cross Abstract: Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Action...
154. AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization ​
Author: Harshkumar Oza, Aritra Sarkar, Syed Naqi Abbas, Rahul Bhowmick, Aryan Prakash, Prateek P Kulkarni, Krishna Kumar Sabapathy
Published: 8/15/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET
arXiv:2608.12936v1 Announce Type: cross Abstract: As quantum computing progresses from proof-of-principle demonstrations toward practical utility, a significant impediment is the need to augment algorithmic feasibility with system-level optimization across heterogeneous hardware and software stacks....
155. The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use ​
Author: Joyjeet Singh
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12959v1 Announce Type: cross Abstract: Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objectiv...
156. Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents ​
Author: Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechanisms into the agent e...
157. Generative Universal Multimodal Retrieval with Dual-role Identifiers ​
Author: Kaipeng Li, Haitao Yu, Xuanchen Zhou
Published: 8/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly. Despite its promise, a numbe...
158. Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language ​
Author: Johan Henriksson
Published: 8/15/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI, cs.SE
arXiv:2608.13029v1 Announce Type: cross Abstract: The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dyna...
159. UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations ​
Author: Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13031v1 Announce Type: cross Abstract: Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain h...
160. Operationalizing Cyber Threat Intelligence with GraphRAG ​
Author: Atul Kabra, Prakhar Paliwal, Manjesh K. Hanawal
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.13050v1 Announce Type: cross Abstract: When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. In practice, most automated attempts at this only extract the simplest clues from the report --- bad IP addresse...
161. TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes ​
Author: Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CL, cs.GT
arXiv:2608.13057v1 Announce Type: cross Abstract: In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU gen...
162. LOB-ID: Evaluating Synthetic Market Data by Inception Distances ​
Author: Andreea Bacalum, Zhuohan Wang, Ollie Olby, Martin Garaj, Namid Stillman
Published: 8/15/2026, 4:00:00 AM
Categories: q-fin.CP, cs.AI, cs.CE
arXiv:2608.13082v1 Announce Type: cross Abstract: Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-lev...
163. Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization ​
Author: Jinhyung Bae
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2608.13087v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything ...
164. EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory ​
Author: Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that ...
165. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation ​
Author: Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB, cs.MA
arXiv:2608.13136v1 Announce Type: cross Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current eval...
166. LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service ​
Author: Zhengzhe Xiang, Yinlin Chen, Fuli Ying, Binbin Zhou, Hailiang Zhao, Schahram Dustdar
Published: 8/15/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.13144v1 Announce Type: cross Abstract: As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern. Existing semantic caching methods largely rely o...
167. Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering ​
Author: Yilin Wang, Yuchun Fan, Weidong Bao, Zili Wei, Shi Feng, Tong Xiao, Zhengtao Yu, Jingbo Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13160v1 Announce Type: cross Abstract: Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved documents into English or...
168. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint ​
Author: Fnu Pramono, John Cai, Sourabh Kulkarni
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedural...
169. GEM: A Generative Embedding Model Bridging Reasoning and Retrieval ​
Author: Zhili Shen, Craig Macdonald
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.13200v1 Announce Type: cross Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely on surface-level matching between queries and documents, resulting in a growing g...
170. NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video ​
Author: Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2608.13210v1 Announce Type: cross Abstract: Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilit...
171. CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport ​
Author: Peng Ling, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Zeyu Hu, Xin Wang, Wenming Yang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13226v1 Announce Type: cross Abstract: While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe computational bottlenecks during inference. Existing token pruning methods primaril...
172. Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales ​
Author: Long Hoang Nguyen, Brice Valentin Kok-Shun, Guangyu Du, Ali Sunyaev
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.13250v1 Announce Type: cross Abstract: Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level...
173. GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport ​
Author: Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Yutong Zhao, Zi Wang, Bo Liu, Huanrui Yang, Sen He
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13255v1 Announce Type: cross Abstract: Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy b...
174. Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data ​
Author: Francesca Pia Panaccione, Sofia Mongardi, Marco Masseroli, Pietro Pinoli
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13256v1 Announce Type: cross Abstract: As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data gene...
175. Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models ​
Author: Paras Balani, Subhrakanta Panda
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13258v1 Announce Type: cross Abstract: Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or h...
176. Into the ORBIT for Time Series: Training Regimes for Foundation Models ​
Author: Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13262v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often poorly controlled wi...
177. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures ​
Author: Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG
arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when ...
178. Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model ​
Author: Mohammed Sabry, Sean Augenstein, Keith Rush, Lucio Dery
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13277v1 Announce Type: cross Abstract: We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular pre-training procedure tha...
179. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks ​
Author: Wojciech Zarzecki, Jaros{\l}aw Arabas
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13296v1 Announce Type: cross Abstract: Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back even to the 1970s. This causes a risk of biasing the development of global optimization methods. We argue that...
180. Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy ​
Author: \v{Z}an Gorenc, \v{Z}iga Gradi\v{s}ar, Felix M"utter, Vanja Suboti'c, Pavle Bo\v{s}koski
Published: 8/15/2026, 4:00:00 AM
Categories: stat.AP, cs.AI
arXiv:2608.13305v1 Announce Type: cross Abstract: Estimating the distribution of relaxation times (DRT) fromelectrochemical impedance spectroscopy (EIS) is an ill-posed inverse problem that is highly sensitive to regularisation choices. We propose a physics-informed convolutional autoencoder that es...
181. Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services ​
Author: Ahmet Bugra Gundogan, Yigit Turkmen, Melih Bastopcu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.SY, eess.SY
arXiv:2608.13315v1 Announce Type: cross Abstract: We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but ...
182. It's How You Ask: Gender-Associated Linguistic Bias in LLMs ​
Author: Katherine Van Koevering, Anjalie Field
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13328v1 Announce Type: cross Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systemati...
183. Training AI Scientists to Replicate Research ​
Author: Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13331v1 Announce Type: cross Abstract: The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified,...
184. Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples ​
Author: Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is dif...
185. Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs ​
Author: Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13368v1 Announce Type: cross Abstract: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discri...
186. Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks ​
Author: Muhammad Hannan Akram, Muhammad Abubakar Rashid, Wassi Haider Kabir, Haejoon Jung, Kapal Dev, Syed Ali Hassan
Published: 8/15/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.MA
arXiv:2608.13394v1 Announce Type: cross Abstract: 6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where thousands of autonomous artificial intelligence (AI) agents are interconnected. The agents are deployed across ...
187. Deliberate Practice: Learning Robot Skills under a Budget ​
Author: Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, George Konidaris
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.13415v1 Announce Type: cross Abstract: We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, \emph{Deliberate Practice (DP)}, that computes a provably \emph{budget-optimal} allocat...
188. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference ​
Author: Zixuan Lan, Yanhong Li, Jiawei Zhou
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.13426v1 Announce Type: cross Abstract: Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference metho...
189. Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity ​
Author: Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13430v1 Announce Type: cross Abstract: Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the...
190. Algebraic Decomposition Theory for Transformer Length Generalization ​
Author: Andy Yang, Blerta Veseli, Corentin Barloy, Micha"el Cadilhac, Andreas Krebs, Charles Paperman, Howard Straubing, Michael Hahn
Published: 8/15/2026, 4:00:00 AM
Categories: cs.FL, cs.AI
arXiv:2608.13433v1 Announce Type: cross Abstract: Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise characterization of which tasks admit length generalization. It is not even known which regular languages transfo...
191. ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models ​
Author: Gehan Zheng, Matthew Johnson-Roberson, Weiming Zhi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.13438v1 Announce Type: cross Abstract: Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close gripper--object views help observe contact, but a poor approach may already push, miss, sli...
192. UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models ​
Author: Yukun Dai, Mingzhe Dai, Tianshi Wang, Fengling Li, Jingjing Li, Lei Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.13453v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them...
193. CAPRI: Contract-Aware Proof Repair for Isabelle ​
Author: Jim Woodcock, Gabriel Leite, Augusto Sampaio, Ran Wei
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LO
arXiv:2608.13459v1 Announce Type: cross Abstract: We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised. We present CAPRI, a contract-awa...
194. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification ​
Author: Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDI...
195. Concept Drift Detection and Adaptive Retraining of Malware Classification Models ​
Author: Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2608.13465v1 Announce Type: cross Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performa...
196. AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models ​
Author: Mohammed Ayman Habib, Rylan Hart, Morteza Fayazi
Published: 8/15/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.13472v1 Announce Type: cross Abstract: Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert intuition. Among recent developments, LLMs have introduced a promising approach by bringing natural language r...
197. Synthetic Persona Pretraining: Alignment from Token Zero ​
Author: Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, onc...
198. Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity ​
Author: Dananjay Srinivas, Saksham Khatwani, Maria Pacheco
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13484v1 Announce Type: cross Abstract: When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertai...
199. DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data ​
Author: Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13517v1 Announce Type: cross Abstract: Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based...
200. The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity ​
Author: Martin J. Wainwright
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT, math.ST, stat.ML, stat.TH
arXiv:2608.13520v1 Announce Type: cross Abstract: We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}). Its local increments directly control Kullback--Leibler (KL) discretization...
201. Vero: Can AI Agents Build Formally Verified Software Repositories? ​
Author: Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO, cs.PL, cs.SE
arXiv:2608.13522v1 Announce Type: cross Abstract: AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offe...
202. LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure ​
Author: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LIT...
203. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark ​
Author: Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.13555v1 Announce Type: cross Abstract: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, p...
204. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design ​
Author: Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.13560v1 Announce Type: cross Abstract: Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors ...
205. MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence ​
Author: Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Emily Mallaber, Harry R. Klein, Gufran Gungor, Matthew Galvin, Michael Deletto, Sabrina Y. Camp, Stephen C. Van Nostrand, James Provencher, Joyce Yu, Naeem Tahir, Jonathan Wischhusen, Olga Kozyreva, Taylor Ortiz, Hande Tuncer, Jad El Masri, Alys Malcolm, Tali Mazor, Ethan Cerami, Kenneth L. Kehl
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2412.17228v4 Announce Type: replace Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. Open-source AI trial matching tools could democratize access to trial options. Methods: We created Matc...
206. Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows ​
Author: Ling Yue, Nithin Somasekharan, Tingwen Zhang, Yadi Cao, Zhangze Chen, Shimin Di, Shaowu Pan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2505.04997v3 Announce Type: replace Abstract: Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi-stage workflow create significant barriers to entry. We present Foam-Agent, a multi-agent framework that lev...
207. Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning ​
Author: Irene Brugnara, Alessandro Valentini, Andrea Micheli
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2505.13372v2 Announce Type: replace Abstract: Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of temporal planners when a domain is fixed and a set of training problems (not plans) is given. The idea is to extra...
208. Identification of Probabilities of Causation: from Recursive to Closed-Form Bounds ​
Author: Xin Shu, Shuai Wang, Ang Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2505.15274v4 Announce Type: replace Abstract: Probabilities of causation (PoCs) are fundamental quantities for counterfactual analysis and personalized decision making. However, existing analytical results are largely confined to binary settings. This paper extends PoCs to multi-valued treatme...
209. PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research ​
Author: Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Rui Ye, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, hep-lat
arXiv:2512.19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation....
210. DomusFM: A Foundation Model for Event-Based Behavioral Monitoring in Smart-Homes ​
Author: Michele Fiori, Gabriele Civitarese, Flora D. Salim, Claudio Bettini
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.01910v2 Announce Type: replace Abstract: Smart-home sensor-based behavioral monitoring holds significant potential for healthcare, independent living, and early detection of functional or cognitive changes. In this setting, tasks like activity recognition, prediction, and pattern discover...
211. Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design ​
Author: Hai Xia, Carla P. Gomes, Bart Selman, Stefan Szeider
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, math.CO
arXiv:2603.08322v2 Announce Type: replace Abstract: We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result in combina...
212. Auditable Agents ​
Author: Yi Nian, Aojie Yuan, Haiyue Zhang, Jiate Li, Li Li, Xiyang Hu, Hua Wei, Xiongye Xiao, Chaowei Xiao, Yue Zhao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.05485v2 Announce Type: replace Abstract: LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the question is no longer only whether harmful actions can be prevented--it is whether those actions remain answer...
213. Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung ​
Author: Thomas Bartz-Beielstein, Eva Bartz
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.23859v3 Announce Type: replace Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical environments. A review of the relevant open-source tooling shows that existing compliance solutions o...
214. From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction ​
Author: Alex Petrov, Alexander Gusak, Denis Mukha, Dima Korolev
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2604.27906v3 Announce Type: replace Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant context later. This design is useful for thematic recall, but it is mismatched to the kinds of memory ...
215. AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design ​
Author: Haoze Lv, Ning Lu, Ziang Zhou, Yew-Soon Ong, Shengcai Liu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2605.08756v2 Announce Type: replace Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show that large language models (LLMs), when integrated into well-designed frameworks (i.e., LLM-AHD),...
216. CEON: Circular Economy Ontology Network ​
Author: Huanyu Li, Els de Vleeschauwer, Robin Keskis"arkk"a, Mikael Lindecrantz, Mina Abd Nikooie Pour, Ying Li, Ben De Meester, Patrick Lambrix, Eva Blomqvist
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.02253v2 Announce Type: replace Abstract: Increasing the circularity of resource use in our society has been recognized as a path to sustainability, i.e., transitioning into a more circular economy. There are many different circular strategies to do so, such as reusing products and compone...
217. Residual Modeling for High-Fidelity Learned Compression of Scientific Data ​
Author: Liangji Zhu, Sanjay Ranka, Anand Rangarajan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.05389v2 Announce Type: replace Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations. Learned compressors can achieve high compression ratios at moderate accuracy targets, but their aggregate reconstruction losses do not guarantee accuracy fo...
218. Learning to Recover Task Experts from a Multi-Task Merged Model ​
Author: Jinwook Jung, Taegyu Kim, Kumju Jo, Sungyong Baik
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.26902v2 Announce Type: replace Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly...
219. Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents ​
Author: Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2607.14573v4 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We intro...
220. Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures ​
Author: Supantho Rakshit, Adele Goldberg, Henry Conklin
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.22699v2 Announce Type: replace Abstract: As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilinguall...
221. Do LLMs Know Their Vulnerable Scenarios? ​
Author: Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2607.23496v2 Announce Type: replace Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed...
222. AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution ​
Author: Junhao Qiu, Zidong Wang, Yansong Sun, Zhitong Ma, Ping Guo, Qingfu Zhang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.26661v3 Announce Type: replace Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally dif...
223. DAPD: Dual-Anchored Policy Distillation ​
Author: Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, Shixiang Tang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01735v2 Announce Type: replace Abstract: On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student learns privilege-dependent behavior it cannot rep...
224. DiffImaginE: Imagine to Verify Entity Types with Diffusion ​
Author: Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03025v3 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visua...
225. Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study ​
Author: Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03172v3 Announce Type: replace Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this onl...
226. Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load ​
Author: Thomas Bartz-Beielstein, Inalbek Akiev, Lalo Mohamad
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.05018v3 Announce Type: replace Abstract: Short-term load forecasting (STLF) plays a vital role in the electric power industry. It is relevant for critical infrastructure. STLF is no longer purely a performance and accuracy problem, because determinism, fail-safe handling, minimal-attack s...
227. Recursive Synthesis for Long-Horizon Terminal Tasks ​
Author: Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.05466v3 Announce Type: replace Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consis...
228. iARCS: Iterative Agentic RL for Controllable 3D Scene Generation ​
Author: Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06161v2 Announce Type: replace Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limit...
229. CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment ​
Author: Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.09164v2 Announce Type: replace Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic...
230. Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models ​
Author: Kevin Murphy
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.09696v3 Announce Type: replace Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data lea...
231. Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research ​
Author: Lin Liao, Peng Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10363v2 Announce Type: replace Abstract: AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for ...
232. Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information ​
Author: Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2608.10766v2 Announce Type: replace Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that ident...
233. Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration ​
Author: Alan Li, Rahul Saha, Anton Xue, Swarat Chaudhuri, Adam Klivans, Pravesh K Kothari, Raghu Meka
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CC, cs.HC, math.FA
arXiv:2608.11195v2 Announce Type: replace Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures t...
234. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation ​
Author: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.11616v2 Announce Type: replace Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus i...
235. OEIS Open: How many conjectures can language models turn into theorems? ​
Author: Tom Adamczewski
Published: 8/15/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11941v2 Announce Type: replace Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code ...
236. The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot ​
Author: Fangchen Song, Ashish Agarwal, Wen Wen
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC, econ.GN, q-fin.EC
arXiv:2410.02091v4 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for developer productivity and participation in software development. To explore its impact on collaborative...
237. Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries ​
Author: Harshavardhan Battula, Jiacheng Liu, Jaideep Srivastava
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2411.16818v2 Announce Type: replace-cross Abstract: To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU) notes are fused with physiology for in-hospital mortality (IHM) prediction, and to determine how much o...
238. Cueless EEG imagined speech for subject identification: dataset and benchmarks ​
Author: Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour, Hamid R. Rabiee
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2501.09700v2 Announce Type: replace-cross Abstract: Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored the use of imagined speech with semantically meaningful words for subject identification, most have rel...
239. Unmasking Conversational Bias in AI Multiagent Systems ​
Author: Erica Coppolillo, Giuseppe Manco, Luca Maria Aiello
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2501.14844v3 Announce Type: replace-cross Abstract: Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application in critical settings. However, the majority of existing methodologies for identifying biases in generat...
240. Yes, Q-learning Helps Offline In-Context RL ​
Author: Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.17666v5 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objecti...
241. Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision ​
Author: Ahmet Bilican, M. Ak{\i}n Y{\i}lmaz, A. Murat Tekalp, R. G"okberk Cinbi\c{s}
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV, eess.SP
arXiv:2505.12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA achieve efficiency through low-rank updates, their discrete rank constraint limits...
242. How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG ​
Author: Qiming Zeng, Hao Luo, Yuhao Lin, Yicheng Jin, Yuxiang Wang, Fangcheng Fu, Xiao Yan, Jiawei Jiang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2506.06331v2 Announce Type: replace-cross Abstract: By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) to generate quality answers for user questions. Many GraphRAG methods have been proposed and reported in...
243. Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights ​
Author: Yuan Zhong, Ruinan Jin, Qi Dou, Xiaoxiao Li
Published: 8/15/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV
arXiv:2506.17337v5 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing specialist medical VLMs requires substantial computational resources and carefully curated datasets, a...
244. Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models ​
Author: Abdullah X
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2508.12220v2 Announce Type: replace-cross Abstract: Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution controls while assig...
245. REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models ​
Author: Ke Fang, Tianyi Zhao, Qianwen Wang, Lu Cheng
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.14390v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications. Existing prompt-based methods treat calibration largely as a one-shot inf...
246. Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs ​
Author: Haneul Yoo, Jiho Jin, Kyunghyun Cho, Alice Oh
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.05678v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languages as LLMs often rely on English-centric latent representations. In this work, we introduce code-switch...
247. StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars ​
Author: Weijian Li, Hong-Yu Chen, Nabeel Rehemtulla, Ved G. Shah, Dongho Kim, Dennis Wu, Qinjie Lin, Adam A. Miller, Han Liu
Published: 8/15/2026, 4:00:00 AM
Categories: astro-ph.SR, astro-ph.IM, cs.AI
arXiv:2510.06200v4 Announce Type: replace-cross Abstract: Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (light curves) are available in immense quantities and exhib...
248. DiffGRM: Diffusion-based Generative Recommendation Model ​
Author: Zhao Liu, Yichen Zhu, Yiqing Yang, Xiao Lv, Guoping Tang, Rui Huang, Qiang Luo, Ruiming Tang, Kun Gai, Guorui Zhou
Published: 8/15/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2510.21805v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predicts the next item by autoregressively generating its SID conditioned on the user's history. However, two...
249. CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning ​
Author: Tianhui Liu, Hetian Pang, Xin Zhang, Jie Feng, Pan Hui, Yong Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2510.22282v2 Announce Type: replace-cross Abstract: Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-moda...
250. SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control ​
Author: Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Casta~neda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi "Jim" Fan, Yuke Zhu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.GR, cs.SY, eess.SY
arXiv:2511.07820v4 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in s...
251. Automated Design Optimization via Strategic Search with Large Language Models ​
Author: Anthony Carreon, Vansh Sharma, Venkat Raman
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.MA
arXiv:2511.22651v2 Announce Type: replace-cross Abstract: Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design parameters are difficult to define. Large language models (LLMs) offer a promising alternative by dynamic...
252. Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models ​
Author: Malte Hellmeier
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2512.13325v2 Announce Type: replace-cross Abstract: Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing control over data when it is being used to train such machine learning models or when distinguishing model...
253. RadarGen: Automotive Radar Point Cloud Generation from Cameras ​
Author: Tomer Borreda, Fangqiang Ding, Sanja Fidler, Shengyu Huang, Or Litany
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO
arXiv:2512.17897v2 Announce Type: replace-cross Abstract: We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen adapts efficient image-latent diffusion to the radar domain by representing radar measurements in bird's-...
254. Learning Latency-Aware Orchestration for Multi-Agent Systems ​
Author: Xi Shi, Mengxin Zheng, Qian Lou
Published: 8/15/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL
arXiv:2601.10560v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations. Existing orchestration metho...
255. Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map ​
Author: Luis F. Rosario Freytes (University of Michigan)
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.11618v3 Announce Type: replace-cross Abstract: Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions answer different identity questions. We study the represented process available at a receiver: an actual...
256. Safe Exploration via Policy Priors ​
Author: Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO
arXiv:2601.19612v4 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated) environments. In this work, we tackle this challenge by utilizing suboptimal yet conservative policies ...
257. MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models ​
Author: Erica Coppolillo, Emilio Ferrara
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2603.00048v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making. This expansion has motivated growing research into the ethical and moral foundation...
258. CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language ​
Author: Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, Li Zhang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2603.14501v2 Announce Type: replace-cross Abstract: Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-resource programming languages primarily focuses on Domain-Specific Languages (DSLs), leaving general-pur...
259. Automatic Termination Strategy of Inelastic Neutron-scattering Measurement Using Bayesian Optimization for Bin-width Selection ​
Author: Kensuke Muto, Hirotaka Sakamoto, Kenji Nagata, Taka-hisa Arima, Masato Okada
Published: 8/15/2026, 4:00:00 AM
Categories: physics.data-an, cs.AI
arXiv:2603.16946v2 Announce Type: replace-cross Abstract: Currently, an excessive amount of event data is being obtained in four-dimensional inelastic neutron-scattering experiments. A method for automatic bin-width optimization of multidimensional histograms has been developed and recently validate...
260. Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI ​
Author: Anna Kozlova, Stanislau Salavei, Pavel Satalkin, Hanna Plotnitskaya, Sergey Parfenyuk, Andy Nkansah
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2603.25821v3 Announce Type: replace-cross Abstract: We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions. Unlike traditional medical benchmarks that rely on solving standardized test questions, the proposed ...
261. In-context superposition: human-like working memory interference in large language models ​
Author: Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.09670v3 Announce Type: replace-cross Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments. This capacity, known as working memory, is fundamental to human reasoning. Yet, human working memory is strikingly limited, ma...
262. A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network ​
Author: Mehdi Hosseinzadeh, Roohallah Alizadehsani, Amin Beheshti, Hamid Alinejad-Roknyd, Lu Chen, Mohammad Sadegh Yousefpoor, Efat Yousefpoor, Muneera Altayeb, Thantrira Porntaveetus, Sadia Din
Published: 8/15/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2604.15489v2 Announce Type: replace-cross Abstract: The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constraints, and diverse QoS requirements. This paper proposes QQMR, a Q-learning-based QoS-aware multipath ro...
263. IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development ​
Author: Jasmine Moreira
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2604.16399v3 Announce Type: replace-cross Abstract: Adoption of AI-assisted development in 2025 exposed a tool-agnostic failure pattern: experienced developers using frontier models were measurably slower while believing they were faster, and 10.3% of applications in one production showcase le...
264. Cat-DPO: Category-Adaptive Safety Alignment ​
Author: Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Henry Peng Zou, Kaize Ding, Xiyang Hu, Yan Liu, Yue Zhao
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a sing...
265. Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference ​
Author: Xiaowei Mao, Bowen Sui, Weijie Zhang, Yawen Yang, Shengnan Guo, Shilong Zhao, Jiaqi Lin, Tingrui Wu, Youfang Lin, Huaiyu Wan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.23724v4 Announce Type: replace-cross Abstract: Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-field vehicles with subtle abnormal motion. Vision-Language Models (VLMs) provide strong semantic reasonin...
266. SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models ​
Author: Chengjie Hong, Feixiang He, Yiheng Zeng, Lulu Kang, He Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.17985v2 Announce Type: replace-cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential for reducing memory use and accelerating inference in large foundation models, it remains un...
267. Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance ​
Author: Jing Chen, Shixiang Pan, Yujie Fan, Haocheng Ye, Haitao Xu, Wenqiang Xu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.18793v3 Announce Type: replace-cross Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing methods face performance bottlenecks, typically yielding only incremental gains and often exhib...
268. Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking ​
Author: Qinwu Xu, Zhuoheng Li, Jessie Salas
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2605.18852v2 Announce Type: replace-cross Abstract: Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be comparable to variability i...
269. INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification ​
Author: Seongjun Lee, Seokhyun Lee, Changhee Lee
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.20088v2 Announce Type: replace-cross Abstract: Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent complexity of time-series classification (TSC) and to make model decision-making processes more transparent....
270. Annealed Softmax Greedy in Many-Armed Bayesian Bandits ​
Author: William Overman, Mohsen Bayati
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.31034v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher...
271. Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems ​
Author: Jonathan Cola\c{c}o Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.00367v2 Announce Type: replace-cross Abstract: Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to specify than scalar rewards, and they express certain goals t...
272. Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection ​
Author: Atmika Bhardwaj, Silvia Vock, Nico Steckhan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.01896v2 Announce Type: replace-cross Abstract: Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expensive, or biased. For hand detection, particularly in occupational safety settings, public datasets mostl...
273. Constitutional On-Policy Safe Distillation ​
Author: Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.03089v3 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervision. Prior work has shown that OPSD can collapse in verifiable...
274. Do Transformers Need Three Projections? Systematic Study of QKV Variants ​
Author: Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.PF
arXiv:2606.04032v3 Announce Type: replace-cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting...
275. Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides ​
Author: Haoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang
Published: 8/15/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.DC
arXiv:2606.07316v2 Announce Type: replace-cross Abstract: Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label? We build a protocol to find out. H-CSC emits one of three typed outcomes per round -- semantic commit, verdict commit...
276. SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation ​
Author: Junghun Oh, Sungyong Baik, Kyoung Mu Lee
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.16454v2 Announce Type: replace-cross Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates with low-rank matrices. In this paper, we investigate the limitations of the LoRA parameterization from a ...
277. The Hidden Evolution of Disguised Visual Context inside the VLM ​
Author: Wish Suharitdamrong, Tony Alex, Xiatian Zhu, Muhammad Awais, Sara Atito
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.20077v2 Announce Type: replace-cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visu...
278. Communication Heterogeneity and Collective Consensus in Neural Cellular Automata ​
Author: Nishit Singh
Published: 8/15/2026, 4:00:00 AM
Categories: cond-mat.dis-nn, cs.AI
arXiv:2606.21202v2 Announce Type: replace-cross Abstract: Reaching global agreement from purely local interactions is a defining problem of collective intelligence, and most models of it assume that all agents share a single communication protocol. We ask what happens when they do not. Using a Neura...
279. Early Warning Signals for OpenVLA Failure under Visual Distribution Shift ​
Author: Dipesh Tharu Mahato, Rachel Ren
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2606.29699v2 Announce Type: replace-cross Abstract: Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with the steps before failure. We freeze the policy, record one MLP a...
280. LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review ​
Author: Ali Hassaan Mughal, Muhammad Bilal
Published: 8/15/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.05031v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Yet two oracles can look identical and rest on different ground: one assertion encodes a written specifi...
281. Scaling Time Series Classification via XAI-Driven Data Reduction ​
Author: Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15774v3 Announce Type: replace-cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel met...
282. Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training ​
Author: Nuemaan Malik
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.19058v2 Announce Type: replace-cross Abstract: Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE language model AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study...
283. Vibe to Code: Elucidating Strategic Oscillation of Tacit Knowledge in Generative AI Design Workflows -- An Exploratory Qualitative Study ​
Author: Daisaku Sato
Published: 8/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.23126v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI tools has created new literacy demands for designers who must verbalize tacit knowledge through natural language prompts. Yet the micro-level cognitive processes by which designers externalize implicit inte...
284. Moral Hazard in Multi-Agent Language Models ​
Author: Dane Malenfant
Published: 8/15/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2607.23982v4 Announce Type: replace-cross Abstract: Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr"om's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theory-grounded controll...
285. A Distributional Robustness Margin For Pathology Foundation Models ​
Author: Cl'ement Grisi, Jeroen van der Laak, Geert Litjens
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.25497v2 Announce Type: replace-cross Abstract: Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions. The Robustness Index (RI} was proposed to asse...
286. SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups ​
Author: Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.27431v4 Announce Type: replace-cross Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requir...
287. Commit Locally, Exit Globally: Coordinating Adaptive Sampling and Early Exit in Diffusion Language Models ​
Author: Chia-Ming Lee, Shao-Kai Liu, Ming-Ching Chang, Xin Li, Yu-Lun Liu, Chih-Chung Hsu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.28166v2 Announce Type: replace-cross Abstract: Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabilizes before the step schedule is exhausted. This creates two acceleration opportunities, leaving a block...
288. Coordinated incentives in AI-generated misinformation governance ​
Author: Qin Li, Gui Zhang, Minyu Feng, Matjaz Perc, Attila Szolnoki
Published: 8/15/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI
arXiv:2608.07070v2 Announce Type: replace-cross Abstract: With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, undermining information credibility and social trust. This study models the strategic interdependence among...
289. Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks ​
Author: Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07335v2 Announce Type: replace-cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm enables off-policy value learning without relying on experience replay...
290. Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction ​
Author: Miki Ueno
Published: 8/15/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.08443v2 Announce Type: replace-cross Abstract: Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a relational microculture. Recent work has also examined how humans and conversational AI nego...
291. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability ​
Author: Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi, Mislav Balunovic, Theophane Weber, Vahab Mirrokni
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.09538v2 Announce Type: replace-cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical comput...
292. Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning ​
Author: Daoyi Li, Yixian Zhang, Wenbo Ding, Yu Wang, Chao Yu
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10473v2 Announce Type: replace-cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online fine-tuning: as the ...
293. Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation ​
Author: Amit Aflalo, Shahaf E. Finder, Roy Amoyal, Eran Treister, Oren Freifeld
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10805v2 Announce Type: replace-cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field exponentially with the number of decomposition levels while keeping the parameter count l...
294. Governing Agentic AI in FinTech ​
Author: Henry Han
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, q-fin.RM
arXiv:2608.11344v2 Announce Type: replace-cross Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the bin...
295. AI Guardrail Survival under Single-Cycle Agentic Self-Summarization ​
Author: Ted Kwartler, Alan Aqrawi, Arian Abbasi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11392v2 Announce Type: replace-cross Abstract: Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (G...
296. Keep the Future, Drop the Rollout: RIFT for World Action Models ​
Author: Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
Published: 8/15/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.11521v2 Announce Type: replace-cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Ac...
297. REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ​
Author: Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
Published: 8/15/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11698v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imit...
298. Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians ​
Author: Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko
Published: 8/15/2026, 4:00:00 AM
Categories: quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI
arXiv:2608.11911v2 Announce Type: replace-cross Abstract: A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical simulation methods. Here we demonstrate that this problem can be effectively amortized across an arbitr...
299. Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge ​
Author: Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi
Published: 8/15/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12218v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help th...