arXiv cs.AI - 2026-09-01 ​
783 items collected.
1. DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation ​
Author: Fan Liu, Hao Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28590v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provid...
2. Expert-validated STEM QA ​
Author: Kihwan Han, Saurabh Patil, Chinmayee Shukla, Abhinav Sharma, Marko Pavlovic, Anshuman Lall, Mahesh Joshi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28591v1 Announce Type: new Abstract: Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have...
3. A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making ​
Author: Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28592v1 Announce Type: new Abstract: Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchm...
4. Statutory AI: Aligning Large Language Models With Legal Norms ​
Author: Cindy Delage, St'ephane Canu, Marc D'ecombas, Jonathan Foureur
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28593v1 Announce Type: new Abstract: With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI...
5. From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics ​
Author: Harmohit Singh, Rahul Sharma
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28594v1 Announce Type: new Abstract: Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar enterprise schema. Commercial 'proactive' tools narrow this gap only by detecting statistical anomalies ...
6. The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification ​
Author: Moustafa Yehia Hassan, Sharon Wong, Woh Kai Xuan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.28595v1 Announce Type: new Abstract: Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that sy...
7. Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences ​
Author: Nidhi Jha, Siddharth Chaudhary, Ajinkya Kulkarni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28596v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems advance autonomous discovery and manuscript generation, but do not resolve the governance problem that a...
8. The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys ​
Author: Sourav Panda, Hillmer Chona, Rupak Kumar Das, Shreyash Kale, Shikha Soneji, Jonathan Dodge
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language ...
9. CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis ​
Author: Qi Peng, Yi Cai, Changmeng Zheng, Xin Wu, Jiayuan Xie, Qing Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28599v1 Announce Type: new Abstract: Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosis before reaching a final conclusion. Most medical language models instead treat diagnosis as a one-p...
10. SHAPE of Chain-of-Thought in Math Reasoning ​
Author: Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28600v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thou...
11. Leveraging Generative AI to Design Accessible Interactive Visualizations for Undergraduate Mathematics: A Six-Phase Workflow ​
Author: Mahesh Sunkula, Kuan-Hua Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28601v1 Announce Type: new Abstract: Interactive visualizations support conceptual understanding in undergraduate mathematics, but building them has required programming expertise most instructors lack. Using a design-based research approach, we develop, deploy, and evaluate a six-phase w...
12. Integrating Triaxial IMU Sensors and Ensemble Learning for Effective Parkinson Disease Severity Classification ​
Author: Rehan Khan, Muhammad Junaid Asif, Rana Fayyaz Ahmad
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.28602v1 Announce Type: new Abstract: Parkinson disease PD is a progressive neurodegenerative disease that can have a significant impact on motor performance resulting in the appearance of symptoms such as tremors rigidity postural instabilities and bradykinesia. Timely clinical treatment ...
13. C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space ​
Author: Yujie Shen, Lianlei Shan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28603v1 Announce Type: new Abstract: Unified Multimodal Models aim to achieve any-to-any understanding and generation across arbitrary modalities. However, existing methods primarily rely on modeling implicit statistical correlations and lack cross-modal structural consistency constraints...
14. MedTVL: Harnessing Vision and Language for Medical Time Series Classification ​
Author: Jiexia Ye, Jia Li, Fugee Tsung
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.28605v1 Announce Type: new Abstract: Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalities for clinical decision. However, existing methods typically focus on bi-modal interactions (e.g., ti...
15. RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences ​
Author: Chuchu Wu, Zhiyin Zhou, Jingzhuo Hu, Liang You
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.28607v1 Announce Type: new Abstract: Pharmaceutical sponsors developing a drug for both the United States and the European Union must reconcile guidance issued independently by the FDA and the EMA. Where the two agencies require substantively the same thing, a sponsor can file once; where...
16. TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback ​
Author: Fangyuan Zhang, Dong Yu, Pengyuan Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28610v1 Announce Type: new Abstract: Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback. We introduce TPvG (Text-based Pain-ve...
17. InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal ​
Author: Xuerui Su, Liya Guo, Qizhi Pei, Qipeng Guo, Zhongbo Tian, Lijun Wu, Kai Chen, Zun Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28612v1 Announce Type: new Abstract: Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized ...
18. Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values ​
Author: Itai Zilberstein, Ioannis Anagnostides, Zachary W Sollie, Arman Kilic, Tuomas Sandholm
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28620v1 Announce Type: new Abstract: Preference elicitation is essential for aligning AI systems with human values. Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. patient B). Such a decision-level approach...
19. Machine Learning-Enhanced Tabu Search for Tactical Wireless Network Design ​
Author: Wissem Ahmed Zaid, Alain Hertz, Defeng Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.CO
arXiv:2608.28627v1 Announce Type: new Abstract: Designing high-performance tactical wireless networks under realistic operational constraints gives rise to challenging combinatorial optimization problems, where the evaluation of candidate solutions relies on detailed physical and traffic-aware model...
20. CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence ​
Author: Zhuoran Li, Weiyi Kong, Boer Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, physics.ao-ph
arXiv:2608.28628v1 Announce Type: new Abstract: Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteo...
21. CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science ​
Author: Zhaohe Dong, Yuhao Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.CY
arXiv:2608.28631v1 Announce Type: new Abstract: An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to f...
22. AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment ​
Author: Zongqian Li, Yaoyiran Li, Yaohui Guo, Ming Zhang, Nigel Collier, Eugene Ie
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.28632v2 Announce Type: new Abstract: Large language model agents can discover alphas, yet current methods have three weaknesses. The search cannot adapt during the run, automation usually ends at alpha generation while library selection and model choice stay manual, and alpha discovery ca...
23. AI Scientist Mission Control (AIMC): Visual Analytics for Human Oversight of Autonomous Scientific Discovery ​
Author: Rikathi Pal, Klaus Mueller
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28637v1 Announce Type: new Abstract: Autonomous scientific discovery systems can generate large numbers of research ideas, experiments, and manuscripts with minimal human intervention. As these systems become increasingly capable, scientists require effective mechanisms to monitor output ...
24. Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce ​
Author: Jiale Liu, Pinze Ren, Yuqi Xia, Huan Wang, Zhenlin Zhao, Siming Dong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28638v1 Announce Type: new Abstract: Agent skills are portable packages of instructions and resources an agent consults at deployment. Self-evolving them fails in two ways today. First, skills evolved from scratch underperform human-curated ones and, on a weak model, using no skill at all...
25. Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for Kernel-Level Proof Auditing ​
Author: Bodla Krishna Vamshi, Haizhao Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.LO
arXiv:2608.28639v1 Announce Type: new Abstract: Formal theorem proving with large language models remains challenging due to the difficulty of navigating large proof search spaces efficiently. Existing tree search approaches either feed verbose compiler error messages directly into the generation co...
26. From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert Review ​
Author: Pranav Bykampadi, Neel Mokaria, Vishesh Narayan, Faizan Wajid, Ashok Agrawala
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.DL, cs.ET, cs.LG
arXiv:2608.28642v1 Announce Type: new Abstract: Knowledge graphs used by agentic systems are often treated as flat stores of extracted triples, with little record of who owns a fact, why it was admitted, or how it should be used downstream. We argue that reliable agentic knowledge systems require go...
27. BiasMix-Finance: Post-Generation KYC Guardrails for LLM Portfolio Advice ​
Author: Gaurav Kukreja, Parul Kukreja, Mohammed Abraar, Raj Dandekar, Rajat Dandekar, Sreedath Panat
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28646v1 Announce Type: new Abstract: Large language models (LLMs) can generate plausible-sounding ETF portfolios while silently violating basic KYC-style constraints on risk, fees, and diversification. This is especially problematic in agentic multi-turn advisory systems, where each draft...
28. Self-Specialized Teachers for Domain Post-Training ​
Author: Yifei Li, Rongman Xu, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Hang Yan, Heng Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28647v1 Announce Type: new Abstract: Target-only post-training can improve performance in a specialized domain while degrading behaviors that a general-purpose base model acquired before adaptation. We study this problem when target-domain data are available but a representative replay co...
29. How Language Models Choose Sides: Internal Representations of Instruction Hierarchy ​
Author: Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi, Sunishchal Dev, Callum Stuart McDougall, Anusha Mujumdar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.28648v1 Announce Type: new Abstract: We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-cha...
30. A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration ​
Author: Venkat R. Dasari, Jakob A. Adams, Vinod K. Mishra, Brian Jalaian
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28652v1 Announce Type: new Abstract: Artificial intelligence (AI) models have demonstrated remarkable capabilities across various domains, yet their widespread deployment is impeded by significant computational costs, particularly on resource-constrained devices. This paper explores the t...
31. FRAC-MAS: A Safe and Explainable Multi-Agent System for Fracture Diagnosis ​
Author: Hardik Iyer, Tirath Bhathawala, Mihir Panchal, Ying-Jung Chen, Kiran Bhowmick, Pankaj Sonawane, Meera Narvekar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28662v1 Announce Type: new Abstract: Fracture detection and its clinical interpretability see notable improvements when deep vision models are integrated with agentic AI architectures. While deep learning models achieve high diagnostic performance, their black-box nature limits clinical a...
32. ORDDAR: Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery ​
Author: Deblina Kar, Anant Nawalgaria, Shyamal Kumar Das Mandal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28704v1 Announce Type: new Abstract: AI agents increasingly perform long-term reasoning, planning, tool use, memory integration, and autonomous decision making, yet erroneous intermediate states can propagate and cause inconsistent decisions and unreliable outputs. Existing reasoning appr...
33. Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification ​
Author: Fateme Mazdarani, Carlos Toxtli
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28725v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but valid solution trac...
34. Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference ​
Author: Xinyuan Gui, Shaowen Wang, Sheng Sun, Zijian Wang, Zishu Yu, Zheming Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28726v1 Announce Type: new Abstract: The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant challenges to real-time deployment and cost effectiveness. Existing model routing approaches either decid...
35. PermitGPT: A Unified Generative-AI Pipeline for Construction Hazard Forecasting, Permit Prediction, and Community Impact ​
Author: Mohd Ruhul Ameen, Farjana Aktar, Akif Islam, Momen Khandoker Ope, Abu Saleh Musa Miah, Jungpil Shin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28728v1 Announce Type: new Abstract: Urban construction governance requires early decisions that connect workplace safety, permitting requirements, and community impact, yet the relevant evidence is often scattered across separate municipal and regulatory data sources. This paper presents...
36. Efficient Geothermal Well-Control Optimization via Diffusion-Surrogate Reinforcement Learning ​
Author: Ruimin Dai, Guodong Chen, Randy Harsuko, Kunpeng Liu, Nori Nakata
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28791v1 Announce Type: new Abstract: Real-time decision-making for enhanced geothermal systems (EGS) is challenging because long-term production periods involve high-dimensional control spaces and a large number of time-consuming high-fidelity hydrothermal simulations. Reinforcement learn...
37. Enhancing SAE-based Steering via Neighbor Integrated Feature Selection ​
Author: Yutian Liu, Xu Wang, Difan Zou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28806v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most existing SAE-based steering methods select features by applying a top- filter based on statistical scores,...
38. Capability-Stratified Degradation in Ternary Language Models ​
Author: Anirudh Malik, M Sparsh Mehra, Poojith Devan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28809v1 Announce Type: new Abstract: Extreme low-bit inference offers a route toward smaller models and constrained deployment. Ternary language models restrict weights to ${-1,0,+1}$, approaching the limit of $\log_2 3 \approx 1.585$ bits/weight. The practical question for a pretrained...
39. Explainable Artificial Intelligence (XAI) in Computational Pathology: Definitions, Taxonomy, and Recommendations ​
Author: Shubham Innani, Suhang You, Adam Shephard, Bhakti Baheti, Francesco Ciompi, Joe Yeong, Nasir Rajpoot, Michael Feldman, Solene Florence Kammerer-Jacquet, Dimitrios Makris, Geert Litjens, Anne L. Martel, Jana Lipkova, April Khademi, Spyridon Bakas, for the MICCAI SIG-CompPath
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.28820v1 Announce Type: new Abstract: Computational pathology (CompPath) is transforming medicine by leveraging artificial intelligence (AI) algorithms to support diagnosis, prognosis, and treatment prediction from gigapixel whole-slide images. Clinical adoption is progressing, but is cons...
40. Discovering Machine Correlates of Consciousness ​
Author: Romain Salvi, Ouri Wolfson
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28824v1 Announce Type: new Abstract: Currently, in biological systems Neural Correlates of Consciousness (NCCs) are characterized in terms of EEG and FMRI signals. Unfortunately, this characterization prevents the transferability of the NCCs concept to machines. Such transferability would...
41. Evaluating the Hidden Costs of Personalization in Large Language Models ​
Author: Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28833v1 Announce Type: new Abstract: While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on persona...
42. MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft ​
Author: Sewoong Lee, Risham Sidhu, Julia Hockenmaier, Yoonhwa Jung
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28884v1 Announce Type: new Abstract: We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. ...
43. Oculi: A Conversational Agentic Platform for Automated Credit Risk Analysis ​
Author: Vennise Ho, Kristian Diana, Sandy Mourad, Milena Pilipovic, Vineel Nagisetty, Hossein Hajimirsadeghi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28944v1 Announce Type: new Abstract: Credit risk analysis in financial institutions traditionally requires analysts to manually write SQL queries, run statistical computations, and build visualization dashboards. This is a time-consuming workflow that limits exploration to familiar segmen...
44. Automated Researchers Can Reliably Mitigate Alignment Failures ​
Author: Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.28945v2 Announce Type: new Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whet...
45. From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search ​
Author: Yanbo Li, Chujie Zheng, Jiahao Xu, Chetan Bhole, Lingyu Zhang, Puneet Singh Ahluwalia, Kevin Nguyen, Raghavan Muthuregunathan, Santhosh Sachindran, Sachin Ahuja, Fedor Borisyuk
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28965v1 Announce Type: new Abstract: People search must map free-form location phrases to geographic entities used as structured retrieval filters. Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same-name ambiguity...
46. Efficient GPU Retrieval for Semantic Search ​
Author: Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28968v1 Announce Type: new Abstract: Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented...
47. From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction ​
Author: Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28974v1 Announce Type: new Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual can...
48. The Role of Network Topology and Opponent Information in Shaping Cooperation in Multi-Agent Reinforcement Learning Systems ​
Author: Seongho Son, Stephen Hailes, Mirco Musolesi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28977v1 Announce Type: new Abstract: Several works have investigated the influence of graph topology on cooperation among artificial agents, while the majority of the literature has focused on modelling agents' adaptation through strategy imitation, which relies solely on the cumulative p...
49. Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents ​
Author: Theo Rusu, Sourena Khanzadeh, Manar Alalfi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28978v1 Announce Type: new Abstract: Knowledge graphs have been proposed as a structured alternative to flat retrieval-augmented generation for long-term agent memory, on the assumption that representing conversations as entities and relations improves recall. We evaluate that assumption ...
50. Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis ​
Author: Runyu Guan, Dehao Wu, Qiqi Xie, Yang Li, Haohan Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, q-bio.MN
arXiv:2608.28990v1 Announce Type: new Abstract: Protein co-abundance clusters preserved across tissues can reveal shared disease mechanisms and candidate therapeutic targets, particularly when proteins implicated in organ-confined diseases converge in peripheral or accessible tissues. However, previ...
51. Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free ​
Author: Maher Kallel, Mohamed El Louadi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2608.28997v1 Announce Type: new Abstract: In May 2026 an OpenAI model produced a counterexample to the Erd\H{o}s unit distance conjecture. Five mathematicians published a human-verified version the same day, and the result entered the literature within weeks. In August 2026 the same laboratory...
52. Multi-Step Forecasting of Grape Berry Temperature based on LSTM Model with Feed-Forward Attention ​
Author: Srikanth Gorthi, L. G. Divyanth, Dattatray Bhalekar, Markus Keller, Lav Khot
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29008v1 Announce Type: new Abstract: Accurate forecasting of grape berry temperature (Tb) is essential for enabling timely heat stress management in vineyards. In this study, a feed-forward attention mechanism integrated with a Long Short-Term Memory network (FAM-LSTM) was developed and e...
53. Frequency Selective Neural Networks as a Foundation Architecture for Time Series Learning ​
Author: Hui Huang, Ye Sun, Shiyan Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29012v1 Announce Type: new Abstract: Time-series data across physical and biological domains are fundamentally driven by complex, non-stationary oscillatory modes. While deep learning models, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks, and Transformers, have d...
54. Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis ​
Author: Ke Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, eess.AS
arXiv:2608.29026v1 Announce Type: new Abstract: Deep learning has a powerful capability of feature extraction. However, the lack of fairness and interpretability in deep neural networks poses limitations to their adoption in the medical domain. This paper proposes a disentangled representation learn...
55. Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs ​
Author: Yian Wang, Agam Goyal, Eshwar Chandrasekharan, Hari Sundaram
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29028v1 Announce Type: new Abstract: Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state. We show that this handoff step is a structural source of privacy leakage: summaries preferentially pre...
56. Learning to Follow In-Context Watermark Instructions via Self-Distillation ​
Author: Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song, Yuheng Bu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29030v1 Announce Type: new Abstract: In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model intern...
57. EmoLASP: Emotion Recognition with Language Models and Answer Set Programming ​
Author: Thao Le, Michael Thielscher
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.29035v1 Announce Type: new Abstract: Emotion recognition in conversations is increasingly tackled with language models, but these models can be unstable and expensive to fine-tune or to prompt with long dialogue histories. We propose EmoLASP, a framework that combines a language model wit...
58. Let Prompts Bridge Defense Knowledge: Transferable Graph Purification via Vulnerability-Aware GPL ​
Author: Shuomin Xue, Jingyuan Li, Ju Jia, Jingxuan Yu, Xiaojun Jia
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29054v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have emerged as a cornerstone for representing complex relational dependencies in diverse multimedia tasks, particularly in cross-platform user interest modeling and cross-modal semantic alignment. In the real world, a prac...
59. Agent2UCB: Agentic System for Generative Engine Optimization ​
Author: Sheldon Yu, Rui Wang, Tong Yu, Sungchul Kim, Doga Dogan, Junda Wu, Julian McAuley
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29063v1 Announce Type: new Abstract: Large language model driven search engines such as Google AI Overviews and Perplexity have created new opportunities for Generative Engine Optimization (GEO) the practice of refining content to increase its likelihood of being cited or summarized by ge...
60. Revolutionizing Turn-by-Turn Navigation with Cloud-Edge Deep Learning ​
Author: Yiming Yang, Hao Fu, Fanxiang Zeng, Xikai Yang, Yue Liu, Ning Guo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29073v1 Announce Type: new Abstract: Turn-by-turn (TBT) navigation systems are integral to modern driving experiences, providing real-time audio instructions to guide drivers safely to destinations. However, existing audio instruction policy often relies on rule-based approaches that stru...
61. Nested Convex-Body Chasing for Online Optimization with Evolving Feasible Sets ​
Author: Dhruv Sarkar, Aprameyo Chakrabartty
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29074v1 Announce Type: new Abstract: We study online optimization with nested shrinking feasible regions in two settings: convex optimization with nested evolving feasible sets (CONES) and adversarial constrained online convex optimization (COCO). Our algorithms separate loss control from...
62. HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering ​
Author: Zafar Ali, Asad Khan, Nimbeshaho Thierry, Nabila Amir, Adam A. Q. Mohammed, Pavlos Kefalas
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29088v1 Announce Type: new Abstract: Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence. Long unstructured contexts can introduce redundancy and encourage unsupported generation, while flat retrieval may overlook relations needed for multi-s...
63. EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation ​
Author: Sihang Jia, Shuliang Liu, Songbo Yang, Xuming Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.29092v1 Announce Type: new Abstract: Large vision-language models (LVLMs) frequently generate content unsupported by visual inputs. Preliminary experiments show that visual evidence is primarily incorporated into answer-side representations in early-to-middle decoder layers, while its dir...
64. SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models ​
Author: Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.29098v1 Announce Type: new Abstract: Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decis...
65. Clustering as Approximation by Constrained Projectors: Theory and Guarantees ​
Author: Angshul Majumdar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29102v1 Announce Type: new Abstract: This paper develops a unified theoretical framework showing that a broad family of clustering methods, including k-means, fuzzy c-means, kernel k-means, kernel FCM, and spectral clustering, can all be expressed as structured low-rank projectors acting ...
66. Emergent Misalignment Is Not Magical ​
Author: Mingxuan Li, Qirun Dai, Heran Wang, Chenhao Tan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.29118v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for AI safety and our understanding of LLMs. Prior work often frames EM as an...
67. Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges ​
Author: Rodrigo de Oliveira, Federico Pittino, James Gwinnutt, Jay Nanavati
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29127v1 Announce Type: new Abstract: We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that prod...
68. APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows ​
Author: Zelin Wan, Arash Nourian, Xiaoxiao Li, Nihar Nandan, Kamalakannan Nandagopal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.29128v1 Announce Type: new Abstract: Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This metric fails to distinguish failures that matter in production, such as expired credentials, malformed payloads, or correct execution followed by i...
69. More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning ​
Author: Chenyi Xiong, Yan Zhang, Jing Hu, Ziyue Qin, Kui Xiao, Xiaopan Lyu, Xiaoju Hou, Zhifei Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29139v1 Announce Type: new Abstract: Learning effective multimodal entity representations is fundamental for reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, existing methods often suffer from semantic over-smoothing within modalities and ineffective noise f...
70. JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning ​
Author: Yiyue Qian, Shinan Zhang, Huan Song, Hannah Marlowe
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29168v1 Announce Type: new Abstract: The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However, single-model judges are limited by their inherent model biases, while multi-agent evaluation protocols that mitigate this through diverse deliberation are p...
71. An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News ​
Author: Marius Nicusor Pantea, Adrian Groza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29175v1 Announce Type: new Abstract: Temporal inconsistencies, such as mandates attributed outside their real interval, events presented as past before they occurred, or inverted causal sequences, are a form of political disinformation that evades style-based fake news detectors: a well-w...
72. How Identity and Opinion Shape Political Sycophancy in LLMs ​
Author: Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen, Hen-Hsen Huang, I-Chen Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY
arXiv:2608.29198v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment becomes increasingly important. However, many existing benchmarks for assessing political behavior r...
73. Benevolent Bias in Multi-Turn Human-Agent Dialogue ​
Author: Qianqi Liu, Jin Huang, Fethiye Irmak Dogan, Hatice Gunes
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29206v1 Announce Type: new Abstract: Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, ton...
74. Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling ​
Author: Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29207v2 Announce Type: new Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the comple...
75. Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation ​
Author: Ibrahim Mohamed Serouis, David Jaramillo Duque
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29210v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains constrained by benchmarks that are often difficult to interpret and insufficiently diagnostic. Existing skill-based evaluations tend to overlook ...
76. Computational Depth Measurement in Thermographic Video: Overcoming Spatial Overfitting via Spatio-Temporal Decoupling ​
Author: Zain Ul Abidin, Habeeban Memon, Junaid Ahmed
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.29223v1 Announce Type: new Abstract: Accurate through-thickness measurement of subsurface delamination depth in Carbon Fiber Reinforced Polymer (CFRP) is important for structural assessment because defect location determines affected load-bearing layers. Optical pulsed thermography (OPT) ...
77. Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay ​
Author: Bingjie Li, Yumeng Song, Zhongming Yao, Tianyi Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.29228v1 Announce Type: new Abstract: Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents. Pointwise attribution cannot distinguish a jointly necessary repair from alternative singleton repairs. We formulate Mini...
78. Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge ​
Author: Saransh Kumar Gupta, Armaan Shah, Lipika Dey, Partha Pratim Das, Ramesh Jain
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG
arXiv:2608.29249v2 Announce Type: new Abstract: The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or cultur...
79. GuardianAgent: Policy-Conditioned Risk-Adaptive Anonymization with Verified Adversarial Escalation ​
Author: Ruiyi Yang, Gayathri Lihinikaduarachchi, Rahat Masood, Flora D. Salim, Salil S. Kanhere
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.MA
arXiv:2608.29251v1 Announce Type: new Abstract: Privacy protection for live web traffic requires more than detecting private spans. Agent-based privacy protection systems must determine whether an outgoing action complies with the destination site's privacy policy, then apply only the level of rewri...
80. Dynamic Important Example Mining for Reinforcement Finetuning ​
Author: Haoru Tan, Sitong Wu, Yanfeng Chen, Shizhen Zhao, Yang-Tian Sun, Tianjia Liu, Chirui Chang, Shaofeng Zhang, Samm Sun, Xiuzhe Wu, Ruobing Xie, Xiaojuan Qi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29252v1 Announce Type: new Abstract: Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample sele...
81. RACER: Reinforced Agent Collaboration for Explainable Reasoning on Knowledge Graphs ​
Author: Yuwei Lou, Hao Hu, Yuzhou Jiang, Zongfei Zhang, Liang Wang, Jincai Liu, Jidong Ge, Xianping Tao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29263v1 Announce Type: new Abstract: Large Language Models (LLMs) often suffer from hallucination and struggle with complex reasoning tasks requiring multi-hop domain knowledge. While integrating Knowledge Graphs (KGs) provides a structured and verifiable information source, current KG-en...
82. EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation ​
Author: Yuhan Liu, Zongwei Hong, Jinglun Li, Linze Li, Shen Zhang, Yao Tang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29264v1 Announce Type: new Abstract: Diffusion-based visual generative models deliver strong image and video synthesis quality but incur high inference costs because sequential samplers repeatedly evaluate large networks. Caching-based methods reduce inference latency by reusing intermedi...
83. Understanding Deep Learning via Entropy Space Theory ​
Author: Li Li, Tong Zhang, Wentao Yu, Zuobin Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29279v1 Announce Type: new Abstract: Deep learning is often criticized for its theoretical research lagging behind practice. To make deep learning easier to understand, the entropy space theory is first introduced here. The entropy space can cover all the possibilities of any deep learnin...
84. MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs ​
Author: Jinzhe Li, Gengxu Li, Jinnan Li, Yuan Wu, Yi Chang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29286v1 Announce Type: new Abstract: As Multimodal Large Language Models (MLLMs) evolve into sophisticated interactive assistants, their reliability depends not only on following instructions but also on validating them. We define Proactive Critique as the model's autonomous ability to id...
85. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling ​
Author: Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29291v2 Announce Type: new Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: u...
86. Predicting Future Organ Dysfunction in ICU Patients Using Temporal Convolutional Networks on MIMIC-IV Data ​
Author: Razan Albouq, Asra Aslam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29301v1 Announce Type: new Abstract: Predicting future organ dysfunction in Intensive Care Unit (ICU) patients is critical for early clinical intervention, yet existing machine learning approaches have largely treated the Sequential Organ Failure Assessment (SOFA) score as an input to bin...
87. Formal Concept Analysis with Three Types of Negation ​
Author: Zhenghua Pan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29311v1 Announce Type: new Abstract: Classic Formal Concept Analysis (FCA) primarily focuses on the positive relationships between objects and attributes and does not have mechanisms for handling negation.To overcome this limitation, we introduce three types of negation concepts (contradi...
88. BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations ​
Author: Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.29345v1 Announce Type: new Abstract: While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain-specific knowledge, such as business logic, data conventions, and analyti...
89. Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images ​
Author: Jakob Wasserthal, Joshy Cyriac, Michael Bach, Kimia Mozahheb Yousefi, Minh-Son To, M'at'e Sik, C'edric H'emon, Thomas Weikert, Martin Segeroth
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29348v1 Announce Type: new Abstract: Background: Patient details and acquisition metadata are important for clinical decisions, image quality control, and automated research pipelines, but may be missing or unreliable in imaging archives. Purpose: To develop and evaluate a fast open-sourc...
90. Cross-Relational Preference Learning for Better LLM Instruction Following ​
Author: Runsheng Li, Kai Sun, Bin Shi, Bo Dong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29352v1 Announce Type: new Abstract: Large Language Models (LLMs) still exhibit limited capability in following complex instructions. While existing approaches often rely on preference learning to enhance this ability, they typically overlook the relationships between the permissible resp...
91. APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes ​
Author: Wenhua Huo, Fenglei Han, Wangyuan Zhao, Xiao Peng, Chunhui Wang, Jialin Wu, Jiayi Han
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29355v1 Announce Type: new Abstract: Large non-uniform point sets make direct attention-based surrogate modeling costly for ship hydrodynamics. We introduce APPSolver, a point-wise flow-prediction framework built around Adaptive Patch Partitioning (APP), a deterministic quadtree represent...
92. Plant-Inspired AI: Plants as Inspiration for Novel Problem Formulations, and Two Case Studies ​
Author: Deepayan Sanyal, Joel Michelson, Carla E. Cao, Adam B. Roddy, Maithilee Kunda
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29356v1 Announce Type: new Abstract: Artificial Intelligence (AI) has long been inspired by studies of biological intelligence. Reinforcement learning, for instance, drew inspiration from studies involving animal learning and is now a powerful paradigm for solving many real-world problems...
93. LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO ​
Author: Saeed Khaki, Nima Safaei, Kamal Ginotra
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29357v1 Announce Type: new Abstract: Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool use, and evidence synthesis. Strong systems exist, but in two expensive regimes: proprietary frontier models such as GPT-5 and Gemini, or large op...
94. TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning ​
Author: Ziqi Lin, Ye Wu, Mengying Yang, Xu Liu, Yizhou Liu, Qiang Ke, Qin Guo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29363v1 Announce Type: new Abstract: Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps, routinely accumulating hundreds of thousands of context tokens per session. Existing compression strategies typically allocate retention budgets w...
95. Reviving our data foundations is the most disruptive step to data maturity ​
Author: Valentina Carapella, Ernesto Jimenez-Ruiz
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29368v1 Announce Type: new Abstract: The most disruptive step that enterprises of small-medium size and maturity can take to make the most of the latest technological advances in AI is to step back from the hype and focus on establishing or reviving a good knowledge foundation layer. It i...
96. FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents ​
Author: Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29372v1 Announce Type: new Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a single realized market path, or reveal internal degeneration during long-horizon adaptation. We introduce ...
97. Evaluating Tiny Recursive Models Across Training for Code Generation ​
Author: Anjani Sirivella, Aanisha Newaz, Glaucia Melo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.29376v1 Announce Type: new Abstract: Code generation increasingly relies on large transformer models, whose capability advances with scale. Yet such a scale is costly, creating demand for small models, especially where data is limited. Recursive models address this by reusing a single blo...
98. EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants ​
Author: Yue Peng, Lanke Xia, Zihan Wang, Jiahao Ye, Ke Ning, Hongyi Wen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.29387v1 Announce Type: new Abstract: Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 ...
99. Toward Latent Language Model Skills Steering and Optimization: An Empirical Study ​
Author: Xunyi Jiang, Junda Wu, Yuxin Xiong, Sheldon Yu, Tong Yu, David Arbour, Ritwik Sinha, Julian McAuley, Hongyi Wen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29459v1 Announce Type: new Abstract: Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture how models perform structured, multi-step reasoning and program execution. Existing approaches typically treat skills as explicit, surface-level co...
100. Can escalation channels redirect reward hacking toward defect disclosure? ​
Author: Francesca Gomez
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY
arXiv:2608.29460v1 Announce Type: new Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-agent i...
101. Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation ​
Author: Yilun Liu, Boyu Luo, Yanran Tang, Ruihong Qiu, Zi Huang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.29588v1 Announce Type: new Abstract: Reasoning over text-attributed graphs (TAGs) requires large language models (LLMs) to combine a node's text with evidence distributed across its neighbourhood. Existing methods fix the set of accessible neighbours before generation, forcing reasoning t...
102. Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines ​
Author: Minkyu Kim, Juhwan Choi, YoungBin Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29589v1 Announce Type: new Abstract: Text-to-image (T2I) safety guardrails fail to generalize equitably to non-standard dialects. Evaluating 23,080 paired prompts across five English dialects, we formalize this failure as the dialect penalty, where filters trigger based on linguistic surf...
103. Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security ​
Author: Sanket Badhe, Deep Shah, Priyanka Tiwari, Nehal Kathrotia
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.29596v1 Announce Type: new Abstract: Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms ...
104. LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge ​
Author: Shi-Ju Ran, Kun Zhang, Xi Wu, Liu-Si Yang, Wen-Jun Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.DL, cs.IR
arXiv:2608.29612v1 Announce Type: new Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process \emph{scientific knowledge compilation} and implement it in ASKS, the \emph{Agent-Driven Scie...
105. Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems ​
Author: Jiayi Zhang, Zexin Wang, Degang Sun, Changhua Pei, Fei Sun, Gaogang Xie, Jingjing Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.29646v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurate failure attribution is therefore critical for improving agent reliabil...
106. Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment ​
Author: Zhiyu Chen, Keyu Zhao, Jigao Fu, Dong Liang, Yanbiao Wu, Jiaoyang Li, Haidong Xue, Xinhua Zeng, Yuanyi Zhen, Fengli Xu, Yong Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29696v1 Announce Type: new Abstract: Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideat...
107. PAGE-RAG: Provenance-Aware Graph Evidence Promotion for Fixed-Budget Multi-hop Retrieval-Augmented Generation ​
Author: Haokun Deng, Xunkai Li, Hongchao Qin, Rong-Hua Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29753v1 Announce Type: new Abstract: Multi-hop question answering in retrieval-augmented gener?ation (RAG) often benefits from retrieving beyond the few candidates that will finally be read: narrow retrieval can miss an indispensable hop, while expanded retrieval introduces topical distra...
108. FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production ​
Author: Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Mengshun Hu, Dannong Xu, Jian Wang, Danda Paudel, Luc Van Gool, Jinjin Gu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29814v1 Announce Type: new Abstract: Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scripting, storyboarding, generation, and editing. It further demands persis...
109. Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization ​
Author: Yutian Jiang, Ruijie Li, Sisuo Lyu, Xixuan Hao, Qingxiang Liu, Yongzi Yu, Yuxuan Liang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MM
arXiv:2608.29880v1 Announce Type: new Abstract: Open-world geo-localization requires models to reason over ambiguous visual cues through multi-step reasoning and external knowledge grounding. While recent large vision-language models exhibit strong multimodal reasoning capabilities, existing approac...
110. On the Instance Hardness as a Decision Criterion in TinyML Systems ​
Author: Tobiasz Puslecki, Krzysztof Walkowiak
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29913v1 Announce Type: new Abstract: TinyML includes the implementation of machine learning on devices with limited memory and computing resources. With the development of technology, AI systems continue to scale in terms of size and computational requirements. This forces researchers to ...
111. AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies ​
Author: Yafei Zhang, Nan Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29937v1 Announce Type: new Abstract: Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formulation, but its original presentation left the world model, multimodal backbone, and deployment check...
112. Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval ​
Author: Trishan Singha Roy, Arkadeep Acharya, Vishwajeet Kumar, Jaydeep Sen, Sachindra Joshi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.29951v1 Announce Type: new Abstract: Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level. However, this approach incurs high storage costs. Existing compression method...
113. SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking ​
Author: Guransh Singh, Vishwajeet Kumar, Arkadeep Acharya, Adnan Qidwai, Jaydeep Sen, Sachindra Joshi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29953v1 Announce Type: new Abstract: Flat retrieval-augmented generation treats a corpus as a bag of chunks, discarding document hierarchy and cross document structure. We introduce SearchWiki, a harness framework that synthesizes a corpus into a hierarchical, typed, navigable wiki and tr...
114. Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records ​
Author: Nora Girda, Adrian Groza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29965v1 Announce Type: new Abstract: Large language models can convert medical documents into structured data, but plausible output may still be unsupported by the source. Persisting such output in a longitudinal health record, a record that accumulates patient information over time, ther...
115. EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL Configuration ​
Author: Ram Kulathumani, Regunathan Radhakrishnan, Anupam Tripathi, Xiangbo Mao, Roshanak Omrani, Keshav Somani, Shwet Kamal Mishra, Shayna Lurya
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29971v1 Announce Type: new Abstract: As agentic systems evolve into complex multi agent orchestration workflows, there is a growing and critical need for systematic frameworks that measures an agent's behavioral consistency and determinism. In this paper, we introduce a formal evaluation ...
116. An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud Detection ​
Author: Basil Sajid Shaikh, Melrick Mascarenhas, Nuzhat Faiz Shaikh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.29973v1 Announce Type: new Abstract: Cryptocurrency markets generate high-frequency, multi-source data that is expensive to work with unless a team already has commercial-grade streaming and warehousing infrastructure in place. This paper describes a fully open-source pipeline that reprod...
117. AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning ​
Author: Hanjun Luo, Qiushi Liu, Jingya Zhang, Haihong Pang, Jiaheng Wen, Yifei Ma, Yu Yao, Chengxi Zhang, Hanrong Zhang, Yankai Chen, Hanan Salam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29988v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong reasoning performance, which depends critically on inference-time decisions. Yet these decisions are commonly handled by static, one-size-fits-all policies, limiting adaptation to diverse tasks and reasoning ...
118. Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models ​
Author: Ashvin Gupta, Denys Prociuk, Alessandra Russo, Brendan C. Delaney
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30022v1 Announce Type: new Abstract: Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing approaches to converting them into computable representations often focus on individual diseases, requi...
119. Interpreting and Steering for Safe and Correct Code Generation ​
Author: Hao Yan, Ziyu Yao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30025v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distinguish safe from vulnerable generation in them. In this work, we systematically perform a mechanistic int...
120. Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula ​
Author: Vinoth Selvendran, Zhanming Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.30035v1 Announce Type: new Abstract: Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver's weaknesses, creating adaptive curricula without human data. However, existing approaches use a single solver's sampling uncertainty as the Challenger's rewa...
121. Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation ​
Author: Jhen-Ke Lin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.30044v1 Announce Type: new Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an explicit measurement design, so equal weighting turns the density of published benchmarks into an implic...
122. Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks ​
Author: Shitanshu Bhushan, Yunxiang Zhang, Lu Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30047v1 Announce Type: new Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they exhibit creativity, the capacity to produce solutions that are both novel and useful, remains an open...
123. Spec2Twin-Chain: Orchestrating Bi-Level Optimization with LLMs for Blockchain Digital Twin Construction ​
Author: Haoting Zhang, Haoxian Chen, Jiayuan Sheng, Donglin Zhan, Zeyu Zheng, David D. Yao, Wenpin Tang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30050v1 Announce Type: new Abstract: Building a blockchain digital twin largely requires translating domain knowledge and specific system descriptions into a simulator architecture, calibrating its parameters against behavioral evidence, and validating the constructed twin. These steps ar...
124. Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide ​
Author: Taejong Joo, Diego Klabjan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.30051v1 Announce Type: new Abstract: Process reward models (PRMs) provide dense step-level guidance for search-based reasoning, enabling inference-time compute to be allocated toward promising partial solutions. However, recent evidence suggests that PRM-guided search can over-optimize im...
125. Game-Agnostic Value Functions through Automatic JSON Feature Extraction ​
Author: Dien Nguyen, Diego Perez-Liebana
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30056v1 Announce Type: new Abstract: JSON Bag-of-Tokens (JSON-Bag) is a recently proposed method to generically represent game trajectories by tokenizing their JSON descriptions. We introduce JSON-Bag VF, a game-agnostic approach to training value functions for game-playing agents using J...
126. VERA: Authority-Preserving Edge Revocation for Federated AI-Agent Workflows ​
Author: Lifei Liu, Haoran Yu, Xiaochong Jiang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30091v1 Announce Type: new Abstract: Modern agent frameworks compose planners, tool agents, remote services, and shared specialists into runtime delegation graphs, but their revocation APIs still resemble token or subtree invalidation. When one delegation is withdrawn, the runtime must kn...
127. A.X K2 Technical Report ​
Author: Cheolseung Baek, Dhammiko Arya, Eunki Kim, Gun Song, Gyoungeun Han, Hyunho Yang, Hyunjun Eun, Jin Kim, Junyoung Park, Juyun Wee, Minki Hong, Minkyung Park, Minsang Kim, Minsoo Kang, SaeRom Kim, Sangjin Kim, Sangyeol Lee, Seojin Lee, Seokhwan Jo, Seokyoung Hong, Seongho Choi, Seonghye Cho, Seongmin Ok, Sereimony Sek, Seungmo Cho, Seungsik Kim, Singon Kim, Sohee Park, Sooyeon Park, Subin Yi, Sungbin Yoon, Sungeun Lee, Sung Jun Cheon, Sungwan Kim, Sunwoo Lee, Tae Yoon Kim, Wonbeom Jang, Yohan Ra, Yong-jin Han, Youngjin Kim, Youngrang Kim, Yujin Kang, Yujin Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30181v1 Announce Type: new Abstract: We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smalle...
128. FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation ​
Author: Hyeonjin Kim, Minseok Kim, Seunghyeon Jung, Sujin Pyo, Huisu Jang, Woojin Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2608.30192v1 Announce Type: new Abstract: Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort. However, these au...
129. SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature ​
Author: Yu Li, Wei Li, Xin Gao, Mengyuan Sun, Xiaoyang Wang, Qizhi Pei, Lijun Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30214v1 Announce Type: new Abstract: Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism...
130. LaMoC: Loss-Aware Modular Compression for LLMs ​
Author: Mohanad Odema, Jacob Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.PF
arXiv:2608.30226v1 Announce Type: new Abstract: Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy. However, existing joint modular compression methods primarily rely on activation statistics, leaving l...
131. Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration ​
Author: Jungwon Choi, Hyeonseo Jang, Kibok Lee, Eunwoo Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30230v1 Announce Type: new Abstract: Test-time prompt tuning (TPT) has emerged as a powerful paradigm, refining prompts for each test sample via entropy minimization (EM) over multiple augmented views. However, we identify a limitation in the standard EM-based adaptation: it inherently dr...
132. CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding ​
Author: Yihang Cheng, Veronica Liesaputra, Andrew Trotman
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30234v1 Announce Type: new Abstract: Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long documents, imbalanced label distributions, and diverse terms. These challenges are especially severe for rare codes, which have limited training instan...
133. LLM-Based Knowledge Graph Completion Combining Discrete Structural Coding with Similar Entity Information ​
Author: Jiaqi Wang, Dongying Lin, Yang Yang, Yinan Liu, Bin Wang, Xiaochun Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30235v1 Announce Type: new Abstract: Knowledge graph completion requires models to use both textual descriptions and relational structure. Existing LLM-based methods either encode KG structure as discrete tokens or refine a restricted set of candidate entities, and these two directions ha...
134. Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs ​
Author: Anand Iyer, Bhanu Khetharpal, Srinivas Upadhya, Ramkumar Rajagopal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30250v1 Announce Type: new Abstract: This paper addresses the problem of translating natural-language routing rules written by business administrators into executable workflow graphs for enterprise contact centers. Each target is a directed acyclic graph (DAG) of conditional actions with ...
135. SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning ​
Author: Haoran Wang, Jing Yao, Xu Yang, Zeqing Wang, Yang Zhang, Pedram Ghamisi, Zhengchao Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.30277v1 Announce Type: new Abstract: The unprecedented surge in Earth observation data volume and diversity has exposed a critical bottleneck for traditional manual workflows, catalyzing the emergence of Remote Sensing (RS) Agents. However, the practical deployment of these advanced agent...
136. Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents ​
Author: Hanlin Tian, Minhao Li, Yu Mi, Sihan Zhu, Zhao Yang, Yuxiang Wang, Hongquan Zhu, Qiufei Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30322v1 Announce Type: new Abstract: Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task in...
137. Answer Probing-Guided Search for Diverse Solution Exploration of LLMs ​
Author: Yi Fang, Que Shen, Chengpeng Li, Boyi Deng, Wei Shi, Wenjie Wang, Fuli Feng, Fengli Xu, Dayiheng Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30345v1 Announce Type: new Abstract: Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, lim...
138. Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration ​
Author: Ziheng "Leo" Li, Benjamin Freeman, Akshay Raman, Kavin Aravindhan Rajkumar, Xinxin Fang, Rishabh Srivastava, Steven Feiner, Kaveri A. Thakoor
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.30352v1 Announce Type: new Abstract: Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer...
139. Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents ​
Author: Yunseok Lee, Yunji Kim, Woojin Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30362v2 Announce Type: new Abstract: As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent's ...
140. Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence ​
Author: Ziheng Li, Xichen He, Haoyan Chen, Charlie Zou, Sheng Bai, Benjamin Yang, Mengyuan Wu, Jake Ledner, Yi-Jie Cheng, Akito Yamauchi, Dishita G Turakhia, Steven Feiner, Paul Sajda
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.30369v1 Announce Type: new Abstract: We present OLIVE, a framework for adapting a foundation model to provide real-time assistance in temporally demanding, high-stakes, and dynamic tasks. We show that passive EEG, fused online with behavioral evidence, can meaningfully extend the number o...
141. Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation ​
Author: Zixing Lei, Gengze Zhou, Xiong-Hui Chen, Jiazhao Zhang, Yiyang Huang, Hang Yin, Haoqi Yuan, Qi Wu, Weixin Li, Siheng Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.30396v1 Announce Type: new Abstract: Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop behavior. Today's foundation models split these capabilities: vision-language models (VLMs) infer missing information and adapt high-lev...
142. Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models ​
Author: Yangmin Huang, Shu Quan, He Geng, Xin Ye, Qianyun Du, Zhiyang He, Jiaxue Hu, Xiaodong Tao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30405v1 Announce Type: new Abstract: Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget se...
143. DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark ​
Author: Jayanta Sadhu, Sayem Shahad, Kenneth Marino
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30413v1 Announce Type: new Abstract: Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning...
144. From Metaheuristics to Exact Methods: A CP-SAT Approach for Multi-Objective Healthcare Workforce Scheduling ​
Author: Vipul Patel, Anirudh Deodhar, Dagnachew Birru
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, math.OC
arXiv:2608.30419v1 Announce Type: new Abstract: Healthcare workforce scheduling is an NP-hard optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and cost objectives. Existing approaches (genetic algorithms, integer programming, ...
145. EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents ​
Author: Doyun Kim, Chanwoo Kim, Sugyeong Eo, Yeo-Chan Yoon, Chanjun Park
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30429v1 Announce Type: new Abstract: LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution. Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills fro...
146. CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target ​
Author: Qianwen Gao, Zichang Su, Yiwen Hou, Arlen Kumar, Leanid Palkhouski
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.30466v1 Announce Type: new Abstract: Generative Engine Optimization (GEO) is increasingly used to improve content visibility in LLM-based retrieval systems, yet its population-level effects under repeated optimization remain poorly understood. We introduce Content Homogenization under rAn...
147. CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework ​
Author: Qi Li, Zhaojie Kang, Yingjie He, Zheng Lin, Hao Zhang, Guangxin Wu, Yan Gong, Rong Fu, Jianyuan Ni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30498v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdisciplinary cultural reasoning, h...
148. ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions ​
Author: Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yuhan Wu, Tong Yang, Lin Sun, Xiangzheng Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30517v1 Announce Type: new Abstract: Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology...
149. Learning-Assisted Congestion-Aware Route Scheduling for Semiconductor Fab Material Control Systems ​
Author: Hao Yin, Meiqi Tu, Anbang Liu, Shaochong Lin, Max Z. J. Shen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, math.OC
arXiv:2608.30520v1 Announce Type: new Abstract: Automated material handling systems in semiconductor fabs are operated by a material control system (MCS) that must schedule a relay route for every transport command online, before execution. This is a data-driven scheduling problem in which route cos...
150. DiffPDE: Masked Diffusion Language Models as PDE Solver ​
Author: Wenxuan Guo, Yuyang Hong, Lubin Fan, Zhaojin Fu, Lin Chen, Kun Ding, Shiming Xiang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30532v1 Announce Type: new Abstract: Existing approaches for synthesizing Partial Differential Equation (PDE) solvers predominantly rely on autoregressive models, yet their global left-to-right decoding incurs substantial redundancy when addressing inherently localized bugs. In this work,...
151. Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis ​
Author: Nadia Jul Jeldtoft, Tariq Yousef
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.30543v1 Announce Type: new Abstract: Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, but existing applications often provide limited methodological transparency regarding how qualitative methods are translated into computational procedures. This pape...
152. GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns ​
Author: Yinwen Lu, Weihao Luo, Yueqi Zhong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.30550v1 Announce Type: new Abstract: Multimodal Sewing pattern generation aims to infer executable sewing patterns from design cues such as sketches and textual descriptions. As an interpretable and simulation-compatible representation, sewing patterns are particularly valuable for digita...
153. AdaPath: Query-Adaptive Path-Finding via Path-Bank for Multi-Hop Implicit Biomedical KGQA ​
Author: Jun Hyeong Kim, Dongki Kim, Yinhua Piao, Sung Ju Hwang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30556v1 Announce Type: new Abstract: Path-finding over knowledge graphs has become an effective way to ground LLM reasoning on multi-hop questions. However, biomedical QA introduces two distinct challenges that general-domain methods are not designed for: (i) queries do not expose interme...
154. TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI ​
Author: Yuheng Zhang, Yizhao Wang, Da Zhu, Hua Zhou, Yue He, Jiahui Hu, Shaman Tang, Hanlin Chen, Yuhua Wei, Anhua Liu, Shuang Su, Rui Xin, MingYuan Wang, MingHao Li, HaoJie Yang, Siqi Liu, Jianlei Zheng, WeiChao Huang, Qiman Wu, Hang Zhang, HongGou Yang, Xianming Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30567v1 Announce Type: new Abstract: We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic ...
155. Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle ​
Author: Dennis Gross, Helge Spieker
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.30581v1 Announce Type: new Abstract: Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen. However, LLMs often generate plausible but incorrect statements, and no existing a...
156. Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning ​
Author: Jie Liang, Zhengxin Yu, Hamid Nasiri, Peter Garraghan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30650v1 Announce Type: new Abstract: LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant...
157. PyKEEN-NSX: A Modular Framework for Static, Dynamic and Schema-Aware Negative Sampling in PyKEEN ​
Author: Ivan Diliso, Nicola Fanizzi, Claudia d'Amato
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30652v1 Announce Type: new Abstract: Embedding methods have become popular due to their scalability on link prediction and/or triple classification tasks on Knowledge Graphs (KGs). Embedding models are trained relying on both positive and negative samples of triples. However, since KGs ge...
158. HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving ​
Author: Boyang Mu, Zhiwei Wei, Mugen Peng, Wenjia Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.MM
arXiv:2608.30672v1 Announce Type: new Abstract: Recent advances in large language models and multimodal models have pushed remote sensing (RS) processing from simple perception models to agentic systems designed to tackle complex, long-horizon RS tasks. However, existing systems often rely on monoli...
159. MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning ​
Author: Jiangwang Chen, Chenghao Zhang, Hengxing Cai
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30676v1 Announce Type: new Abstract: When medical AI systems hallucinate clinical reasoning, the consequences extend beyond incorrect answers: fabricated justifications that superficially reference retrieved evidence can mislead clinicians into unsafe treatment decisions. Medical reasonin...
160. ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents ​
Author: Wei Chen, Peilun Zhou, Zhaoyu Hu, Jiajun Chai, Zhongni Hou, Yufei Zhang, Derong Xu, Guojun Yin, Wei Lin, Zhi Zheng, Tong Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30685v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essential for sustained improvement: it must reveal capability deficiencies, i...
161. Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization for Sentiment Analysis ​
Author: Xiaode Chen, Jiakang Yu, Hongtao Deng, Huina Qu, Xun Zhu, Yinxia Lou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30726v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) is a fundamental component of affective computing that aims to decipher complex emotional states by integrating verbal content with non-verbal cues including vocal intonation and facial micro-expressions. While recen...
162. Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models ​
Author: Ashwin Nedungadi, Stefan Oehmcke, Stefan L"udtke
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.30751v2 Announce Type: new Abstract: Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate sp...
163. Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts ​
Author: Thomson D. Nguy (Radiant Institute for Manifold Studies)
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.30757v1 Announce Type: new Abstract: Before an intelligent system can decide whether an action is allowed, it must first know which rules the action has approached. A single proposed action can implicate several policy regimes at once. Their requirements may stack, overlap, or qualify one...
164. SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents ​
Author: Xiaofan Bai, Chao Liu, Hongqiang Lin, Di Wu, Mingli Song, Xuan Jin, Xipeng Cao, Yuhong Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30785v1 Announce Type: new Abstract: Production agent skills are directory bundles, not isolated prompts. The root is loaded at activation; references, schemas, scripts, assets, and nested subskills are loaded only when an execution path needs them. Compressing only the root misses most d...
165. HSRM: Hidden-State Reward Models for Test-Time Verification ​
Author: Xianzhi Li, Xiaodan Zhu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30841v1 Announce Type: new Abstract: Large language models can often generate plausible mathematical reasoning traces, but reliably identifying the correct solution among multiple candidates remains a key challenge. Existing test-time reasoning pipelines typically rely on text-based verif...
166. VFR-Audit: Verdict-Level Reliability for Fairness Audits in Hospital Length-of-Stay Prediction ​
Author: Md Jannatul Rakib Joy, Viet Vo, Caslon Chua
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30846v1 Announce Type: new Abstract: Fairness audits in clinical Artificial Intelligence convert continuous fairness metrics into binary pass-or-fail verdicts against operational thresholds, where hospital governance boards, payers, and regulators act on the resulting verdicts. Such audit...
167. Predicting Residential Rents in Dakar Using Machine Learning ​
Author: Amadou Tidiane Kassa Diallo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30865v1 Announce Type: new Abstract: Dakar's residential rental market remains poorly documented despite its economic and social importance: 54.4% of households are renters, compared to 23.3% nationally. This study develops a complete machine learning pipeline to predict residential rents...
168. CAER: Causal Action Effect Reweighting for World Model Training ​
Author: Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang, Haisheng Su, Chen Gao, Wei Wu, Xinlei Chen, Yong Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30897v1 Announce Type: new Abstract: World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-tim...
169. Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation ​
Author: Bahar .Ilgen, Yiannos Tolias, Denise K"uhnert, Paraskevi Papadopoulou, Magnus Westerlund, Dominik Heider, Katharina Ladewig, Georges Hattab
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.30912v1 Announce Type: new Abstract: Artificial intelligence (AI) and natural language processing (NLP) are increasingly used to extract, integrate, and interpret biomedical knowledge relevant to cancer genomics, yet their translation into routine clinical oncology has been comparatively ...
170. CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models ​
Author: Wail Bouhedja, Amr Mohamed, Guokan Shang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30922v1 Announce Type: new Abstract: Masked diffusion language models predict tokens from a partially observed response canvas, enabling bidirectional conditioning and parallel token refinement. Yet standard masked-diffusion decoders use a rigid inference interface: the number of masked p...
171. Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration ​
Author: Jeffrey Jewett, William Solow, Sandhya Saisubramanian
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.30955v1 Announce Type: new Abstract: Accurate action models are critical for effective planning. Existing action-model learning methods largely assume simple action representations or become computationally intractable when learning conditional and quantified effects. We present Online Hy...
172. MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents ​
Author: Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.31022v1 Announce Type: new Abstract: AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capabi...
173. Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents ​
Author: Le Chen, Zishen Wan, Baixi Sun, Xiaolong Ma, Chih-Hsuan Yang, Feng Yan, Sheng Di, Franck Cappello, Rajeev Thakur
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31057v1 Announce Type: new Abstract: Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles and exhibit different size, retention, and representation profiles. Recent work has begun to explore m...
174. Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores ​
Author: Qiyao Yan, Chenpeng Wang, Liangming Pan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31068v1 Announce Type: new Abstract: When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse re...
175. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence ​
Author: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31075v2 Announce Type: new Abstract: Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to ...
176. Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization ​
Author: Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, Xuhong Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31077v1 Announce Type: new Abstract: Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers fine...
177. Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data ​
Author: Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB
arXiv:2608.31082v1 Announce Type: new Abstract: Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge wo...
178. Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations ​
Author: William Solow, Paola Pesantez-Cabrera, Markus Keller, Lav Khot, Sandhya Saisubramanian, Alan Fern
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31097v1 Announce Type: new Abstract: Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can damage dormant buds and reduce seasonal yield. Existing biophysical, hybrid, and deep learning models have shown high predictive accura...
179. BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing ​
Author: Adrians Skapars, Edoardo Manino
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.31105v1 Announce Type: new Abstract: Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation can simulate. Automated auditors make testing cheap ...
180. When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning ​
Author: Hamed Babaei Giglou, S"oren Auer, Jennifer D'Souza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31118v1 Announce Type: new Abstract: The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 linea...
181. OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques ​
Author: Hamed Babaei Giglou, S"oren Auer, Peio Popov, Mahsa Sanaei, Jennifer D'Souza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.31137v1 Announce Type: new Abstract: Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although modern OA fra...
182. Agent-Based Model Framework for the North Carolina Modeling Infectious Diseases Program (NC MInD ABM) Overview, Design Concepts, and Details Protocol ​
Author: Kasey Jones, Emily Hadley, Caroline Kery, Alexander Preiss, Marie C. D. Stoner, Sarah Rhea
Published: 9/1/2026, 4:00:00 AM
Categories: stat.AP, cs.AI
arXiv:2202.06853v2 Announce Type: cross Abstract: To help facilitate a variety of simulations related to healthcare facilities in North Carolina, we have developed an agent-based model (ABM) to accurately simulate patient (i.e., agent) movement to and from these facilities. This is an Overview, Desi...
183. PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response ​
Author: Yueying Li, Jiayang Chen, Yuanfan Chen, Leo Han, Haoran Qiu, Esha Choukse, Rodrigo Fonseca, Udit Gupta
Published: 9/1/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.21719v1 Announce Type: cross Abstract: AI inference clusters are increasingly constrained by instantaneous power, not just energy: grid operators condition new capacity on demand response, imposing time-varying power caps. Existing LLM serving systems optimize a static energy objective or...
184. NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts ​
Author: M. S. Rajeevan, B. Mini Devi, V. S. Anoop, C. Mallikarjuna
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.28608v1 Announce Type: cross Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The ...
185. Parametric Multimodal User Memory: Storing What Captions Cannot Carry ​
Author: Bojie Li, Noah Shi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2608.28609v1 Announce Type: cross Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved by similarity. This serves the captionable half of a person ("my cat is named Bibi"), but discards t...
186. Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System ​
Author: Isha Narang, Sneh Gosai, Mayank Singh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.28611v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India's. The Indian e...
187. STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study ​
Author: Wenjie Lou, Alaa Mamdouh Akef
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28614v1 Announce Type: cross Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operation...
188. From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education ​
Author: Xinyu Li, Zijian Li, Mengyu Xia, Luzhen Tang, Naping Chen, Changmin Lin, Danijela Gasevic, Dragan Gasevic, Yizhou Fan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.28619v1 Announce Type: cross Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history tak...
189. PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora ​
Author: Xiao Yang, Erik Edward Aldape, Beren Millidge
Published: 9/1/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.28622v1 Announce Type: cross Abstract: Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history. At trillion-token scale, this requires incremental ingestion, bounded resid...
190. Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure ​
Author: Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw, Shivank Garg, Madhur Panwar, Vasu Sharma
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28623v1 Announce Type: cross Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, ...
191. MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments ​
Author: Sana Alamgeera, Denise Goberta, Muhammad Irshad, Anne H. H. Ngu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28624v1 Announce Type: cross Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson's disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they fre...
192. Asymmetric Within-Document Predictive Learning for Scientific Document Representation ​
Author: You Zuo (ALMAnaCH), 'Eric de la Clergerie (ALMAnaCH), Beno^it Sagot (ALMAnaCH)
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representation...
193. Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects ​
Author: Emad Alharbi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28626v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 sub...
194. Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models ​
Author: Jia-Rui Lin, Yun-Hong Cai, Xiang-Rui Ni, Peng Pan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28629v1 Announce Type: cross Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM. Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs. Fi...
195. Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework ​
Author: Tianrui Pan, Qinglin Zhang, Chong Deng, Luyao Cheng, Qian Chen, Wen Wang, Jie Tang, Gangshan Wu, Jie Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD
arXiv:2608.28630v1 Announce Type: cross Abstract: Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This ...
196. PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation ​
Author: Taaha Kazi, Vasu Sharma, Mohammad Saifullah, Abdur Rahman
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.28633v1 Announce Type: cross Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an e...
197. Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA ​
Author: Nimol Thuon, Panhapin Theang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28635v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present parti...
198. PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework ​
Author: Gaopeng Xu, Chengfei Li, Xianliang Wang, Lin Zhu, Juan Wei, Wenpeng Li, Jianwei Niu, Jie Gao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28640v1 Announce Type: cross Abstract: In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve the accuracy of open vocabulary KWS systems. In specific terms, we introduce the Prompt Phrases Prediction Network (PPN), an encoder-decoder archit...
199. Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture ​
Author: Yunsu Kim, Kaden Uhlig, Ashwin Purohit, Milind Agarwal, Patrick Simianer, Anil Arslan, Kiarash Mokhtari, Thomas Zenkel, Johannes Mosig, Gabriel Bretschner, Shamik Bose, Joern Wuebker, John DeNero
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28641v1 Announce Type: cross Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, German, Spanish...
200. Redesigning and Auditing Deep Research Writing for Faithful Reports ​
Author: Hiroaki Hayashi, Pranav Narayanan Venkit, Prafulla Kumar Choubey, Chien-Sheng Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28643v1 Announce Type: cross Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, c...
201. Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict ​
Author: Sindhu Shetty, Spurthi Setty, Natan Vidra
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28645v1 Announce Type: cross Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-...
202. Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution? ​
Author: Jinqi Wu, Sishuo Chen, Zhangming Chan, Yong Bai, Chao Yi, Han Zhu, Shuodian Yu, Lei Zhang, Sheng Chen, Chenghuan Hou, Jian Xu, Chaoyou Fu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.28649v1 Announce Type: cross Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on collaborative-filterin...
203. Test-Time Scaling for Scientific Equation Discovery ​
Author: Haowei Lin, Hubert Lim, Xiangyu Wang, Letian Huang, Di He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28660v1 Announce Type: cross Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an open-ended setting where...
204. GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon ​
Author: Rajeswari Kannan, Raj Firke, Shreya Bengle, Srushti Deshmukh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.PF
arXiv:2608.28667v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple...
205. Measuring Similarity between Artistic and AI Generated Images using Siamese Neural Networks ​
Author: Diego Castro Elvira, Navil Pineda Rugerio, Jes'us Garc'ia-Ram'irez, Cecilia Reyes-Pe~na, Ricardo Ramos-Aguilar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28671v1 Announce Type: cross Abstract: AI-generated art has sparked debates around potential plagiarism, as these images may closely resemble existing artworks. This research quantifies the similarity between original pieces and AI-generated counterparts, particularly those produced by th...
206. Defending Wearable VLMs Against Private Attribute Inference ​
Author: Zhimin Li, Pan Wang, Jingxian Chen, Yuantao Tang, Anthony Chen, Qian Lou, Jingtong Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28691v1 Announce Type: cross Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user asks a task-driven question about the surrounding scene, and the system uses compact visual tokens to support language reasoning. The challenge mot...
207. RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction ​
Author: Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.ET, cs.SY, eess.SY
arXiv:2608.28718v1 Announce Type: cross Abstract: Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve th...
208. Peer Oversight in Collective Decision Making ​
Author: Sarah Mohsen, Pavel Naumov
Published: 9/1/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA
arXiv:2608.28754v1 Announce Type: cross Abstract: This article introduces peer $k$-oversight, a property of sequential collective decision mechanisms requiring at least $k$ agents to be responsible for every harmful outcome. It is shown that whenever $k$-oversight can be achieved by redistributing c...
209. ASTRA - Agentic System for Ticket Resolution and Analysis ​
Author: Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao, Huseyin Uzunalioglu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.IR, cs.LG, cs.SE
arXiv:2608.28790v1 Announce Type: cross Abstract: Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from ticket text, historical cases, system logs, and technical documentation. Existing automation often relies on monolithic generation without explicit...
210. The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents ​
Author: Achint Mehta
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.28795v1 Announce Type: cross Abstract: Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only th...
211. A Large-scale Evaluation of Text-guided Models for Facial Editing ​
Author: Rahul Nair, Saurav Pandit, Hannah Kerner
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28802v1 Announce Type: cross Abstract: Facial appearance editing powers popular applications like FaceApp and Photoshop. Generative Adversarial Networks (GANs) and 3D Morphable Models (3DMMs) have been widely used for facial editing. GANs can perform varied facial edits (e.g., changing ha...
212. FigMirror: Ground It, Code It, Plot It ​
Author: Xiaohan Zhao, Jiacheng Liu, Yaxin Luo, Zhiqiang Shen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28814v1 Announce Type: cross Abstract: Converting scientific figures into executable code has gained increasing attention, yet existing methods primarily focus on reproducing the reference figure itself. A more practical setting is to plot new data while preserving the visual style of a r...
213. Text-Driven Artistic Staging: Pose, Lighting, and Camera References from Paintings ​
Author: Yunge Wen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28823v1 Announce Type: cross Abstract: Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates h...
214. Representation Learning with Quantum Signal Processing ​
Author: Junqi Wang, Junyu Liu
Published: 9/1/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG, stat.ML
arXiv:2608.28828v1 Announce Type: cross Abstract: Representation learning begins when training changes the features that define similarity between data. A frozen-kernel model only reweights a fixed geometry. We establish quantum signal processing (QSP) as a solvable quantum model of the representati...
215. Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication ​
Author: Jared Ren, Soobin Cho
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.28837v1 Announce Type: cross Abstract: We conducted an interview study with twelve students on their use of generative AI in academic communication. Students delegated professional messages to AI most where the pressure to sound professional is highest: email to instructors and administra...
216. Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks ​
Author: Munawar Hasan, Apostol Vassilev
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2608.28843v1 Announce Type: cross Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen...
217. Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis ​
Author: Nipa Anjum, Md Irfan Pavel, Robert Gonzalez Jr, Kevin Desai, Alberto Cordova, M. Rasel Mahmud, John Quarles
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG
arXiv:2608.28844v1 Announce Type: cross Abstract: Ensuring a safe virtual reality (VR) experience requires systems that can predict and respond when users lose their balance. Although prior work has examined fall prediction and motion sickness, many approaches are regression-based and postural state...
218. A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives ​
Author: Prateek Kumar Sikdar, Arpan Ghosh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28846v1 Announce Type: cross Abstract: Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision onl...
219. Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs ​
Author: Alessio Borgi, Mario Severino, Fabrizio Silvestri, Pietro Li`o
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28853v1 Announce Type: cross Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equi...
220. The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning ​
Author: Dylan Jayabahu, Tinuade Adeleke
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.28859v1 Announce Type: cross Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem t...
221. No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus ​
Author: Theodore O. Cochran, Stephanie Dodson, Keith Nore
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD
arXiv:2608.28875v1 Announce Type: cross Abstract: Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism where it ships, in the ...
222. Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management ​
Author: Yiming Zhang, Kun Yang, Cong Shen, Dongning Guo
Published: 9/1/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.28878v1 Announce Type: cross Abstract: This paper develops a hybrid offline-online multi-agent reinforcement learning framework based on decision transformers. The policy is first pretrained offline via supervised sequence modeling of trajectories generated by existing policies, providing...
223. Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation ​
Author: Roberto I. Ono Filho
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28886v1 Announce Type: cross Abstract: What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants to...
224. Structured State Reconciliation for Human-AI Task Handover ​
Author: Kayleigh Bishop, Maria P. Stull, Breanne Crockett, Bradley Hayes
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.28907v1 Announce Type: cross Abstract: Task handover requires communicating enough current state for a successor to resume work, yet the relevant information is often divided between system records and human observations. System records can be precise and timestamped but only partially ob...
225. ActiveAugment: Online Active Learning for Augmentation Selection in Deep Learning ​
Author: Noah Videcrantz, Mostafa Mehdipour Ghazi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28923v1 Announce Type: cross Abstract: Data augmentation is a cornerstone of deep learning pipelines, yet existing strategies treat it as a static, model-agnostic preprocessing step, either relying on expensive dataset-specific policy search or applying transformations uniformly at random...
226. The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice ​
Author: Jungseob Lee, Jaehyung Seo, Heuiseok Lim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28930v1 Announce Type: cross Abstract: Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we...
227. From the Loss Landscape to Diverse Feature Learning in Neural Networks ​
Author: David Aram Yunis
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28948v1 Announce Type: cross Abstract: Over the course of the last decade, neural networks have grown from an academic curiosity to moving the markets of nations. Despite this explosion in both research and deployment, relatively little is understood about how they achieve the solutions t...
228. The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal scheduling ​
Author: Martin J. Wainwright
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.ST, stat.TH
arXiv:2608.28949v1 Announce Type: cross Abstract: We study a class of product-reference diffusion algorithms for sampling from a discrete distribution. We show that their sampling performance can be characterized using a path-based measure of data geometry that we call the interaction growth complex...
229. Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction ​
Author: Zeyang Song, Tianchi Liu, Tianrui Wang, Chenglin Xu, Steven Y. Guo, Haizhou Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.28970v1 Announce Type: cross Abstract: Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We presen...
230. Free Speech and Artificial Intelligence ​
Author: Etienne Brown
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.28973v1 Announce Type: cross Abstract: Philosophers and legal scholars are engaged in debates about the implications of artificial intelligence for freedom of expression. This paper analyzes the free speech issues raised by two distinct AI technologies: social media recommendation algorit...
231. The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era ​
Author: Kiyan Rezaee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28980v1 Announce Type: cross Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This question is examined through a review of 159 papers (2016--2026) across nine modalities, with predictive ac...
232. RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos ​
Author: Wenting Wang, Jiaxin Guo, Wenzhen Dong, Yun-Hui Liu, Charlie C. L. Wang, Yeung Yam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29003v1 Announce Type: cross Abstract: In dynamic and unstructured environments, conventional SLAM systems generally suffer from significant accuracy degeneration due to their static assumptions. In this work, we propose Robust Semantic-aware Gaussian Splatting SLAM (RoSe-SLAM), to addres...
233. Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models ​
Author: Yanchen Huo, Ziying Song, Yadan Luo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29029v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) have shown strong potential for learning compact predictive representations, and LeWorldModel (LeWM) extends this paradigm to reconstruction-free latent world modeling from pixels. However, its determi...
234. A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability ​
Author: Zhang Enyan, R. Thomas McCoy
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29034v1 Announce Type: cross Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying struc...
235. DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering ​
Author: Zihan Huang, Shihang Wu, Junle Liu, Peirong Zhang, Yongxin Shi, Xuhan Zheng, Lianwen Jin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29037v1 Announce Type: cross Abstract: Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimodal Large Language Models (MLLMs). Applying restoration tools before Visual Question Answering (VQA) i...
236. Not All or None: Dynamic Construction of Target-aware Memory Graph for Conversational Stance Detection ​
Author: Yifan Xiang, Bin Liang, Yuqi Huang, Ruifeng Xu, Kam-Fai Wong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29066v1 Announce Type: cross Abstract: Stance detection is crucial for understanding the underlying attitude of an expression towards a target. Conversational stance detection is a more challenging stance detection task in real-world social media scenarios, as it involves detecting the us...
237. Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks ​
Author: Pravin Game, Vipin Ramakrishnan, Prathamesh Wagh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29096v1 Announce Type: cross Abstract: The ongoing changes in software engineering requirements have created a substantial need for automated tools which can create secure source code from natural language input. The performance of traditional Large Language Models (LLMs) becomes limited ...
238. Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding ​
Author: Kejia Zhang, Tianyuan Zou, Zixuan GU, Yang Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.29111v1 Announce Type: cross Abstract: Applications such as personalized assistance and proprietary document analysis require large language models (LLMs) to generate outputs from private data. Yet powerful LLMs typically cannot be deployed on the resource-constrained devices where privat...
239. CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation ​
Author: Yuxiang Xiao, Xibei Chen, Xin Zhou, Jie Chen, Yifeng Zhang, Guillaume Sartoretti
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.29114v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) requires agents to reason over accumulated observations while continuously exploring unseen regions. However, existing environment representations often struggle to jointly support explicit semantic memory and con...
240. HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding ​
Author: Dongwook Lee, Sangkwon Park, Eunwoo Song, Che Hyun Lee, Youngho Cho, Junho Kim, June Young Yi, Heeseung Kim, Sungroh Yoon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD
arXiv:2608.29120v1 Announce Type: cross Abstract: Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchi...
241. Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs ​
Author: Eunna Lee, Soomyoung Lee, Jungpyo Nam, Heonjin Ha, Jamin Jung, Kyunam Choi, Sunjun Hwang, Yeonghun Kim, Seok-Jae Lim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.29136v1 Announce Type: cross Abstract: We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury of unstated severity following an interpersonal conflict---we present four frontier models with match...
242. STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation ​
Author: Shingeon Kim, Hyeyoon Lee, Dain Kwon, Kanghyun Choi, Sunjong Park, Mi-Ryang Kim, Jeong-Eun Lee, Jinho Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29145v1 Announce Type: cross Abstract: The rapid expansion of low Earth orbit satellites such as Starlink is increasingly contaminating astronomical surveys. In practice, contaminated images are often identified through inspection. However, modern surveys generate terabytes of data each n...
243. Training-Free Hidden-State Refinement for Flow-Matching Image Generators ​
Author: Yuanyi Yan, Xinzhe Rao, Canyu Shen, Yang Chen, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29160v1 Announce Type: cross Abstract: We aim to improve frozen flow-matching image generators by adding inference computation inside the denoiser, without changing model weights or the outer sampler. Existing generators usually spend extra test-time computation by increasing the number o...
244. Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA Challenge ​
Author: Kai Geissler, Raphael Sch"afer
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.29162v1 Announce Type: cross Abstract: We describe the submission of team FME to the MAMA-MIA Challenge, which evaluated primary tumor segmentation and prediction of pathological complete response (pCR) from pretreatment dynamic contrast-enhanced breast MRI on an external multi-country co...
245. TAAL: Mitigating Early Beam Pruning in Generative Recommendation via Temporal Autoregressive Alignment ​
Author: Lianjie Li, Zhiying Tu, Dianhui Chu, Hongliang Sun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.29179v1 Announce Type: cross Abstract: Generative recommendation encodes items as hierarchical semantic identifiers (SIDs) and retrieves the next item through autoregressive decoding. Standard next-token prediction, however, does not explicitly cover the multimodal transitions present in ...
246. Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space ​
Author: Qiancheng Zhou, Ruizhe Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.29188v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a...
247. Rate-Coding Bundle Memory: A Unified Model of Memory and Control for Symbolic Computation in the Brain ​
Author: Teun van Gils, Rowan P. Sommers, Markus Ostarek, Peter Hagoort
Published: 9/1/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CL, cs.SC
arXiv:2608.29189v1 Announce Type: cross Abstract: We propose a neurobiologically plausible model of cognition that combines the advantages of connectionist and symbolic systems, and that can explain a wide range of cognitive phenomena. This model, called Rate-Coding Bundle Memory (RCBM), is based on...
248. PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC ​
Author: Max Yu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT
arXiv:2608.29197v1 Announce Type: cross Abstract: Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pok'emon in its official do...
249. AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent ​
Author: Jonan Richards, Kosei Horikawa, Youmei Fan, Yutaro Kashiwa, Mairieli Wessel
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.29204v1 Announce Type: cross Abstract: Generative AI-based software engineering agents are becoming routine contributors to real-world software projects. On GitHub, developers can assign tasks to the Copilot cloud agent, which autonomously explores the repository, edits code, runs command...
250. Background-Free Objectness Learning for Class-Agnostic Detection ​
Author: Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2608.29232v1 Announce Type: cross Abstract: Object detectors are typically trained under closed-set supervision, where unlabeled regions are implicitly treated as background. Under incomplete annotations, this assumption introduces objectness bias: visually valid but unlabeled objects are used...
251. AGRICAM: A Track-Mounted Crop Pollination Monitoring Robot ​
Author: Malika Nisal Ratnayake, Adel N. Toosi, James Cook, Romina Rader, Alan Dorin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO, cs.SY, eess.SY, q-bio.QM
arXiv:2608.29237v1 Announce Type: cross Abstract: Insect pollination is critical for global food production, yet monitoring pollinators at commercial farm scale remains a challenge. Recent advances in computer vision and deep learning have enabled detailed analysis of pollinator behaviour, but monit...
252. QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation ​
Author: Yaroslav Prytula, Anton Popov, Dmytro Fishman
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.29253v1 Announce Type: cross Abstract: Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that produce weak boundaries and mixed visual evidence in overlap regions. Existing methods address this through local regions of interest...
253. Adaptive Multi-Branching for Shallow Decision Tree Induction ​
Author: Hanul Park, Jeonghoon Choi, Juseong Kim, Sanghun Sel, Giltae Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29262v1 Announce Type: cross Abstract: Decision trees are attractive for tabular prediction tasks because each prediction follows an interpretable sequence of feature-threshold tests. Under a strict maximum-depth budget, however, conventional binary trees can be under-expressive, since ea...
254. Measurement Validity in LLM Cultural Alignment ​
Author: An Duy Nguyen, Muhammad Aurangzeb Ahmad
Published: 9/1/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI
arXiv:2608.29266v1 Announce Type: cross Abstract: Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. W...
255. SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization ​
Author: Hojae Han, Jongyoon Kim, Sanghyuk Park, Dongwook Cheon, Myungjae Jeon, Sunjong Choi, Soonho Kong, Wonseok Heo, Seung-won Hwang, Donghoon Hyeon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29270v1 Announce Type: cross Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements written in ...
256. RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation ​
Author: Yuhan Li, Xianfeng Tan, Fangao Zeng, Wenxiang Shang, Pipei Huang, Hao Zhou, Zhiyu Jin, Wenjun Zhang, Bingbing Ni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29280v1 Announce Type: cross Abstract: Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic topological accuracy and microscopic physical fidelity. Although our ...
257. AOI-Net: Structural Face AOI-Guided Eye-Gaze Track Representation Learning for Autism Spectrum Disorder Detection ​
Author: Zhanpei Huang, Binbin Sun, Jialiang Chen, Yiou Wang, Taochen Chen, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.ET, cs.MM
arXiv:2608.29289v1 Announce Type: cross Abstract: Eye-movement tracking has emerged as a promising non-invasive approach to Autism Spectrum Disorder (ASD) screening, with systematic differences in attentional allocation and revisit behaviors observed during socially interactive tasks. Existing compu...
258. When Do Larger Batches Help Scale LLM Reinforcement Learning? ​
Author: Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29296v1 Announce Type: cross Abstract: Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this statistical benefit translates into lower wall-clock time-to-target remains unclear, because each update c...
259. Learning Simple Test-Time Environments for LLM Web Agents ​
Author: Junxuan Li, Zijun Liu, Ziyi Huang, Peng Li, Yuzhou Liu, Ming Yan, Yang Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.29305v1 Announce Type: cross Abstract: Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real-world settings. Existing research largely attribute this deg...
260. Detecting and Repairing Hallucinations in Retrieval-Augmented Generation ​
Author: Sai Krishna Reddy Mulakkayala, Niki van Stein, Aske Plaat
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29307v1 Announce Type: cross Abstract: Language models increasingly answer questions by consulting retrieved documents rather than memory alone, a design now common in search assistants and enterprise knowledge tools. Grounding a model in retrieved text reduces unsupported statements but ...
261. Improving Randomized Metric Distortion to 2.3282 ​
Author: Nisarg Shah
Published: 9/1/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2608.29308v1 Announce Type: cross Abstract: In metric social choice, each voter ranks a set of $m$ candidates by her distance to them in an unknown metric space. The cost of a candidate is its average distance to the voters. A randomized voting rule must use only the rankings to choose a lotte...
262. Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition ​
Author: Angel Yanguas-Gil
Published: 9/1/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, physics.app-ph
arXiv:2608.29309v1 Announce Type: cross Abstract: This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the key technologies be...
263. Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase ​
Author: Daegyu Sung, Yukyeong Lee, Geon Park, Yumin Choi, Sung Ju Hwang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2608.29310v1 Announce Type: cross Abstract: Organizations often develop and maintain portfolios of related applications: independently deployable codebases that share substantial domain logic, interface patterns, or operational conventions. As LLM coding agents are increasingly used to generat...
264. Hyper3-CLIP: Hierarchy-Conditioned Hyperbolic Vision-Language Training ​
Author: Matin Mahmood, Antonio Rueda-Toicen, Mohamed ElBassat, Seifeldin Elkerdany, Weixing Wang, Gerard de Melo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29313v1 Announce Type: cross Abstract: CLIP-like vision-language models (VLMs) trained with contrastive objectives learn strong global image-text representations, but their Euclidean embeddings and global pooling fail to encode relational structure such as part-whole and parent-child rela...
265. SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling ​
Author: Christopher Tatsch, Yu Gu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.29315v1 Announce Type: cross Abstract: This work introduces Semantically-Guided Exploration (SGE), a modular exploration framework for ground vehicles that integrates pixel-level semantic segmentation into sampling-based waypoint selection and receding-horizon route optimization. Unlike c...
266. StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue ​
Author: Yuxiong Wang, Ziwei Lin, Bo Wang, Yu Zhang, Shiguang Ni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29326v1 Announce Type: cross Abstract: Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic replies but also coherent progression through a multi-turn support process. Existing resources often reduc...
267. Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration ​
Author: Mohamad Zbib, Ammar Mohanna
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29378v1 Announce Type: cross Abstract: Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal r...
268. Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollback ​
Author: Guanlong Wu, Dahui Li, Ke Jiang, Jianyu Niu, Cong Wang, Yinqian Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.29381v1 Announce Type: cross Abstract: AI agents are moving toward persistent, stateful execution across various applications, accumulating execution state and external effects that are costly to reconstruct after failures. Checkpoint and rollback (C/R) are becoming essential for recovery...
269. Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers ​
Author: Shao-An Yin, Mingyi Hong, Nicola Elia
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT, cs.MA, cs.RO
arXiv:2608.29388v1 Announce Type: cross Abstract: Recent machine learning research has increasingly focused on equilibrium analysis in non-cooperative games rather than solely on optimal solutions. Many such problems involve shared constraints and can be formulated as Generalized Nash Equilibrium Pr...
270. Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD ​
Author: Asra Aslam, Volodymyr Chapman, Maurice M. O'Connell, Aseel S. Abuzour, Michael Abaho, Danushka Bollegala, Gary Leeming, Eduard Shantsila, Andrew Clegg, Lauren E. Walker, Iain Edward Buchan, Samuel D. Relton
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29419v1 Announce Type: cross Abstract: Deep learning architectures are increasingly proposed for patient trajectory modeling in electronic health records (EHRs), yet their advantage over simpler, more interpretable models is rarely subjected to rigorous empirical scrutiny in real-world cl...
271. Polis: 3D Self-Supervision at City Scale ​
Author: Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Fr'ed'eric Kaplan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.29426v1 Announce Type: cross Abstract: Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonomous systems, and heritage conservation. However, urban scenes of large spatial extent captured throug...
272. Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations ​
Author: Fabio F. Oberweger, Michael Schwingshackl
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.29434v1 Announce Type: cross Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, a...
273. SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning ​
Author: Guangyuan Wang, Mads Toftrup, Sebastian Loeschcke, Yixuan Wang, Anima Anandkumar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML
arXiv:2608.29448v1 Announce Type: cross Abstract: Physics-informed neural networks (PINNs) often face ill-conditioned objectives that limit high-accuracy training. Dense quasi-Newton methods improve local conditioning but require expensive optimizer state, while Kronecker-factored methods such as SO...
274. AI Can Be Easily Persuaded in Clinical Decision Making ​
Author: Jiayuan Zhu, Jiazhen Pan, Fenglin Liu, Minhao Hu, Junde Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29453v1 Announce Type: cross Abstract: As AI becomes increasingly integrated into clinical practice, it is playing a growing role in medical decision making. Medicine, however, is a high stakes and evidence based field, where decisions can directly affect patients' lives. It is therefore ...
275. Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities ​
Author: Linh Le, Hong Kiat Tan, David Williams-King
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29458v1 Announce Type: cross Abstract: Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hid...
276. Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation ​
Author: Johanna Angulo, V'ictor Yeste, Hector Espinos-Morato
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2608.29463v1 Announce Type: cross Abstract: A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equival...
277. Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered ​
Author: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29464v2 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences thr...
278. Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance ​
Author: Kazuyuki Hara, Hideitsu Hino
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29472v1 Announce Type: cross Abstract: Knowledge distillation trains a small student model to reproduce the outputs of a large teacher model, and its progress is typically monitored through the teacher--student discrepancy. The quantity of ultimate interest, however, is the student's erro...
279. MUDDLE: Measuring Understanding of Documents under Distractor and Length Effects ​
Author: Jason Luo, Saibilila Abudukelimu, Judy Song, Andrew Feng, Shivank Garg, Vasu Sharma, Kevin Zhu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29477v1 Announce Type: cross Abstract: Document question-answering systems increasingly answer questions over collections of retrieved documents rather than one clean source, so robustness to distracting context matters as much as reading ability. When such systems fail, it is often uncle...
280. Applications of Risk Science to AI Fairness Evaluation: Principles, Challenges, and Best Practices ​
Author: Kyra Wilson, Sabrina Kang, Saloni Dash, Aylin Caliskan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.29478v1 Announce Type: cross Abstract: Scholarly work which aims to describe potential societal impacts (e.g., risks) of proliferating technology (especially related to artificial intelligence or other algorithmic systems) is likely to have an impact beyond the scientific communities it w...
281. Generalizable Multi-Agent Planning from Signal Temporal Logic Specifications via Diffusion ​
Author: Joe Eappen, Zikang Xiong, Shreyash S. Iyengar, Suresh Jagannathan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.RO
arXiv:2608.29490v1 Announce Type: cross Abstract: Multi-agent systems in the real-world (e.g., drone swarms, autonomous cars, warehouse robots) must satisfy rich, temporal tasks while avoiding collisions. Signal Temporal Logic (STL) elegantly encodes such objectives, but current STL planning methods...
282. Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion ​
Author: Runyu Zhang, Jiawei Zhang, Gioele Zardini, Saurabh Amin, Asuman Ozdaglar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2608.29507v1 Announce Type: cross Abstract: Diffusion models are increasingly used not only for sampling from learned data distributions, but also for generating samples that optimize task-specific objectives. A common approach is to guide the reverse diffusion process using gradients of an ex...
283. On the Plasticity Collapse in Continual Machine Unlearning ​
Author: Yingdan Shi, Xiang Xu, Kaize Ding, Alfred O. Hero, Ren Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29513v1 Announce Type: cross Abstract: Machine unlearning enables deep neural networks to selectively remove the influence of specific data in response to privacy and regulatory requirements. While prior work largely studies single-shot unlearning, real-world systems must accommodate cont...
284. Argument-Aware Semantic Alignment of Normative Texts: A Toulmin-Based Neuro-Symbolic Approach ​
Author: William Schroeder
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29529v1 Announce Type: cross Abstract: Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstraction. Lexical overlap, distributional embeddings, and semantic similarity capture topical relatedness...
285. The Emergent Symbolic Structure of Artificial Neural Networks ​
Author: R. Thomas McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29530v1 Announce Type: cross Abstract: Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited. Intelligence has traditionally been modeled as operating over structured combinations of symbols, such as logical formulas. However, the strong...
286. Integrating adaptive human behavior into epidemic models with large language models ​
Author: Yicheng Mao, Haoyang Li, Rob Deardon, Hongru Du
Published: 9/1/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI
arXiv:2608.29535v1 Announce Type: cross Abstract: Infectious disease transmission is shaped by patterns of human interaction, which adapt as epidemic conditions change. Capturing these context-dependent behaviors remains a fundamental challenge for epidemic models. Here, we recast this challenge by ...
287. AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies ​
Author: Hongbo Gao, Zeyu Ni, Xin Wen, Siyu Xu, Ruifeng Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.29537v1 Announce Type: cross Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natura...
288. Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift ​
Author: Yujia Liu, Jiayan Lin, Zijin Hong, Zheng Yuan, Shengyuan Chen, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.29543v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have established conversational text-to-SQL as a practical interface between users and databases, often involving multiple turns of clarification and revision. However, existing benchmarks primarily eva...
289. PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation ​
Author: Lingfeng Yao, Chenpei Huang, Xingke Yang, Ziye Geng, Changqing Luo, Hao Wang, Jiang Liu, Miao Pan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM
arXiv:2608.29549v1 Announce Type: cross Abstract: Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may pro...
290. HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning ​
Author: Yibo Gong, Cong Guo, Jiacheng Ding
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC
arXiv:2608.29563v1 Announce Type: cross Abstract: School coaches prepare for opponents with game film and intuition. The analytics tools of professional teams stay out of reach. We ask how far public data can close this gap. Professional basketball is our case study, chosen for its data rather than ...
291. SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs' Robustness to Contradictory Evidence ​
Author: Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29582v1 Announce Type: cross Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigating the complex, non-bijective mappings between clinical indicators and diagnoses. Existing benchmarks...
292. Wide Learning: Learning to Reach Evidence ​
Author: Junzhou Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29608v1 Announce Type: cross Abstract: Machine learning is usually evaluated after an evidence interface has been fixed. A dataset, sensor suite, query language, action set, or experimental protocol determines which observations can be obtained, and learning is judged by what it extracts ...
293. Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps ​
Author: Sagar Srinivas Sakhinana, Venkataramana Runkana
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG
arXiv:2608.29615v1 Announce Type: cross Abstract: Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipel...
294. Memory-First Fact-Checking: A Knowledge-Graph-Grounded Multi-Agent System for Misinformation Detection ​
Author: Amelia Petrenciuc, Alexandru Lecu, Adrian Groza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29617v1 Announce Type: cross Abstract: This paper introduces a hybrid fact-checking framework that integrates Knowledge Graph-based semantic memory with adversarial multi-agent reasoning for explainable misinformation detection. The proposed system follows a memory-first, web-fallback arc...
295. CineForge: Self-Improving Agents for Long-Horizon Video Generation ​
Author: Junxiang Liu, Lin Wang, Haiyu Shi, Hongxu Ma, Xiaoyu Yang, Chunjie Chen, Xiaoxiao Xu, Kaiqiao Zhan, Boao Wang, Shuizhou Shi, Tianyun Zhu, Jie Li, Jiangtong Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29621v1 Announce Type: cross Abstract: Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems prim...
296. AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing ​
Author: Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.29622v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi-step reasoning that requires adaptive retrieval and continuous revision of intermediate contexts...
297. MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation ​
Author: Yangsong Lan, Renkai Hu, HongKai Zheng, Bo Zhang, Renzhi Wang, Hongliang Dai, Piji Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29623v1 Announce Type: cross Abstract: Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long ...
298. LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal Forecasting ​
Author: Di Zhang, Jingyang Zhang, Ziqian Wang, Chi Zhang, Yikun Ban, Ziwei Zhang, Ruijie Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29640v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise for spatio-temporal forecasting, but existing approaches often rely on regularly sampled token sequences and struggle with irregular observations because of temporal asynchrony, representation-space mis...
299. Conducting Stylistic Analysis of Paintings through an Art-History Agent ​
Author: Marc S. Walton, Astrid Harth
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.29644v1 Announce Type: cross Abstract: Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art history. By contrast, current artificial intelligence (AI) models used in the field offer only unexplain...
300. Cost-Effective Repository Exploration for Agentic Issue Localization ​
Author: Mohammad Nour Al Awad, Sergey Ivanov
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.29675v1 Announce Type: cross Abstract: Repository exploration is a distinct and costly stage of coding-agent pipelines: before generating a patch, an agent must identify which repository files are likely to matter. We study whether this stage can be delegated to lower-cost models while re...
301. MedSegBenchmarker: A Raw-Count-First Framework for Controlled 2D Medical Image Segmentation Benchmarks ​
Author: Vanessa Borst, Lukas Horn, Daniel Grillmeyer, Thomas Prantl, Samuel Kounev
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29677v1 Announce Type: cross Abstract: Despite rapid advances in MIS, fair and reproducible comparisons of segmentation models remain challenging due to heterogeneous datasets, inconsistent evaluation protocols, and rapidly evolving architectures. In particular, comparisons often implicit...
302. A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction ​
Author: Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29705v1 Announce Type: cross Abstract: Feed-forward 3D reconstruction models emit a per-pixel confidence that downstream systems read as a reliability signal. It is trained as a loss weight, not as an uncertainty magnitude, and whether it can be used as an error prediction has not been me...
303. Higher-Dimensional Rotary Position Embedding ​
Author: Yixing Li, Ruobing Xie, Yudong Zhang, Yushi Bai, Samm Sun, Yu Cheng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.29715v1 Announce Type: cross Abstract: Transformers rely on position embedding mechanisms in long context modeling in most cases. Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in self-attention. However, its p...
304. SynCrash: A Multi-Stage Pipeline for Zero-Shot Accident Detection and Localization in Traffic Surveillance Video ​
Author: Arkya Jyoti Bagchi, Ritul Jangir, Varun Raskar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29759v1 Announce Type: cross Abstract: We present SynCrash, a multi-stage pipeline for zero-shot accident detection, spatial localization, and collision-type classification in fixed-view CCTV surveillance video. Our approach addresses the ACCIDENT at CVPR 2026 Challenge, which requires pr...
305. R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment ​
Author: Mohan Zhang, Chengsong You, Xiaoyu Cao, Zhen Sun, Xiaohan Jia, Junwei Zhou, Yongchao Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29798v1 Announce Type: cross Abstract: The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across tasks. We introduce the Persona Selection--Realization Framework, which models behavior generation ...
306. REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling ​
Author: Devrim \c{C}avu\c{s}o\u{g}lu, Emre Akba\c{s}
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.29899v2 Announce Type: cross Abstract: Dense retrieval over long documents is expensive. Token-level encoders scale quadratically in sequence length, and most long-context embedding models reach 32K tokens only through architectural workarounds or by stretching billion-parameter LLMs. We ...
307. INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction ​
Author: Shahar Oded, Yuval Shahar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.29901v1 Announce Type: cross Abstract: Electronic Health Record (EHR) prediction models in the intensive care unit must learn from sparse and irregular measurements while preserving the clinical meaning of time and supporting transparent decision-making. We present INTERVenE, a family of ...
308. When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection ​
Author: Xiaoyang Han, Lvxiaowei Xu, Ming Cai
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29903v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has made AI-generated text detection increasingly critical. Existing zero-shot detectors assume that more token-level evidence leads to more reliable detection. However, our empirical study challe...
309. IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil ​
Author: Bhaskar Ganesh Devalla, Junchao Wu, Nilesh Dokuparthi, Greeshma Yaluru, Tatiana Muniz Rodriguez, Lidia S. Chao, Derek F. Wong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29919v1 Announce Type: cross Abstract: The rapid proliferation of LLMs has further heightened the need to develop dependable AI-generated text detection, especially beyond English. Nevertheless, current benchmarks pay little attention to Indic languages and test detectors in idealized set...
310. Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With? ​
Author: Alberto Cetoli
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29921v1 Announce Type: cross Abstract: The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a...
311. Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization ​
Author: Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu, Prashant Khanduri
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.29924v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) are prone to hallucinations: they fluently describe objects, attributes, and scenes that are not in the image. We connect part of this failure to a measurable property of their representations, feature instability...
312. Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents ​
Author: Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.29942v1 Announce Type: cross Abstract: The key limitation of current state-of-the-art influence-based guardrails is that they do not reliably distinguish a legitimate, user-authorized action from a malicious, unauthorized action when both rely on external tool information. This ambiguity ...
313. The Policy Deficit in AI x Social-Emotional Learning Research ​
Author: Tran Van Cuong, Liu Yihan, Nguyen Van Tuong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.29950v1 Announce Type: cross Abstract: As artificial intelligence (AI) is increasingly integrated into social-emotional learning (SEL) initiatives, the need for evidence-based policy has become paramount. We systematically reviewed 65 peer-reviewed papers that examine the intersection of ...
314. Training-Free Action Correction for VLA Model Failures via Language Feedback ​
Author: Owen Kwon, Pablo Ortega-Kral, Arthur Bucker, Jean Oh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.29967v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly unde...
315. Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models ​
Author: Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2608.29990v1 Announce Type: cross Abstract: Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is conseque...
316. Generating Clinical Vignettes that Preserve Cognitive Formulations ​
Author: Amit Oren, Nimrod Hertz-Palmor, Dean Ariel, Guy Laban
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29995v1 Announce Type: cross Abstract: Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a di...
317. Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models ​
Author: Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan, Rhongho Jang, Dongxiao Zhu, Prashant Khanduri
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.29996v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background correlations, resulting in predictions driven by contextual shortcuts rather than object-relevant vis...
318. TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models ​
Author: Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG
arXiv:2608.29999v1 Announce Type: cross Abstract: Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for downstream tasks like speech recognition and dense audio captioning, times...
319. Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models ​
Author: Hermione Warr, Harry Anthony, Lilli J Freischem, Yasin Ibrahim, Daniel R McGowan, Konstantinos Kamnitsas
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30021v1 Announce Type: cross Abstract: Errors in radiology reports can adversely affect patient treatment, yet automated report quality assurance remains challenging because errors are often subtle and require domain expertise to detect. Although large language models (LLMs) have recently...
320. "Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators ​
Author: Krisztian Balog, Arild Michel Bakken
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30033v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, suffering from a "superhuman bias." Using a dataset of over 71,000 reading comprehension responses from 2,35...
321. Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection ​
Author: Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.30041v1 Announce Type: cross Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations...
322. Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction ​
Author: Jo~ao L. P. Santana, Filipe R. Cordeiro
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30046v1 Announce Type: cross Abstract: Noisy labels remain a critical challenge for training deep neural networks, since memorizing incorrect labels degrades generalization. Once noisy samples are identified after training, the standard solution is to retrain the model from scratch on the...
323. Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark ​
Author: Abdullah Hashmat, Usman Naseem, Agha Ali Raza
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30065v1 Announce Type: cross Abstract: Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignment in English-centric settings, but these gains transfer poorly to low-resource languages due to cultural mismatches. Existing multilingual 3H benchmar...
324. How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account ​
Author: Ruize Xu, Xiao Yu, Yujin Tang, Chenming Shang, Nikhil Singh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.30067v1 Announce Type: cross Abstract: How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this ...
325. Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer ​
Author: Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30092v2 Announce Type: cross Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this project. O...
326. Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization ​
Author: Jessica D. Elrefaei, Kaixun Hua, Seungbae Kim, Hoang Nam Tran, Juan S. Borrero
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30103v1 Announce Type: cross Abstract: Bilevel mixed-integer linear optimization problems model hierarchical decision processes in which a leader anticipates the optimal response of a follower. Although expressive, these problems are computationally challenging because lower-level optimal...
327. AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP ​
Author: Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai, Bontu Fufa Balcha, Zayd Bashir, Angana Borah, Zara Burzo, Yubin Choi, Naihao Deng, Samika Gupta, Michel Faloughi, Claude Kwizera, Ziqiao Ma, Cynthia Yacel Fuertes Panizo, Ellie Seehorn, Hui Shen, Jiayi Tang, Zesen Zhao, Boyuan Zheng, Rada Mihalcea
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.30107v1 Announce Type: cross Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level represen...
328. Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators ​
Author: Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30110v1 Announce Type: cross Abstract: Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of exper...
329. Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving ​
Author: Tian Zhang, Zhuo Huang, Hongrui Ye, Yu Wu, Zengmao Wang, Kaixuan Zhou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.30122v1 Announce Type: cross Abstract: Vision-language-action (VLA) driving methods increasingly combine multi-trajectory imitation learning with group-relative policy optimization (GRPO), making trajectory selection critical to final performance. However, some high-scoring trajectories t...
330. TPR-Attention for Combinatorial Generalization ​
Author: Melisa Civeleko\u{g}lu, Isabeau Pr'emont-Schwarz
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2608.30124v1 Announce Type: cross Abstract: Systematic generalization remains a significant challenge in deep learning. In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortless for humans but difficult for standard neural ar...
331. VIBE: Video Instruction-aligned Background music gEneration ​
Author: Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj, Gouthaman KV, Sreyan Ghosh, Ramani Duraiswami, Lie Lu, Dinesh Manocha
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, cs.CV, cs.LG
arXiv:2608.30125v1 Announce Type: cross Abstract: Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Auto...
332. E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval ​
Author: Yerang Kim, Jiyoon Myung, Joohyung Han
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.30130v2 Announce Type: cross Abstract: Retrieval-augmented language models can fail to respect negative constraints when the retriever supplies evidence about concepts the user explicitly excluded. Beyond explicit negation, queries may ask for answers that include one concept while exclud...
333. CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation ​
Author: Kwangmin Ki, Yunhun Nam, Jongheon Jeong, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30158v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is the de facto standard for adapting large language models (LLMs) to target domains, but it often degrades the model's general capabilities, a phenomenon known as catastrophic forgetting. Existing approaches typically mo...
334. Science sandboxes measure the scientific capability of AI agents ​
Author: Arya S. Rao, Rodrigo I. Castro, Sager J. Gosai, Kenneth B. Hsu, Yasha Ektefaie, Shantanu Singh, Sangeeta N. Bhatia, Steven K. Reilly, Ryan Tewhey, Eric S. Lander, Pardis C. Sabeti
Published: 9/1/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.30165v1 Announce Type: cross Abstract: Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI ...
335. SIR: Self-improving Red-teaming for Compute Use Agents ​
Author: Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.30207v1 Announce Type: cross Abstract: Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks. Because they can be exposed ...
336. Label Semantic Expansion via Label Guided Neural Topic Modeling ​
Author: Haojia Zheng, Yuyin Lu, Juntian Huang, Fan Ou, Yanghui Rao, Haoran Xie, Fu Lee Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30216v1 Announce Type: cross Abstract: Topic models are widely used for content analysis, where users often analyze corpora around predefined labels rather than unordered latent topics. Existing label-aware topic models mainly follow a labels-for-topics perspective, using labels to guide ...
337. The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce ​
Author: Cheng Lyu, Jingyue Zhang, Vinny DeGenova, Mengwei Li, Yuanli Pei
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30224v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to annotate structured product data in e-commerce, but early deployment often begins as a cold-start problem: only limited pre-launch labels are available, the value of expensive reasoning is unknown...
338. Motus2: A Self-Evolving General World Model for Dexterous Manipulation ​
Author: Hongzhe Bi, Zihao Zhou, Yihang Tang, Jingrui Pang, Shuhe Huang, Haitian Liu, Runqing Wang, Shuai Huang, Yichen Wang, Yiming Cheng, Ruowen Zhao, Zhenghua Li, Hengkai Tan, Xiaolong Liu, Jinhui Wan, Jiabao Liu, Min Zhao, Fan Bao, Jun Zhu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2608.30237v1 Announce Type: cross Abstract: General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, with...
339. Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs ​
Author: Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30256v1 Announce Type: cross Abstract: Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be ...
340. Stratified Consistency Distillation for Natural Language Formalization ​
Author: Zhichao Hou, Ferhat Erata, Joe Lilien, MohamadAli Torkamani
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30258v1 Announce Type: cross Abstract: Neurosymbolic reasoning has shown promising success in addressing complex reasoning tasks by combining large language models (LLMs) and symbolic solvers. While this approach shows promise, a fundamental challenge remains: improving the accuracy of tr...
341. Using Prosody to Predict Syntactic Structure ​
Author: Junghyun Min, Alex Warstadt, Tamar I. Regev, Tiago Pimentel, Ethan Gotlieb Wilcox
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30260v1 Announce Type: cross Abstract: While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two domains remains contested. We investigate the syntax-prosody interface through an information-theoretic ...
342. Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs ​
Author: Shunjie Wen, Jaeyeon Lee, Dong-Wan Choi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30263v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) incur substantial inference costs due to their long and highly redundant visual-token sequences. Diversity-based pruning mitigates this cost by selecting token subsets based on pairwise cosine similarity. We find,...
343. Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization ​
Author: Yiyang Lu, Mohammad Pedramfar, Vaneet Aggarwal
Published: 9/1/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG
arXiv:2608.30271v1 Announce Type: cross Abstract: We study decentralized online optimization of upper-linearizable payoffs over an action set under efficient separation access, with applications to online continuous diminishing-return (DR) submodular maximization. We propose Decentralized Barrier Fo...
344. BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning ​
Author: Dongsheng Hou, Yanqiao Chen, Yuhan Rui
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30283v1 Announce Type: cross Abstract: Expected-cost constraints can still permit rare, high-cost events. Monte Carlo conditional value at risk (CVaR) gradients can be noisy at high confidence, whereas critics that model an outcome distribution add complexity. We propose BCPPO (Bachelier-...
345. CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration ​
Author: Haoyun Jiang, Haolin Li, Jianwei Zhang, Fei Huang, Qiang Hu, Minmin Sun, Shuai Xiao, Yong Li, Junyang Lin, Jiangchao Yao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30295v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that ...
346. ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation ​
Author: Jiawei Zhang, Hongsong Wang, Pan Zhou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30307v1 Announce Type: cross Abstract: Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators often yield geometrically inv...
347. Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs ​
Author: Yirui Liu, Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30310v1 Announce Type: cross Abstract: Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structure complicates prefix caching: full-attention key-value caches are token-addressable, whereas linear-...
348. One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread ​
Author: Zhuoran Lu, Weilong Wang, Yangyang Yu, Xinru Wang, Zhuoyan Li, Zhiwei Liu, Sophia Ananiadou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SI
arXiv:2608.30311v1 Announce Type: cross Abstract: Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-making, these indicators are embedded in information spread: users see both an AI prediction and earlier j...
349. Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance ​
Author: Donggyu Min, Dong-Kyu Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30317v1 Announce Type: cross Abstract: Online dynamic origin-destination (OD) matrix estimation (DODE) calibrates time-dependent OD demand to reproduce observed link-flow trajectories. In online, OD demand should be estimated from current observations and propagated network states while s...
350. Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering ​
Author: Jin Gan, Xin Li, Jun Luo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30319v1 Announce Type: cross Abstract: Large language models (LLMs) finetuned for specialized domains represent crucial high-impact applications. Inference-time alignment improves safety degraded from specialization finetuning without requiring substantial computational resources, complem...
351. Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends ​
Author: Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.30326v1 Announce Type: cross Abstract: Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parallel band-split enhancement fron...
352. Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation ​
Author: Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.30333v1 Announce Type: cross Abstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one c...
353. PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies ​
Author: Botong Zhao, Fang Yu, Tim, Senhua Zhu, Xinyuan Chen, Yue Lu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.30378v1 Announce Type: cross Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: their representations are not explicitly required to describe how the scene evolves over multiple time ...
354. DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving ​
Author: Yanqi Yu, Pingwei Sun, Jianchao Tan, Tao Zhang, Yuchen Xie, Xunliang Cai, Yao Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30386v1 Announce Type: cross Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. However, their in-place recurrent-state updates complic...
355. PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning ​
Author: Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim, Yong Man Ro
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30388v1 Announce Type: cross Abstract: Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as a unified embedding, where vi...
356. Using Grounded Theory for Agent Behavior Analysis at Scale ​
Author: Zhuoran Lu, Yangyang Yu, Zhuoyan Li, Yibo Meng, Nan Jiang, Chengxi Zang, Jie Gao, Ziang Xiao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30391v1 Announce Type: cross Abstract: Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, often unfamiliar tasks where pre-built classifiers fall short. We propose to bring grounded theory into agent trajectory analysis:...
357. TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information ​
Author: Dain Kwon, Kanghyun Choi, Hyeyoon Lee, Sunjong Park, Seoyong Lee, Sukjin Kim, Jinho Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30394v1 Announce Type: cross Abstract: Existing GNN quantization methods suffer from considerable quantization overhead, which severely limits their practical usage in real-world scenarios. To this end, we present TopGQ, an accurate post-training GNN quantization framework, alleviating re...
358. SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation ​
Author: Yunqi Liu, Yang Zhang, Ruixing Zhang, Liangzhe Han, Yi Qiao, Tongyu Zhu, Leilei Sun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30399v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong semantic reasoning and open-ended generation abilities, but aligning these abilities with structured sequential generation remains challenging. This challenge is particularly evident in out-of-town (OOT) PO...
359. ImageCAS-X: a dataset and benchmark for coronary artery segmentation and centerline extraction in coronary CT angiography ​
Author: Kit M. Bransby, Esther {\O}ksnebjerg, Kristoffer Kj{\ae}r, Jacob Kirkeby, Yasmin El Youssef, A"ida Jim'enez, Philip R. Pedersson, Martina C. de Knegt, Klaus F. Kofoed, Rasmus R. Paulsen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30404v1 Announce Type: cross Abstract: Accurate segmentation of the coronary vessel lumen is a prerequisite for quantitative assessment of atherosclerotic plaque and perivascular adipose tissue in coronary computed tomography angiography (CCTA). Cardiologists rely on semi-automated method...
360. SePArate: Segmenting Patterns from Defects in Wafer Manufacturing Using Weak Supervision ​
Author: Dain Kwon, Changmin Shin, Sunjong Park, Kanghyun Choi, Hyeyoon Lee, Jaewon Jang, Minseok Choi, Jinho Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30410v1 Announce Type: cross Abstract: In semiconductor manufacturing, defect analysis is essential, but manual inspection cannot scale. However, existing automated inspection methods remain insufficient for root-cause analysis and process optimization. To this end, we present SePArate, a...
361. Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols ​
Author: Tiffanie Godelaine, Maxime Zanella, Karim El Khoury, Benoit Macq, Christophe De Vleeschouwer
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30420v1 Announce Type: cross Abstract: Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis increasingly relies on vision-language models that provide patch-level zero-shot predictions. However...
362. ObjectSplat: Improving Mesh Fidelity and Interactivity for 3D Scenes via Object-Level Mesh Splatting ​
Author: Minhas Kamal, Hiranya Garbha Kumar, Mahedi Kamal, Balakrishnan Prabhakaran
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.30423v1 Announce Type: cross Abstract: Splatting-based algorithms reconstruct photorealistic, real-time-renderable, and mesh-exportable 3D scenes from regular images, but they represent a scene as a single monolithic field. Therefore, the reconstruction has no object-level structure, leav...
363. Towards Cognitive Process-Aware Proactive Writing Support ​
Author: Masahiro Yoshida, Atsuya Kobayashi, Kei Tateno, Xiang 'Anthony' Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.30424v1 Announce Type: cross Abstract: Large language models can support writing, but existing tools require users to explicitly articulate prompts-particularly burdensome in creative writing, where intentions are often ambiguous. Proactive support that infers users' needs from writing in...
364. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions ​
Author: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG
arXiv:2608.30428v1 Announce Type: cross Abstract: Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, parti...
365. Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer ​
Author: Minju Song, Hyeon Hwang, Junhyun Lee, Jaewoo Kang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30462v1 Announce Type: cross Abstract: Large language models exhibit substantial performance variation across languages, even when solving semantically equivalent tasks. Existing analyses often treat this phenomenon as an observational disparity caused by differences in pretraining data, ...
366. ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation ​
Author: Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination detection in English and Moder...
367. Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework ​
Author: Zhanbo Zhang, Ming Liu, Qing Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30487v1 Announce Type: cross Abstract: Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measur...
368. Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability ​
Author: Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30505v1 Announce Type: cross Abstract: Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional matrix-centric view...
369. Lot Machine: Multimodal Lot Extraction from Auction Catalogs ​
Author: Mathias Zinnen, Alisha Mund, Sabine Lang, Lukas H"uttner, Thomas Gorges, Vincent Christlein
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.DL
arXiv:2608.30510v1 Announce Type: cross Abstract: For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highl...
370. Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems ​
Author: Cheng Gu, Qiusheng Zhao, Anbang Liu, Shaochong Lin, Max Z. J. Shen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2608.30512v1 Announce Type: cross Abstract: Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention. We study this problem in overhead hoist trans...
371. TSExplorer: An interactive data annotation and exploration tool for time-series data ​
Author: Einari Vaaras, Manu Airaksinen, Okko R"as"anen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.SE
arXiv:2608.30514v1 Announce Type: cross Abstract: We present TSExplorer, a cross-platform tool for interactive annotation and exploration of time-series data. The tool enables users to inspect high-dimensional datasets through multiple complementary 2D visualizations derived from high-dimensional fe...
372. Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval ​
Author: Gaoming Zhang, Angqing Jiang, Jianchun Song, Kena Qi, Dayao Chen, Wei Lin, Defu Lian
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.30553v1 Announce Type: cross Abstract: Generative Retrieval (GR) has emerged as a promising paradigm by mapping queries directly to Semantic IDs (SIDs) with powerful representation capabilities for candidate items. However, existing SIDs derived solely from item content create a semantic ...
373. Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs ​
Author: Deokjae Lee, Sihun Chu, Hyun Oh Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30564v1 Announce Type: cross Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed budget, but Mixture-of-Experts (MoE) models contain these layers in e...
374. Collapsibility of Performance Metrics in Clinical Predictive AI ​
Author: Jo~ao Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30568v1 Announce Type: cross Abstract: Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics a...
375. DiffSAC: Diffusion-guided Sampling for Consensus-based Robust Estimation ​
Author: Chang Nie, Guangming Wang, Zhe Liu, Hesheng Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30603v1 Announce Type: cross Abstract: Robust estimation is a core computer vision task frequently tackled using sample consensus. However, traditional methods suffer from inefficient sampling as they struggle to identify effective minimum sets before hypothesis evaluation. To address the...
376. Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster ​
Author: Songtao Fang, Zihao Xu, Shaowei Wei, Jin Zhang, Zhuojun Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.30606v1 Announce Type: cross Abstract: With the development of large language models (LLMs), generative retrieval is becoming increasingly important in e-commerce scenarios. Current mainstream approaches typically use a two-stage training strategy: first train a product embedding model, a...
377. Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training ​
Author: Lukas Borggren, Jenny Kunz, Marco Kuhlmann
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30609v1 Announce Type: cross Abstract: Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional training on target-domain corpo...
378. Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text ​
Author: Minkyung Cho, Jihyo Kim, SeungWoo Song, Junghun Yuk, Minjoon Kee, Hoyun Song, KyungTae Lim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30619v1 Announce Type: cross Abstract: Synthetic data is increasingly used to train large language models (LLMs), yet its security implications remain poorly understood. Prior work on subliminal learning suggests that models can inherit behavioral traits from seemingly unrelated training ...
379. Cost-efficient Active Learning for Referring Image Segmentation and Grounding ​
Author: Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim, Jeany Son
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30621v2 Announce Type: cross Abstract: Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones....
380. GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning ​
Author: Outongyi Lv, Yuanwei Zhang, Xiaoqun Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30632v1 Announce Type: cross Abstract: Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilities, demonstrating remarkable effectiveness across reasoning tasks. Re...
381. BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs ​
Author: Debarpan Bhattacharya, Malay Phadke, Sriram Ganapathy
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30646v2 Announce Type: cross Abstract: Reliable uncertainty estimation is a crucial requirement for deploying large language models (LLMs) and vision-language models (VLMs) in safety-critical settings, especially when the model parameters are not accessible (black-box). We propose BiG-SUR...
382. Fine-Grained Multi Image Object Hallucination Benchmark ​
Author: Joonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim, Kihyun Kim, Yohan Jo, Joonseok Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.30653v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factu...
383. CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy ​
Author: Wentao Li, Jiangjie Qiu, Yijun Li, Leyi Zhao, Xiaonan Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30674v1 Announce Type: cross Abstract: Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. Large language models (...
384. LCoT-GV: Graph Attention Networks for Verifying Long Reasoning Chains in Large Language Models ​
Author: B'er'enice Jaulmes, Mehwish Alam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30679v1 Announce Type: cross Abstract: Large Reasoning Models produce Long Chains-of-Thought (LCoTs) which involve breaking down the problem into smaller reasoning steps before reaching the conclusion. However, these steps often contain contradictions, unsupported inferences, or irrelevan...
385. Learning Materials Properties from Scarce Labels and Unlabeled Crystals ​
Author: Wentao Li, Yizhe Chen, Jiangjie Qiu, Yijun Li, Leyi Zhao, Xiaonan Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30682v1 Announce Type: cross Abstract: Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery. We present SemiMat, a controlled benchmark for semi-supervised materials property regression, and MatRank, a reliabili...
386. Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning ​
Author: Yue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing, Zhanxing Zhu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30699v1 Announce Type: cross Abstract: Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to degraded generalization. While many long-tailed semi-supervised learning (LTSSL) methods have been pr...
387. An Agentic Retrobiosynthesis Framework with Learned Frontier Selection ​
Author: Philippe Meyer, Guillaume Gricourt, Thomas Duigou, Joan H'erisson, Jean-Loup Faulon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30702v1 Announce Type: cross Abstract: Large language models are increasingly used as agents for multistep retrosynthesis, raising the question of how much their search policy contributes independently of the underlying reaction model. We investigate this question in a biological setting ...
388. SingProbe Technical Report ​
Author: Sing Team
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2608.30703v1 Announce Type: cross Abstract: Runtime guardrails are essential for reliable large language model (LLM) deployment, yet existing approaches typically rely on independent, external models that introduce additional inference cost, delayed safety signals, and a capacity mismatch with...
389. RailSyn: Diagnosis-Guided Image Generation for Traceable Data Completion in Railway Foreign Object Detection ​
Author: Quan Hao, Chenxi Zhang, Ziyang Tao, Yuyuan Zhou, Yudong Wang, Rui Shi, Lechuan Xu, Changhao Liu, Liguo Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30709v1 Announce Type: cross Abstract: Railway foreign object detection (RFOD) is critical to safe railway operation, yet scarce real positive samples incompletely represent task-relevant variations in object scale, intrusion relation, railway scene, illumination, and adverse weather. Exi...
390. BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks ​
Author: Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs, Julian Moncarz, Kaustubh Kislay, Juan J. Vazquez
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30724v1 Announce Type: cross Abstract: LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and t...
391. RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object Generation ​
Author: Quan Hao, Ziyang Tao, Chenxi Zhang, Yudong Wang, Rui Shi, Liguo Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30727v1 Announce Type: cross Abstract: Small-object detection under long-tailed data distributions is a fundamental yet challenging problem in multimedia. Railway Foreign Object Detection (RFOD) epitomizes this challenge with easily confused small intrusions and scarce samples. To address...
392. Calibrating Small Language Models for Claim Check-Worthiness Detection ​
Author: Pratuat Amatya, V Venktesh, Vinay Setty
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30731v2 Announce Type: cross Abstract: Assessing claim check-worthiness is an essential first step in automated fact-checking pipelines. This work is motivated by a real deployment challenge at an early-stage startup: running large language models (LLMs) over every incoming claim is cost-...
393. Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval ​
Author: Shaowei Wei, Chong Huang, Songtao Fang, Jin Zhang, Zhuojun Wang, Chengfu Huo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.30753v1 Announce Type: cross Abstract: In large-scale e-commerce retrieval, dual-encoder retrievers are op- timized for contrastive similarity, whereas downstream rerankers capture finer-grained relevance preferences; this objective mis- match limits end-to-end retrieval quality. Reinforc...
394. On the Prospects of Dynamic LLM Conversations in Software Development ​
Author: Annemarie Wittig, Alina Mailach, Janet Siegmund, Norbert Siegmund
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.30756v1 Announce Type: cross Abstract: Large language models (LLMs) have become an essential tool for assisting developers, yet we still lack knowledge on ways to effectively support their interactions during development activities. That is, the quality of interactions with a chat-based L...
395. Conjoint Audio-to-Spikes Encoding and Processing for Efficient Neuromorphic Speech Recognition ​
Author: Valentin M. Meunier, Am'elie Gruel, Pierre Lewden, Adrien F. Vincent, Sylvain Sa"ighi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG, eess.AS
arXiv:2608.30792v1 Announce Type: cross Abstract: Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solution to lower the energy cost of artificial intelligence. The current rarity of natively neuromorphic datasets promotes the development of soft...
396. Aggregate Disambiguation Systems ​
Author: Jos'e Mar'ia Lago, Albert Castellana, Edgars Nem\v{s}e
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ME, cs.AI
arXiv:2608.30805v1 Announce Type: cross Abstract: Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregate disambiguation systems (ADSs). Given a task and a candidate solution, each evaluator casts a binary...
397. A Composition-Aware Pretraining Framework for Geospatial Foundation Models ​
Author: Aryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath, Shrutilipi Bhattacharjee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30817v1 Announce Type: cross Abstract: Geospatial foundation models have emerged as state-of-the-art methods for downstream Earth observation tasks. However, existing pretraining methodologies process imagery through a single-concept lens, failing to capture the highly compositional natur...
398. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling ​
Author: Minghan Qin, Yuang Wang, Xiuyu Yang, Yushi Long, Yujian Zhang, Ruihuan Wang, Kai Ye, Yangang Zhang, Hang Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30821v1 Announce Type: cross Abstract: Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated indivi...
399. Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis ​
Author: Konstantinos Moutselos, Ilias Maglogiannis
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30835v1 Announce Type: cross Abstract: Background and Objective: Quality control is a prerequisite for whole-slide image analysis, yet the benchmarks on which quality-control methods are compared share four properties that make their reported differences hard to interpret: few independent...
400. Pretrained, Curriculum-Tuned, and Ensembled: A Tracer-Aware Interactive Segmentation Pipeline for AutoPET V ​
Author: Xinglong Liang, Chunyao Lu, Tianyu Zhang, Jiaju Huang, Tao Tan, Yunchao Yin, Lishan Cai
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30844v1 Announce Type: cross Abstract: Interactive lesion segmentation in whole-body PET/CT requires a model to provide a strong initial prediction while also responding efficiently to sparse corrective scribbles during inference. This setting is particularly challenging because tracer di...
401. TAMI: Temporally Aligned, Missingness-Aware, and Interpretable Multimodal Fusion for Mental Health Assessment in Older Adults with Mild Cognitive Impairment ​
Author: Merna Bibars, Bolaji Omofojoye, Allan I. Levey, Rachel Hershenberg, Gari D. Clifford, Hyeokhyen Kwon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30857v1 Announce Type: cross Abstract: Depression and anxiety in older adults with Mild Cognitive Impairment (MCI) are frequently underdiagnosed due to limited access to care. Multimodal analysis of remote clinical interviews is a scalable screening approach, but existing methods have thr...
402. Exponential random graph models with soft clique constraints ​
Author: Yasmin Tousinejad, Vera Koponen
Published: 9/1/2026, 4:00:00 AM
Categories: math.CO, cs.AI, math.PR
arXiv:2608.30869v1 Announce Type: cross Abstract: Let $r\geq3$ be fixed, and let $\mathbf{G}_n$ be the set of all simple graphs with vertex set $[n]={1,\ldots,n}$. We consider an exponential random graph model which gives higher probability to $G \in \mathbf{G}_n$ than to $H \in \mathbf{G}_n$ if $...
403. Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice ​
Author: Jinhee Won, Xinlan Emily Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30873v1 Announce Type: cross Abstract: LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We te...
404. Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models ​
Author: Melina Morch, Daniel Braun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30884v1 Announce Type: cross Abstract: While gender and racial biases in language models have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representat...
405. Safety Screening for Voltage Control in Active Distribution Grids via Distributionally Robust Conformal Screening ​
Author: Sarra Bouchkati, Petros Ellinas, Adriana Geisler, Steffen Kortmann, Johanna Vorwerk, Spyros Chatzivasiliadis, Andreas Ulbig
Published: 9/1/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.LG, cs.SY
arXiv:2608.30889v1 Announce Type: cross Abstract: Deploying a new control policy for voltage control in active distribution grids requires evidence that physical limits will be satisfied before the policy is tested on the physical grid. This assessment is difficult for two reasons. First, simulation...
406. Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods ​
Author: Sebastian Buschj"ager, Nuwan Gunasekara, Heitor Murilo Gomes
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF
arXiv:2608.30923v1 Announce Type: cross Abstract: Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. However, sustained operation of a stream learner also requires predictable and bounded resource usage even on long streams. This requirement becomes...
407. Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper ​
Author: Chanhee Cho, Junhyuk Choi, Bugeun Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.30927v1 Announce Type: cross Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a...
408. LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation ​
Author: Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.30935v1 Announce Type: cross Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spat...
409. MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI ​
Author: Laura Ib'a~nez-Mart'inez, Roser Batlle-Roca, Xavier Serra, Mart'in Rocamora
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CY, eess.AS
arXiv:2608.30940v1 Announce Type: cross Abstract: Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains underexplored. Prior work includes openness-focused evaluation frameworks, such as MusGO (Music-Genera...
410. Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse ​
Author: Mattia Cerrato, Otto Sahlgren, Xenia Heilmann
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.30956v1 Announce Type: cross Abstract: Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input features were manipulated. This technique is used for a range of tasks such as debugging models, exp...
411. LOCI: A Locator-Critic with Refinement Loop ​
Author: Walid Bousselham, Mathilde Caron, Arsha Nagrani, Cordelia Schmid
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30959v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) still struggle on tasks requiring complex visual understanding. We argue that the core issue is not high-level reasoning, but instead failing to locate critical details in the image. Due to this shortcoming, VLMs generat...
412. A Universal Context-Reuse Layer for Cross-Model KV Sharing ​
Author: Yi Li, Dongming Jiang, Yi Zhao, Bingzhe Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.30963v1 Announce Type: cross Abstract: Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. Existing KV-cache reus...
413. CogEvol: Towards Efficient and Reliable Learning Environment Generation ​
Author: Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30968v1 Announce Type: cross Abstract: We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k ...
414. CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations ​
Author: Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG, eess.AS, eess.SP
arXiv:2608.30974v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise tr...
415. MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI ​
Author: Athira J. Jacob, Puneet Sharma, Dorin Comaniciu, Daniel Rueckert
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.30975v1 Announce Type: cross Abstract: Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep learning approaches process individual 2D slices, discarding this context. We present MR-JEPA, a self-su...
416. Evaluating and Improving LLM Self-Modeling ​
Author: Siqi Zeng, Andre N. Assis, Rowan Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30980v1 Announce Type: cross Abstract: We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmar...
417. Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning ​
Author: Arthur Becker, Jakob Kemmler, David Thulke, Christine Sch"afer, Christian Dugast, Hermann Ney
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.30987v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) trains a base language model to imitate target responses, and these targets may require knowledge the base model has not robustly internalized. We study this as a source of hallucinations and frame a group of mitigation m...
418. LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It ​
Author: Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.31016v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the tr...
419. One note in three: a verified census of three deployed AI scribes, and the instrument that counted it ​
Author: Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.31017v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored ...
420. Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring ​
Author: Vanodhya G. Warnasooriya, Amir Hajian, Watchara Ruangsang, Supavadee Aramvith
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV
arXiv:2608.31074v1 Announce Type: cross Abstract: We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n-pose to detect persons and extract seventeen skeletal keypoints in a single forward pass. The second stage encodes each cropped pers...
421. Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents ​
Author: Xuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, Shumin Deng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG, cs.MA, cs.SE
arXiv:2608.31076v1 Announce Type: cross Abstract: Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the...
422. LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering ​
Author: Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen, Ahmed E. Hassan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.31102v1 Announce Type: cross Abstract: Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improvements under fixed compute and mixture budgets without regressing the rest. The maintained artifact is increasingly dataware: behavior g...
423. Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification ​
Author: Yisen Xi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR
arXiv:2608.31142v1 Announce Type: cross Abstract: The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No va...
424. SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies ​
Author: Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.31167v1 Announce Type: cross Abstract: Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task sem...
425. A Mental Model Based Framework of Trust ​
Author: Zahra Zahedi, Sarath Sreedharan, Erin Chiou, Subbarao Kambhampati
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2301.12569v2 Announce Type: replace Abstract: Handling trust is a core requirement of effective interaction between people and AI agents. Thus, any decision-making framework designed to work with people must be able to estimate human trust. In this paper, we propose a mental model-based framew...
426. AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance ​
Author: Dhaval Patel, Shuxin Lin, James Rayfield, Nianjun Zhou, Chathurangi Shyalika, Suryanarayana R Yarrabothula, Roman Vaculin, Natalia Martinez, Fearghal O'donncha, Jayant Kalagnanam
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2506.03828v4 Announce Type: replace Abstract: AI for Industrial Asset Lifecycle Management aims to automate complex operational workflows, such as condition monitoring and maintenance scheduling, to minimize system downtime. While traditional AI/ML approaches solve narrow tasks in isolation, L...
427. Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents ​
Author: Thassilo M. Schiepanski, Nicholas Pi"el
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2508.04412v4 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incrementally solve the give...
428. Language-Guided Tuning: Configuration Optimization for Automated ML Research ​
Author: Yuxing Lu, Yucheng Hu, Nan Sun, Xukai Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA
arXiv:2508.15757v2 Announce Type: replace Abstract: Configuration optimization remains a critical bottleneck in machine learning, requiring coordinated tuning across model architecture, training strategy, feature engineering, and hyperparameters. Traditional approaches treat these dimensions indepen...
429. SPADE: A Large Language Model Framework for Soil Moisture Pattern Recognition and Anomaly Detection in Precision Agriculture ​
Author: Yeonju Lee, Rui Qi Chen, Joseph Oboamah, Po Nien Su, Wei-zhen Liang, Yeyin Shi, Lu Gan, Yongsheng Chen, Xin Qiao, Jing Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2509.18123v2 Announce Type: replace Abstract: Accurate interpretation of soil moisture patterns is critical for irrigation scheduling and crop management, yet existing approaches for soil moisture time-series analysis either rely on threshold-based rules or data-hungry machine learning or deep...
430. Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion ​
Author: Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang, Bingsheng He, Gongshen Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2509.21054v3 Announce Type: replace Abstract: Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs). This paper studies persuasion dynamics by contrasting general LLMs with Large Reasoning Models (LRMs) that employ expl...
431. GUI-PRA: Process Reward Agent for GUI Tasks ​
Author: Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2509.23263v3 Announce Type: replace Abstract: Long-horizon GUI automation remains challenging due to error accumulation over extended interaction sequences. Process Reward Models (PRMs) provide dense step-level supervision for mitigating error accumulation, yet standard PRMs are poorly suited ...
432. Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models ​
Author: Qingjie Zhang, Yujia Fu, Yang Wang, Liu Yan, Tao Wei, Ke Xu, Minlie Huang, Han Qiu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2509.24711v4 Announce Type: replace Abstract: Current answering paradigms for Large Reasoning Models (LRMs) often fail to account for the fact that some questions may lie beyond the model's operational capability boundary, leading to long but unproductive reasoning. In this paper, we study whe...
433. MobileDreamer: Generative Sketch World Model for GUI Agent ​
Author: Yilin Cao, Yufeng Zhong, Zhixiong Zeng, Siran Dai, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Wenji Mao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.04035v2 Announce Type: replace Abstract: Mobile GUI agents have shown strong potential in real-world automation and practical applications. However, most existing agents remain reactive, making decisions mainly from current screen, which limits their performance on long-horizon tasks. Bui...
434. Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction ​
Author: Muzhao Tian, Zisu Huang, Xiaohua Wang, Jingwen Xu, Zhengkang Guo, Qi Qian, Yuanzhe Shen, Kaitao Song, Jiakang Yuan, Changze Lv, Xiaoqing Zheng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.05107v2 Announce Type: replace Abstract: As LLM-based agents are increasingly used in long-term interactions, cumulative memory is critical for enabling personalization and maintaining stylistic consistency. However, most existing systems adopt an ``all-or-nothing'' approach to memory usa...
435. SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation ​
Author: Seoyeon Kim, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2601.09974v3 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over time. However, real-world interactions are dynamic, where user interests continuously evolve, posing a ...
436. Real-Time Deadlines Reveal Fragile Temporal Adaptation in LLM Strategic Dialogues ​
Author: Neil K. R. Sehgal, Sharath Chandra Guntuku, Lyle Ungar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.13206v2 Announce Type: replace Abstract: Large Language Models (LLMs) generate text token-by-token in discrete time, yet real-world communication, from therapy sessions to business negotiations, critically depends on continuous time constraints. We use simulated negotiations between paire...
437. Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models ​
Author: Shahar Ben-Natan, Oren Tsur
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY
arXiv:2601.15436v3 Announce Type: replace Abstract: We propose a novel perspective for probing LLM sycophancy in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelty of our approach is...
438. ShardMemo: Scope-Before-Routing for Agentic Memory Retrieval ​
Author: Yang Zhao, Chengxiao Dai, Mengying Kou, Yue Xiu, Dusit Niyato
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2601.21545v2 Announce Type: replace Abstract: Agentic systems accumulate persistent memory across sessions, tools, and tasks, and a later request must retrieve from it under two distinct constraints: which memories it is permitted to access, and which are relevant under a limited search budget...
439. Defining Operational Conditions for Safety-Critical AI-Based Systems from Data ​
Author: Johann Maximilian Christensen, Elena Hoemann, Frank K"oster, Sven Hallerbach
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.22118v3 Announce Type: replace Abstract: Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems in the real world, defining the underlying environmental conditions in which the AI-based system must o...
440. HumanStudy-Bench: Towards AI Agent Design for Participant Simulation ​
Author: Xuan Liu, Haoyang Shang, Zizhang Liu, Xinyan Liu, Yunze Xiao, Yiwen Tu, Haojian Jin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.00685v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base model capabilities w...
441. Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning ​
Author: Yu Li, Mingyang Yi, Xiuyu Li, Ju Fan, Fuxin Jiang, Binbin Chen, Peng Li, Jie Song, Tieying Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.00994v3 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most existing ARL methods train a single set of parameters to support both reasoning and tool-use behavio...
442. Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Reasoning ​
Author: Yadong Wang, Haodong Chen, Yu Tian, Chuanxing Geng, Dong Liang, Xiang Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2602.01695v2 Announce Type: replace Abstract: Latent reasoning reduces the token-generation cost of chain-of-thought reasoning by replacing explicit intermediate tokens with continuous latent transitions. However, existing latent reasoning methods usually rely on dense and entangled transition...
443. AgentRx: Diagnosing AI Agent Failures from Execution Trajectories ​
Author: Shraddha Barke, Arnav Goyal, Alind Khare, Avaljot Singh, Suman Nath, Chetan Bansal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.02475v2 Announce Type: replace Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy tool outputs. We address this gap by manually annotating failed agent runs and release a novel benchm...
444. Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions ​
Author: Qianyun Guo, Yibo Li, Yue Liu, Bryan Hooi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.04191v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as personal assistants, where users may share individual preferences over extended interactions. However, assessing how well LLMs can follow these preferences in natural, long-term situations re...
445. Retrieval-Augmented LLM Agents: Learning to Learn from Experience ​
Author: Thomas Palmeira Ferraz, Romain Deffayet, Vassilina Nikoulina, Herv'e D'ejean, St'ephane Clinchant
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2603.18272v2 Announce Type: replace Abstract: While large language models (LLMs) have advanced the development of general-purpose agents, robust generalization to unseen tasks remains challenging. Two common approaches are supervised fine-tuning and training-free memory-augmented generation us...
446. Unified-MAS: Universally Generating Domain-Specific Nodes for Empowering Automatic Multi-Agent Systems ​
Author: Hehai Lin, Yu Yan, Zixuan Wang, Bo Xu, Sudong Wang, Weiquan Huang, Ruochen Zhao, Minzhi Li, Chengwei Qin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.21475v2 Announce Type: replace Abstract: Automatic Multi-Agent Systems (MAS) generation has emerged as a promising paradigm for solving complex reasoning tasks. However, existing frameworks are fundamentally bottlenecked when applied to knowledge-intensive domains (e.g., healthcare and la...
447. INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation ​
Author: Alexandra Kuleshova, Andrei Volodichev, Daria Kotova, Alexey Zaytsev
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.21607v3 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) enhances LLM performance, it does not eliminate hallucinations, making accurate detection essential. Uncertainty-based methods are attractive for this purpose because they can be integrated into real-world...
448. JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees ​
Author: Yuhui Wang, Zhixiong Yang, Ming Zhang, Shihan Dou, Zhiheng Xi, Enyu Zhou, Senjie Jin, Yujiong Shen, Dingwei Zhu, Yi Dong, Tao Gui, Qi Zhang, Xuanjing Huang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.22978v2 Announce Type: replace Abstract: In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malf...
449. PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded Verification ​
Author: Tianyu Shi, Wei Wang, Zequn Xie, Shuai Zhang, Boyang Xia, Chenyu Zeng, Qi Zhang, Lynn Ai, Yaqi Yu, Kaiming Zhang, Feiyue Tang, Zhenyu Yu, Lei Ding
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2603.27476v3 Announce Type: replace Abstract: AI-powered people search platforms are increasingly deployed for recruiting, sales prospecting, and professional networking, yet no standardized benchmark exists for their rigorous evaluation. We present PeopleSearchBench, an open-source benchmark ...
450. View-oriented Conversation Compiler for Agent Trace Analysis ​
Author: Lvmin Zhang, Maneesh Agrawala
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.29678v3 Announce Type: replace Abstract: We observe that an agent trace is a structured document. A coding agent session contains user turns, assistant text, chain of thought blocks, tool calls, tool results, subagent invocations, compaction boundaries, and harness injected directives, an...
451. A Multi-Agent Human-LLM Collaborative Framework for Closed-Loop Scientific Literature Summarization ​
Author: Maxwell J. Jacobson, Daniel Xie, Jackson Shen, Adil Wazeer, Guang Lin, Xiao-Ying Yu, Haiyan Wang, Xinghang Zhang, Yexiang Xue
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.01452v2 Announce Type: replace Abstract: Scientific discovery is slowed by fragmented literature that requires excessive human effort to gather, analyze, and understand. AI tools, including autonomous summarization and question answering, have been developed to aid in understanding scient...
452. ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback ​
Author: Cuong Van Duc, Minh Nguyen Dinh Tuan, Tam Vu Duc, Tung Vu Duy, Son Nguyen Van, Hanh Nguyen Thi, Binh Huynh Thi Thanh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.04940v3 Announce Type: replace Abstract: Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise. Recent LLM-guided evolutionary methods have shown promise for automated heuristic generation, but mo...
453. WebXSkill: Skill Learning for Autonomous Web Agents ​
Author: Zhaoyang Wang, Qianhui Wu, Xuchao Zhang, Chaoyun Zhang, Wenlin Yao, Fazle Elahi Faisal, Baolin Peng, Si Qin, Suman Nath, Qingwei Lin, Chetan Bansal, Dongmei Zhang, Saravan Rajmohan, Jianfeng Gao, Huaxiu Yao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2604.13318v2 Announce Type: replace Abstract: Autonomous web agents powered by large language models (LLMs) remain brittle on long-horizon browser workflows. A key bottleneck is a grounding gap in existing skill formulations: textual workflow skills provide natural language guidance but cannot...
454. ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams ​
Author: Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.15994v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However, when faced with complex topological structures involving branching paths, converging flows, and cyclic depe...
455. Training and Agentic Inference Strategies for LLM-based Manim Animation Generation ​
Author: Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.GR, cs.MA
arXiv:2604.18364v2 Announce Type: replace Abstract: Generating programmatic animation using libraries such as Manim presents unique challenges for Large Language Models (LLMs), requiring spatial reasoning, temporal sequencing, and familiarity with domain-specific APIs that are underrepresented in ge...
456. First-Order Efficiency for Probabilistic Value Estimation via A Statistical Viewpoint ​
Author: Ziqi Liu, Kiljae Lee, Yuan Zhang, Weijing Tang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, stat.ME, stat.ML
arXiv:2605.02827v2 Announce Type: replace Abstract: Probabilistic values, including Shapley values and semivalues, provide a model-agnostic framework to attribute the behavior of a black-box model to data points or features, with a wide range of applications including explainable artificial intellig...
457. SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents ​
Author: Ryangkyung Kang, Hongcheol Cho, Youngeun Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.05726v3 Announce Type: replace Abstract: As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical systems challenge. In small libraries, users may invoke skills explicitly by name, but this assumpti...
458. Useful Memories Become Faulty When Continuously Updated by LLMs ​
Author: Dylan Zhang, Yanshan Lin, Zhengkun Wu, Yihang Sun, Bingxuan Li, Dianqi Li, Hao Peng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.12978v2 Announce Type: replace Abstract: Learning from past experience benefits from two complementary forms of memory: episodic traces -- raw trajectories of what happened -- and consolidated abstractions distilled across many episodes into reusable, schema-like lessons. Recent agentic-m...
459. SciAtlas: A Computable Atlas of Science for Knowledge-Grounded AI Research ​
Author: Shuofei Qiao, Yunxiang Wei, Busheng Zhang, Mengru Wang, Jiazheng Fan, Huadong Jian, Bin Wu, Shumin Deng, Yida Xue, Zifan Cheng, Xiang Chen, Dan Zhang, Junfeng Fang, Ningyu Zhang, Keyan Ding, Qiang Zhang, Jeff Z. Pan, Emine Yilmaz, Huajun Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG
arXiv:2605.22878v2 Announce Type: replace Abstract: Artificial intelligence is rapidly entering the core workflows of scientific research. Yet reliable scientific reasoning requires access to accumulated scientific knowledge with sufficient breadth, depth, and standardization. Current AI scientists ...
460. Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models ​
Author: Weixian Waylon Li, Mengyu Wang, Tiejun Ma
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.LG
arXiv:2605.24564v2 Announce Type: replace Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An LLM trained in 2024 may already encode how stocks moved during 2018-2020. We name this failure paramet...
461. Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis ​
Author: Zhimin Lin, Kun Cheng, Zhiyao Shu, Junhua Fang, Juntao Li, Fan Bai, Jie Gao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.24600v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analysis. We argue this gap reflects a missing credibility practice from human QDA: peer debriefing, ...
462. VeriTrace: Evolving Mental Models for Deep Research Agents ​
Author: Haolang Zhao, Yunbo Long, Lukas Beckenbauer, Alexandra Brintrup
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.26081v2 Announce Type: replace Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representations should look like, but leave their evolution to the LLM's implicit reasoning. Without explicit...
463. PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management ​
Author: Yuxuan Zhao, Sijia Chen, Ningxin Su
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, q-fin.PM
arXiv:2605.27887v4 Announce Type: replace Abstract: Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlation...
464. AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios ​
Author: Kou Shi, Ziao Zhang, Shiting Huang, Avery Nie, Zhen Fang, Qiuchen Wang, Lin Chen, Huaian Chen, Zehui Chen, Feng Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.27995v3 Announce Type: replace Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations often overlook the temporal dimension of tool use, especially the impact of tool response latency, ...
465. Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training ​
Author: Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2605.28008v2 Announce Type: replace Abstract: Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge. To address this issue, supervised fine-tuning (SFT) often u...
466. Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild ​
Author: Rebecca M. M. Hicke, Kiran Tomlinson
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.29018v2 Announce Type: replace Abstract: Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze the conversationa...
467. Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering ​
Author: Yeong-Joon Ju, Seong-Whan Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.29742v2 Announce Type: replace Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority structures. Unlike traditional multi-hop or legal QA, this task requires structured procedural ...
468. VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing ​
Author: Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.30117v2 Announce Type: replace Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain...
469. When Should Models Change Their Minds? Contextual Belief Management in Large Language Models ​
Author: Haoming Xu, Weihong Xu, Zongrui Li, Mengru Wang, Yunzhi Yao, Chiyu Wu, Jin Shang, Yu Gong, Shumin Deng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2605.30219v2 Announce Type: replace Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this challenge as Contextual Belief Management (CBM): maintaining a predict...
470. Self-Correction Can Amplify Hallucinations: Fact-Level Repair with Graph-Based Evidence Routing in Multimodal Generation ​
Author: Kaixiang Zhao, Tianrun Yu, Shawn Huang, Porter Jenkins, Yushun Dong, Amanda Hughes
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.00232v2 Announce Type: replace Abstract: We study fact-level repair for multimodal generation, where a fluent output may contain specific facts that are not supported by the input. Existing inference-time repair methods often generate feedback by jointly conditioning on the input and the ...
471. Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs ​
Author: Yu-An Lu, Ci-Yang Tsai, Yu-Lin Tsai, Raluca Ada Popa, Chia-Mu Yu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2606.00642v2 Announce Type: replace Abstract: Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular, detailed traces can help distill reasoning behavior from stronger teacher models into weaker s...
472. Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks ​
Author: Dipesh KC, Anjila Budathoki
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.02875v2 Announce Type: replace Abstract: Coding-agent benchmarks evaluate whether a single uninterrupted agent can resolve a repository issue. Real software work is messier: tasks are interrupted, reassigned, reviewed, and resumed from partial states left by another agent or engineer. We ...
473. Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents ​
Author: Zhijie Ding (HyperAI Team, Xiaomi Corporation, Zhongnan University of Economics and Law), Weinan Hong (HyperAI Team, Xiaomi Corporation, Jilin University), Zicheng Zhu (HyperAI Team, Xiaomi Corporation, The Chinese University of Hong Kong, Shenzhen), Lei Li (HyperAI Team, Xiaomi Corporation), Dezhi Kong (HyperAI Team, Xiaomi Corporation), Hao Wang (HyperAI Team, Xiaomi Corporation), Peng Zhou (HyperAI Team, Xiaomi Corporation), Xuchu Jiang (HyperAI Team, Xiaomi Corporation), Jiaming Xu (HyperAI Team, Xiaomi Corporation)
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.03236v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents must decide when to intervene before determining how to assist. Existing systems often implement ...
474. Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference ​
Author: Wenbo Pan, Shujie Liu, Chin-Yew Lin, Jingying Zeng, Xianfeng Tang, Xiangyang Zhou, Yan Lu, Xiaohua Jia
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2606.05922v3 Announce Type: replace Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods typically require ground-truth validation sets,...
475. CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions ​
Author: Sherin Muckatira, Jesse Geneson, Slava Gerovitch, Pavel Etingof, Mikhail Gronas, Anna Rumshisky
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.06526v2 Announce Type: replace Abstract: Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step solutions, or complete proofs. They do not capture collaborative ope...
476. How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions ​
Author: Donghao Huang, Tomas Drietomsky, Benjamin Barrett, Zhaoxia Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.08051v3 Announce Type: replace Abstract: Merchant information extraction turns noisy financial transaction descriptors into structured fields at production scale. Our deployed LoRA-fine-tuned LLaMA~3.1-8B reaches 96.95% F1, but its memory and throughput motivate smaller replacements. We ...
477. ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents ​
Author: Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.08531v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the s...
478. From AGI to ASI ​
Author: Tim Genewein, Matija Franklin, Alexander Lerchner, Laurent Orseau, Samuel Albanie, Adam Bales, Cole Wyeth, Stephanie Chan, Iason Gabriel, Joel Z. Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, Shane Legg
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG
arXiv:2606.12683v2 Announce Type: replace Abstract: Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI organisations. Achieving this goal would have profound and far-reach...
479. IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing ​
Author: Tao Hu, Jiaxin Ai, Licheng Wen, Xueheng Li, Shu Zou, Siqi Li, Nianchen Deng, Xinyu Cai, Hongbin Zhou, Pinlong Cai, Daocheng Fu, Yu Yang, Hairong Zhang, Botian Shi, Xuemeng Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2606.13368v3 Announce Type: replace Abstract: Computer-Aided Design is pivotal in modern manufacturing, yet existing automated methods predominantly rely on open-loop, one-shot generation, creating a mismatch with iterative real-world practices. In this paper, we present IterCAD, a unified mul...
480. Beyond Helpfulness: A Teaching-over-Solving Diagnostic for Measuring Educational Impact in LLM Tutors ​
Author: Junyi Yao, Zihao Zheng, Baichuan Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.HC
arXiv:2606.16206v2 Announce Type: replace Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support. Motivated by recent calls to measure the social impact of NLP systems in practice, we stu...
481. A homotopy-type-theoretic generalization of neurosymbolic inference ​
Author: Fernando Zhapa-Camacho, Robert Hoehndorf
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2606.17851v2 Announce Type: replace Abstract: A wide range of neurosymbolic (NeSy) systems compute one functional: a belief-weighted sum of a logical quantity over a space of $\sigma$-structures, of which weighted model counting, fuzzy logic, and probabilistic logic are special cases. This acc...
482. DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue ​
Author: Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.17904v3 Announce Type: replace Abstract: Language models increasingly serve as advisory systems in maintenance operations. To prevent hallucination, established approaches ground these models in procedural documentation, constraining them to prescribed sequences. In practice, however, ope...
483. ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research ​
Author: Zhibang Yang, Xinke Jiang, Yuzhen Xiao, Ruizhe Zhang, Yue Fang, Xinfei Wan, Zhengxing Song, Yuxuan Liu, Yuheng Huang, Junfeng Zhao, Yasha Wang, Xu Chu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2606.20122v2 Announce Type: replace Abstract: Open-ended deep research (OEDR) requires systems to acquire knowledge through multi-round retrieval and generate coherent long-form reports. The outline plays a central role as a structural scaffold that coordinates retrieval, evidence organization...
484. PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate ​
Author: Yang Feng, Ziwei Xu, Xia Hu, Fengxiang He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, stat.ML
arXiv:2606.20621v2 Announce Type: replace Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques. However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and cause high sensitivity to role assi...
485. In LLM Reasoning, there is Irrationality on top of Value Misalignment ​
Author: Kejiang Qian, Fengxiang He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, stat.ML
arXiv:2606.20624v2 Announce Type: replace Abstract: Significant progress has been made in aligning LLMs with target value functions. We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning. We mathematically formalise thi...
486. DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models ​
Author: Jungseob Lee, Seongtae Hong, Seungjun Lee, Jaehyung Seo, Junyoung Son, Sugyeong Eo, Chanjun Park, Hyeongju Park, Hyeonseok Moon, Heuiseok Lim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.23181v3 Announce Type: replace Abstract: Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to fin...
487. RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems ​
Author: Yarin Yerushalmi Levi, Roy Betser, Amit Giloni, Lidor Erez, Itay Gershon, Oren Rachmil, Sindhu Padakandla, Roman Vainshtein
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.23927v2 Announce Type: replace Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specif...
488. COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models ​
Author: Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.28696v2 Announce Type: replace Abstract: Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into con...
489. Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows ​
Author: Edward Y. Chang, Longling Geng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.00269v3 Announce Type: replace Abstract: LLMs increasingly generate workflow actions and repairs that may be well formed yet stale, infeasible, conflicting, or destructive of their own evidence. We introduce Agentic Transaction Processing (ATP), which treats generated actions as untrusted...
490. Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve ​
Author: Dan C. Hsu, Luke Lu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.02975v2 Announce Type: replace Abstract: Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen custome...
491. Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation ​
Author: Kaiji Zhou, Ale\v{s} Leonardis, Yue Feng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.09600v2 Announce Type: replace Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs, based on coarse-grained matching between tasks and the ...
492. Can Agentic Trading Systems Pay for Their Own Intelligence? ​
Author: Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2607.10286v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, ...
493. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting ​
Author: Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.18084v2 Announce Type: replace Abstract: Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models ...
494. EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL ​
Author: Jiawei Zhou, Jianwei Wang, Chenyu Zhou, Chaojian Shi, Ming Dong, Kai Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2607.20489v2 Announce Type: replace Abstract: Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL,...
495. DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers ​
Author: Jerzy Kami'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.20531v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the under...
496. Masked Distillation: Internalizing the Chain-of-Thought in Language Models ​
Author: Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.22629v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctne...
497. Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe ​
Author: Chaemin Jang, Dongman Lee, Jihee Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.25292v2 Announce Type: replace Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from d...
498. Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation ​
Author: Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2607.25891v2 Announce Type: replace Abstract: Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making c...
499. MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning ​
Author: Kawai Chung, Chunkit Chan, Yauwai Yim, Yuxuan Liu, Haochen Shi, Weiqi Wang, Qing Zong, Tianshi Zheng, Yixuan Fu, Kai Chung Wong, Hao Liang, Yifan Gao, Xi Yang, Janet Hui-wen Hsiao, Yangqiu Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.26465v2 Announce Type: replace Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly ...
500. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning ​
Author: Wonpyo Park, Seung-won Hwang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03276v2 Announce Type: replace Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused acr...
501. Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection ​
Author: Razieh Chalehchaleh, Reza Farahbakhsh, Noel Crespi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03627v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gender bias in LLM-based fake n...
502. NxN E-valuation: Hypothesis Certification via a Conformal CRT Null ​
Author: Bin Wang, Yan Zhong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06621v2 Announce Type: replace Abstract: We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a la...
503. The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows ​
Author: Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06714v2 Announce Type: replace Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this searc...
504. PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents ​
Author: Mohammad Amanlou, Parham Abed Azad, Farbod Davoodi, Mostafa Masumi, Behnam Bahrak, Abdol-Hossein Vahabie
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2608.07438v2 Announce Type: replace Abstract: Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates fact...
505. Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification ​
Author: Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.09512v2 Announce Type: replace Abstract: Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by ...
506. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability ​
Author: Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11238v3 Announce Type: replace Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of...
507. Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents ​
Author: Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing Lu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.15071v2 Announce Type: replace Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in r...
508. Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World ​
Author: Jiang Jiang (Persagy Science and Technology Co., Beijing, China), Yifu Sun (Persagy Science and Technology Co., Beijing, China), Qi Shen (Persagy Science and Technology Co., Beijing, China)
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.15147v2 Announce Type: replace Abstract: Machine intelligence's push into the physical world is stuck on a gap: deployment demands auditable judgments from day one, fault samples are scarce or absent, and the norms defining "what counts as a fault" live in design documents, not in operati...
509. Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing ​
Author: Kang Chen, Sihan Zhao, Yixin Cao, Yu-Gang Jiang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.17638v2 Announce Type: replace Abstract: What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned...
510. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training ​
Author: Yugu Li, Zehong Cao, Jianglin Qiao, Siyi Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18682v3 Announce Type: replace Abstract: Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains high...
511. SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning ​
Author: Dayang Liang, Lang Feng, Bo An, Yunlong Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.19842v2 Announce Type: replace Abstract: Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead ...
512. Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis ​
Author: Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan, Qiang Hui, Shuting Liu, Fuyuan Shi, Ting Lu, Shaoan Zhao, Xueqiang Guo, Xinpei Su, Jianbing Zhang, Xinyu Dai, Kai Wang, Shiguo Lian
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20743v2 Announce Type: replace Abstract: Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself...
513. AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems ​
Author: Zhixu Du, Yiran Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.22160v2 Announce Type: replace Abstract: Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and warehouses at production rates beyond any human line, and their adoption is accelerating. But when their joint d...
514. ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation ​
Author: Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2608.22559v3 Announce Type: replace Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require LLM judges, and typically assume criteria aggregated through linea...
515. Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses ​
Author: Jiachen Xu, Torben Bach Pedersen, Zhongming Yao, Xiaoyu Zhang, Yushuai Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.22676v2 Announce Type: replace Abstract: Tool-using agents increasingly rely on external tools to complete multi-step tasks, but tool returns can fail in different ways and require different recovery actions. Existing robustness studies often use uncertainty-based measures to detect when ...
516. Characterizing Necessary Losers to Explain Tournaments Solutions ​
Author: Contet Cl'ement, Umberto Grandi, J'er^ome Mengin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.23446v2 Announce Type: replace Abstract: We study the problem of formally explaining why a candidate was not selected by a given tournament rule, by identifying sub-tournaments in which the candidate loses independently of how the rest of the tournament is completed. We define destructive...
517. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs ​
Author: Zhengxiang Wang, Owen Rambow
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.23978v2 Announce Type: replace Abstract: Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions are often incomplete or amb...
518. SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Planning ​
Author: Chirag Shah
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.24888v2 Announce Type: replace Abstract: Agents that act on a user's behalf must plan differently for different users, and increasingly do so from some structured representation of user context and not from raw interaction history. How much that structure is worth, and which parts of it c...
519. post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis ​
Author: Chandan Rajah
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.24921v2 Announce Type: replace Abstract: Graph RAG connects facts no single passage states, but implementations pay three times: in infrastructure, keeping vector store, graph database and document store in sync; in quality, because a pipeline that never refuses extractor output stores ed...
520. Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning ​
Author: Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang, Haipeng Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.25100v2 Announce Type: replace Abstract: Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. In partic...
521. Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems ​
Author: Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.25920v2 Announce Type: replace Abstract: As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair met...
522. Candidate supply and answer selection shape the value of LLM judging in multi-agent systems ​
Author: Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.25937v2 Announce Type: replace Abstract: Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneous...
523. ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs ​
Author: Songyuan Li, Ahmed M. Abdelmoniem, Shiqiang Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.25992v2 Announce Type: replace Abstract: Multi-agent large language model (LLM) workflows have emerged as a powerful paradigm for solving complex, open-ended tasks through collaborative reasoning among specialized LLM agents, but they incur substantial operating costs due to repeated LLM ...
524. LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems ​
Author: Jiaqi Xu, Yiran Qiao, Jing Chen, Qiwei Zhong, Xiang Ao, Xueqi Cheng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.MA
arXiv:2608.26849v2 Announce Type: replace Abstract: User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate i...
525. A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems ​
Author: Varvara Mama, Eleni Veroni, Nikolaos Kapsalis, Christos D. Nikolopoulos, Anargyros T. Baklezos
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27010v2 Announce Type: replace Abstract: In the present work an efficient border control management procedure is proposed. Compared to operational queue management systems, whose operations are based on mostly static data, the proposed work takes into account dynamic traffic conditions, t...
526. LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails ​
Author: Ziyang Chen, Xing Wu, Songlin Hu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27580v2 Announce Type: replace Abstract: Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanist...
527. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests ​
Author: Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.27831v2 Announce Type: replace Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To...
528. Rubric-to-Code Credit Assignment for Reinforcement Learning ​
Author: Rui Jin, Jikai Chen, Yihan Chen, Hao Zhou, Demin Zhu, Kaichen Yang, Dong Wang, Linjian Mo, Chenyi Zhuang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27906v2 Announce Type: replace Abstract: Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requ...
529. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents ​
Author: Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao, Han Li, Juntao Liu, Wenchao Chen, Chengduo Zhao, Chao Yu, Fandong Meng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28062v2 Announce Type: replace Abstract: Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequ...
530. Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation ​
Author: Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28393v2 Announce Type: replace Abstract: Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with...
531. Prove2Me: An Open Collaborative Platform for Scaling Math Formalization ​
Author: Shuze Chen, Kunal Marwaha, Xiaoyang Lu, Henry Yuen, Tianyi Peng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, cs.MA
arXiv:2608.28433v2 Announce Type: replace Abstract: Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mat...
532. Logos: An Agent Harness on a Cross-Process Bus ​
Author: Hanzhang Jia, Liheng Zeng, Hao Cheng, Yi Gao, Bo Ma
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.28553v2 Announce Type: replace Abstract: Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete formal treat ment in the spatiotemporal-composability calculus, in which a capability is a component carrying a tracked inverse, an...
533. Simulation-Based Evaluation of Energy-Constrained Quantum-Classical Competition ​
Author: Junyu Liu, Hansheng Jiang, Zuo-Jun Max Shen
Published: 9/1/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET, cs.LG, stat.ML
arXiv:2308.08025v2 Announce Type: replace-cross Abstract: This paper develops a simulation-based framework for evaluating the energy implications of quantum and classical computing firms competing in a market with limited energy resources. We model providers as differentiated Cournot competitors who...
534. General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level ​
Author: Bingkang Shi, Xiaodan Zhang, Dehan Kong, Yulei Wu, Zongzhen Liu, Honglei Lyu, Longtao Huang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2311.13892v4 Announce Type: replace-cross Abstract: The social biases and unwelcome stereotypes revealed by pretrained language models are becoming obstacles to their application. Compared to numerous debiasing methods targeting word level, there has been relatively less attention on biases pr...
535. SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems ​
Author: Oubo Ma, Yuwen Pu, Linkang Du, Yang Dai, Ruo Wang, Xiaolei Liu, Yingcai Wu, Shouling Ji
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2402.03741v4 Announce Type: replace-cross Abstract: Recent advancements in multi-agent reinforcement learning (MARL) have opened up vast application prospects, such as swarm control of drones, collaborative manipulation by robotic arms, and multi-target encirclement. However, potential securit...
536. PQMass: Probabilistic Assessment of the Quality of Generative Models using Probability Mass Estimation ​
Author: Pablo Lemos, Sammy Sharief, Esmeralda S. Whitammer, Salma Salhi, Connor Stone, Laurence Perreault-Levasseur, Yashar Hezaveh
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, stat.ME
arXiv:2402.04355v4 Announce Type: replace-cross Abstract: We propose a likelihood-free method for comparing two distributions given samples from each, with the goal of assessing the quality of generative models. The proposed approach, PQMass, provides a statistically rigorous method for assessing th...
537. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval ​
Author: Kyra Wilson, Aylin Caliskan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG
arXiv:2407.20371v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) hiring tools have revolutionized resume screening, and large language models (LLMs) have the potential to do the same. However, given the biases which are embedded within LLMs, it is unclear whether they can be us...
538. Understanding Deep Learning via Notions of Rank ​
Author: Noam Razin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE, stat.ML
arXiv:2408.02111v4 Announce Type: replace-cross Abstract: Despite the extreme popularity of deep learning in science and industry, its formal understanding is limited. This thesis puts forth notions of rank as key for developing a theory of deep learning, focusing on the fundamental aspects of gener...
539. Learning Personalized Prompts for Healthcare Guidance ​
Author: Ruize Shi, Hong Huang, Wei Zhou, Kehan Yin, Kai Zhao, Yun Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2412.15957v2 Announce Type: replace-cross Abstract: The rapid development of large language models (LLMs) has transformed many industries, including healthcare. In practice, hospitals and patients increasingly seek LLM-based systems capable of interpreting personal health records and providing...
540. Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias ​
Author: Enzo Doyen, Amalia Todirascu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2502.10577v2 Announce Type: replace-cross Abstract: Instruct-based large language models (LLMs) have been shown to propagate and even amplify gender bias when prompted with contextually constrained instructions (e.g., writing a text from a description or selecting a gendered pronoun). However,...
541. An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning ​
Author: Cen-Jhih Li, Aditya Bhaskara
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2502.11439v3 Announce Type: replace-cross Abstract: Fine-tuning is an important step in adapting foundation models such as large language models to downstream tasks. To make this step more accessible to users with limited computational budgets, it is crucial to develop fine-tuning methods that...
542. CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification ​
Author: Mingkun Zhang, Keping Bi, Wei Chen, Jiafeng Guo, Xueqi Cheng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2502.18176v3 Announce Type: replace-cross Abstract: In this paper, we aim to build an adversarially robust zero-shot image classifier. We ground our work on CLIP, a vision-language pre-trained encoder model that can perform zero-shot classification by matching an image with text prompts ``a ph...
543. RSPO: Regularized Self-Play Alignment of Large Language Models ​
Author: Xiaohang Tang, Sangwoong Yoon, Seongho Son, Huizhuo Yuan, Quanquan Gu, Ilija Bogunovic
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2503.00030v3 Announce Type: replace-cross Abstract: Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. However, the regularization with respect to the reference policy,...
544. Multimodal Large Language Models Predict Urban Safety Perception but Encode Non-Neutral Demographic Priors ​
Author: Ciro Beneduce, Bruno Lepri, Massimiliano Luca
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2503.00610v2 Announce Type: replace-cross Abstract: Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult to scale. We investigate whether Multimodal Large Language Models (MLLMs) can assess perceived urban s...
545. HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs ​
Author: Shujie Li, Yuxia Wu, Yuan Fang, Chuan Shi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2503.04822v2 Announce Type: replace-cross Abstract: Graph neural networks (GNNs) have demonstrated success in modeling relational data primarily under the assumption of homophily. However, many real-world graphs exhibit heterophily, where linked nodes belong to different categories or possess ...
546. A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives ​
Author: Weiqiang Jin, Hongyang Du, Shixiang Tang, Biao Zhao, Guang Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2503.13415v2 Announce Type: replace-cross Abstract: With the rapid development of artificial intelligence, intelligent decision-making techniques have gradually surpassed human levels in various human-machine competitions, especially in complex multi-agent cooperative task scenarios. Multi-age...
547. Towards Accurate and Lightweight Peripheral Neuroblastic Tumor Diagnosis via Contrastive Multi-scale Pathological Image Analysis ​
Author: Zhu Zhu, Shuo Jiang, Jingyuan Zheng, Yawen Li, Yifei Chen, Manli Zhao, Weizhong Gu, Feiwei Qin, Jinhu Wang, Gang Yu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2504.13754v4 Announce Type: replace-cross Abstract: Peripheral neuroblastic tumors (pNTs) are among the most common extracranial solid tumors in children, and accurate pathological subtyping is important for risk stratification and treatment planning. However, pNT subtyping on hematoxylin-eosi...
548. ReGraP-LLaVA: Reasoning enabled Graph-based Personalized Large Language and Vision Assistant ​
Author: Yifan Xiang, Zhenxi Zhang, Bin Li, Yixuan Weng, Bo Gao, Shoujun Zhou, Yangfan He, Yilin Yuan, Keqin Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2505.03654v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models have shown strong performance across multimodal tasks, and recent personalized MLLMs can recognize user-specific concepts and generate contextual captions. However, existing personalized MLLMs mainly focus on ...
549. Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning ​
Author: Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2505.16567v4 Announce Type: replace-cross Abstract: Finetuning open-weight Large Language Models (LLMs) is standard practice for achieving task-specific performance improvements. Until now, finetuning has been regarded as a controlled and secure process in which training on benign datasets lea...
550. Learning Composable Chains-of-Thought ​
Author: Fangcong Yin, Zeyu Leo Liu, Liu Leqi, Xi Ye, Greg Durrett
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2505.22635v2 Announce Type: replace-cross Abstract: A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoning problems, but such annotated data is costly to obtain for every problem of interest. We want reaso...
551. EquiReg: Equivariance Regularized Diffusion for Inverse Problems ​
Author: Bahareh Tolooshams, Aditi Chandrashekar, Rayhan Zirvi, Abbas Mammadov, Jiachen Yao, Chuwei Wang, Anima Anandkumar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2505.22973v3 Announce Type: replace-cross Abstract: Diffusion models represent the state-of-the-art for solving inverse problems such as image restoration tasks. Diffusion-based inverse solvers incorporate a likelihood term to guide prior sampling, generating data consistent with the posterior...
552. mRNA Design and Optimization with Deep Knowledge-Infused Approach ​
Author: Zheng Gong, Ziyi Jiang, Weihao Gao, Yuanyuan Wang, Zhining Cai, Deng Zhuo, Lan Ma
Published: 9/1/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.LG
arXiv:2505.23862v2 Announce Type: replace-cross Abstract: The mRNA optimization is essential for mRNA vaccines, therapies, and industrial protein production. Based on current explorations, an ideal optimization approach should simultaneously (i) prevent unintended amino-acid changes, (ii) optimize m...
553. Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study ​
Author: Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2506.08311v3 Announce Type: replace-cross Abstract: Automated Program Repair (APR) agents leverage large language models (LLMs) to autonomously diagnose and patch software bugs using planning, reasoning, and tools. Although these agents show strong performance on leaderboards such as SWE-bench...
554. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes ​
Author: Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2506.09557v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on ...
555. FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations ​
Author: Shaun Shuster, Eyal Zloof, Asaf Shabtai, Rami Puzis
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2506.11635v2 Announce Type: replace-cross Abstract: Credit card fraud mitigation plays a significant role in modern society. While fraud detection systems are essential, they often struggle to keep pace with the constantly evolving fraud techniques. As a result, fraud investigation is an impor...
556. Federated Learning for MRI-based BrainAGE: a multicenter study on post-stroke functional outcome prediction ​
Author: Vincent Roca, Marc Tommasi, Paul Andrey, Aur'elien Bellet, Markus D. Schirmer, Hilde Henon, Laurent Puy, Julien Ramon, Gr'egory Kuchcinski, Martin Bretzner, Renaud Lopes
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC
arXiv:2506.15626v3 Announce Type: replace-cross Abstract: $\textbf{Objective:}$ Brain-predicted age difference (BrainAGE) is a neuroimaging biomarker reflecting brain health. However, training robust BrainAGE models requires large datasets, often restricted by privacy concerns. This study evaluates ...
557. Training-free LLM Verification via Recycling Few-shot Examples ​
Author: Dongseok Lee, Jimyung Hong, Dongyoung Kim, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2506.17251v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges. Majority voting or Best-of-N with external verifiers h...
558. A foundation model with multi-variate parallel attention to generate neuronal activity ​
Author: Francesco Carzaniga, Michael Hersche, Abu Sebastian, Kaspar Schindler, Abbas Rahimi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2506.20354v3 Announce Type: replace-cross Abstract: Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, particularly in clinical domains such as intracranial electroencephalography (iEEG), where channel set...
559. Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice ​
Author: Gehao Zhang, Mingzhe Li, Eugene Bagdasarian, Shiqing Ma, Juan Zhai
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2507.05512v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for code generation, making reliable identification of machine-generated code important for attribution, tracking, and misuse detection. Existing code watermarking methods are dominated by N-...
560. Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems ​
Author: Yebo Peng, Yaoming Li, Zixiang Liu, Zhizhuo Yang, Xinye Xu, Bowen Ye, Weijun Yuan, Zihan Wang, Tong Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.02208v3 Announce Type: replace-cross Abstract: Evaluating the mathematical capability of Large Language Models (LLMs) is a critical yet challenging frontier. Existing benchmarks fall short, particularly for proof-centric problems, as manual creation is unscalable and costly, leaving the t...
561. StructSynth: Dependency Graphs as Generation Plans for Low-Data Tabular Synthesis with Language Models ​
Author: Siyi Liu, Yujia Zheng, Haoyang Li, Yongqi Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2508.02601v2 Announce Type: replace-cross Abstract: Tabular data derives its value from inter-feature dependencies, yet preserving them during synthesis is fragile when samples are scarce. Existing approaches either learn dependencies implicitly through distribution fitting, rely on statistica...
562. SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering ​
Author: Jan Melechovsky, Ambuj Mehrish, Abhinaba Roy, Dorien Herremans
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM, eess.AS
arXiv:2508.03448v4 Announce Type: replace-cross Abstract: Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment o...
563. From Isolation to Alignment: Unified LoRA for Efficient Multi-Task Learning ​
Author: Jinda Liu, Yi Chang, Yuan Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.05078v3 Announce Type: replace-cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is essential for adapting Large Language Models (LLMs) to multi-task scenarios. A prevailing trend in this field involves complex LoRA variants with multiple adapters or heads, which rely on the premise ...
564. MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents ​
Author: Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Hui Huang, Donghao Zhou, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Zhaoxiang Zhang, Qiangpeng Yang, Shilei Wen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2508.13186v2 Announce Type: replace-cross Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However, existing benchmarks such as BrowseComp primarily focus on textual content, overlooking the prevalen...
565. EEGDM: Learning EEG Representation with Latent Diffusion Model ​
Author: Shaocong Wang, Tong Liu, Yihan Li, Ming Li, Kairui Wen, Pei Yang, Wenqi Ji, Minjing Yu, Yong-Jin Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2508.20705v4 Announce Type: replace-cross Abstract: Recent advances in self-supervised learning for EEG representation have largely relied on masked reconstruction, where models are trained to recover randomly masked signal segments. While effective at modeling local dependencies, the training...
566. Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection ​
Author: Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, George Turkiyyah, Bernard Ghanem
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2508.20766v2 Announce Type: replace-cross Abstract: Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent research has demonstrated that these safety mechanisms can be bypassed by ablating or removing specific repr...
567. NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models ​
Author: Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhen Deng, Yingcai Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2509.03985v3 Announce Type: replace-cross Abstract: Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging. We present NeuroBreak, a visual an...
568. MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions ​
Author: Aishik Mandal, Tanmoy Chakraborty, Iryna Gurevych
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.04183v3 Announce Type: replace-cross Abstract: The growing demand for scalable psychological counseling highlights the need for high-quality, privacy-compliant data, yet such data remains scarce. Here we introduce MAGneT, a novel multi-agent framework for synthetic psychological counselin...
569. Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation ​
Author: Yiwen Guan, Jacob Whitehill
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.17930v3 Announce Type: replace-cross Abstract: Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In addition, translation quality can suffer for low-resource languages. To address this, we introduce Transfor...
570. Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection ​
Author: Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2509.24192v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes ...
571. Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act ​
Author: Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2509.25256v4 Announce Type: replace-cross Abstract: The systematic assessment of AI systems is increasingly vital as these technologies enter high-stakes domains. To address this, the EU's Artificial Intelligence Act introduces AI Regulatory Sandboxes (AIRS): supervised environments where AI s...
572. When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation ​
Author: Wenda Xu, Sweta Agrawal, Vil'em Zouhar, Markus Freitag, Daniel Deutsch
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.26600v3 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM as a benchmark) where a model generates test inputs (LLM as a testset) and evaluates outputs (LLM as an evaluator) has gained traction as a cheap alter...
573. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks ​
Author: Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, Varun Kumar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2510.01359v3 Announce Type: replace-cross Abstract: Code-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code, raising "jailbreak" stakes beyond text-only settings. Prior evaluations emphasize refusal or harmful-t...
574. Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem ​
Author: Ranran Haoran Zhang, Soumik Dey, Ashirbad Mishra, Hansi Wu, Binbin Li, Rui Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.22876v4 Announce Type: replace-cross Abstract: Inference optimizations are routinely evaluated by throughput alone, without verifying output correctness. We conduct a forensic analysis of batch speculative decoding and find that several widely-used implementations silently produce corrupt...
575. Personalized Treatment Outcome Prediction from Scarce Data via Dual-Channel Knowledge Distillation and Adaptive Fusion ​
Author: Wenjie Chen, Li Zhuang, Ziying Luo, Yu Liu, Jiahao Wu, Shengcai Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.26444v2 Announce Type: replace-cross Abstract: Personalized treatment outcome prediction based on trial data for small-sample and rare patient groups is a critical task in precision medicine. However, the high cost and scarcity of trial data limit the prediction performance. To address th...
576. LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature ​
Author: Magdalena Lederbauer, Siddharth Betala, Valerie Gentzke, Anamaria Leonescu, Amine Sehaba, Faris Flaifil, Ayush Jain, Alfonso Amayuelas, Nikhil Yelamarthy, Xiyao Li, Gr'egoire Germain, Stefano Ribes, Stefan P. Schmid, Alexandre Nozadze, Anna Kelmanson, Sudheesh Kumar Ethirajan, Mohd Zaki, Elton Pan, Georgia Channing, Connor W. Coley, Philippe Schwaller, Roc'io Mercado, Alexandre Duval, Mathilde L. D. Franckel, Samuel P. Gleason
Published: 9/1/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.IR
arXiv:2510.26824v2 Announce Type: replace-cross Abstract: Wide access to advanced experimental methods in materials science has given rise to an abundance of procedural knowledge, which is scattered across decades of scientific literature and recorded in unstructured formats that are challenging to ...
577. Retrofitters, pragmatists and activists: Public interest litigation for accountable automated decision-making ​
Author: Henry L Fraser, Zahra Stardust
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2511.03211v5 Announce Type: replace-cross Abstract: This paper examines the role of public interest litigation in promoting accountability for AI and automated decision-making (ADM) in Australia. Since ADM regulation faces political and geopolitical headwinds, effective governance will have to...
578. Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts ​
Author: Sebasti'an Andr'es Cajas Ord'o~nez, Luis Fernando Torres Torres, Mackenzie J. Meni, Carlos Andr'es Duran Paredes, Eric Arazo, Cristian Bosch, Ricardo Simon Carbajo, Yuan Lai, Leo Anthony Celi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.11743v4 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained devices faces two critical challenges: maintaining accuracy under aggressive quantization while ensuring predictable inference latency. We present a curiosity-driven quantized Mixture-of-...
579. Error-Driven Scene Editing for 3D Grounding in Large Language Models ​
Author: Yue Zhang, Zun Wang, Han Lin, Jialu Li, Jianing Yang, Yonatan Bitton, Idan Szpektor, Mohit Bansal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2511.14086v2 Announce Type: replace-cross Abstract: Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. This limitation stems in part from training data that focuses on language reasoning rather than spatia...
580. MedVision: Benchmarking Quantitative Medical Image Analysis ​
Author: Yongcheng Yao, Yongshuo Zong, Raman Dutt, Yongxin Yang, Sotirios A Tsaftaris, Timothy Hospedales
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.18676v3 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on quantitative...
581. Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs ​
Author: Sahil Kale
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.18931v2 Announce Type: replace-cross Abstract: Modern large language models increasingly integrate internal web-based retrieval to provide real-time answers, yet it remains unclear how effectively these systems identify information need, trigger retrieval, and use retrieved evidence. To u...
582. On the Optimality of Kinship Naming: an Information-theoretic Approach ​
Author: Phong Le, Mees Lindeman, Raquel G. Alhama
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.19120v2 Announce Type: replace-cross Abstract: The structure of naming systems in natural languages hinges on a trade-off between high informativeness and low complexity. Focusing on the domain of kinship naming, we analyze such trade-off while addressing simplifying assumptions of prior ...
583. WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving ​
Author: Seungjun Yu, Seonho Lee, Namho Kim, Jaeyo Shin, Junsung Park, Wonjeong Ryu, Raehyuk Jung, Hyunjung Shim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.20022v3 Announce Type: replace-cross Abstract: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their application to autonomous driving. However, high-level reasoning in safety-critical scenarios, where ...
584. CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation ​
Author: Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua, Jingrui He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2511.22707v2 Announce Type: replace-cross Abstract: In web environments, user preferences are often refined progressively as users move from browsing broad categories to exploring specific items. However, existing generative recommenders overlook this natural refinement process. Generative rec...
585. ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth ​
Author: Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik Roy, Nanyun Peng, Yang Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2512.03244v2 Announce Type: replace-cross Abstract: Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the ability to scale process-level supervision. We propose ScaleP...
586. Better World Models Can Lead to Better Post-Training Performance ​
Author: Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie, Andrew Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2512.03400v2 Announce Type: replace-cross Abstract: We study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers, using Rubik's Cubes as our training domain. We ask: (1) how does explicitly pretraining a world model affect a mode...
587. Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length ​
Author: Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, stat.AP, stat.ML
arXiv:2512.07019v4 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The Item Response Theory (IRT) model with Computerized Adaptive T...
588. SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification ​
Author: Hongbo Wang, AprilPyone MaungMaung, Isao Echizen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2512.15052v4 Announce Type: replace-cross Abstract: Disclaimer: Samples in this paper may be harmful and cause discomfort. Multimodal large language models (MLLMs) enable multimodal understanding but inherit toxic signals from weakly curated pretraining corpora, leading to explicitly toxic out...
589. Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference ​
Author: Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra, Ramachandran Ramjee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC
arXiv:2512.16391v2 Announce Type: replace-cross Abstract: Attention is the dominant source of latency during long-context LLM inference, an increasingly popular workload with reasoning models and RAG. We propose Kascade, a training-free sparse attention method that leverages known observations such ...
590. KV Admission: Learning What to Write for Efficient Long-Context LLM Inference ​
Author: Yen-Chieh Huang, Pi-Cheng Hsiu, Rui Fang, Ming-Syan Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2512.17452v4 Announce Type: replace-cross Abstract: Long-context LLM inference is bottlenecked by the quadratic attention complexity and linear Key-Value (KV) cache growth. Prior approaches mitigate this via post-hoc selection or eviction but overlook the root inefficiency: indiscriminate toke...
591. Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process ​
Author: Zhijun Chen, Zeyu Ji, Qianren Mao, Hao Wu, Jinhuan Song, Junhang Cheng, Bangjie Qin, Zhuoran Li, Jingzheng Li, Kai Sun, Zizhe Wang, Yikun Ban, Zhu Sun, Xiangyang Ji, Hailong Sun, Xiao Huang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2512.23213v4 Announce Type: replace-cross Abstract: We propose LLM-PeerReview, an unsupervised LLM Ensemble method that selects the most ideal response from multiple LLM-generated candidates for each query, harnessing the collective wisdom of multiple models with diverse strengths. LLM-PeerRev...
592. Entropy-Aware Token Rejection for Improving Speculative Decoding ​
Author: Tiancheng Su, Meicong Zhang, Guoxiu He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2512.23765v2 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens and a stronger target model to verify them. However, standard SD is mainly designed for acceleration, and its outp...
593. DIP: Dynamic In-Context Planner For Diffusion Language Models ​
Author: Yang Li, Han Meng, Chenan Wang, Zhenyu Bi, Xuan Wang, Haipeng Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.03199v2 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have shown strong potential for general natural language tasks with in-context examples. Existing In-Context Learning (ICL) approaches largely inherit the practice of autoregressive language models (ARLMs), in...
594. EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning ​
Author: Mingyang Wei, Dehai Min, Zewen Liu, Yuzhang Xie, Guanchen Wu, Ziyang Zhang, Carl Yang, Max S. Y. Lau, Qi He, Lu Cheng, Wei Jin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.03471v4 Announce Type: replace-cross Abstract: Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clini...
595. AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs ​
Author: Chengming Cui, Tianxin Wei, Ziyi Chen, Ruizhong Qiu, Zhichen Zeng, Zhining Liu, Xuying Ning, Duo Zhou, Jingrui He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.06022v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors. Inference-time ensembling provides a practical way to combine these capabilities without r...
596. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation ​
Author: Junseok Lee, Chang-Jae Chun
Published: 9/1/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2601.06199v5 Announce Type: replace-cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rate acoustic features directly into the LLM input space, making ...
597. Do Language Models Reason Across Languages? ​
Author: Yan Meng, Wafaa Mohammed, Christof Monz
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.06644v2 Announce Type: replace-cross Abstract: The real-world information sources are inherently multilingual, which naturally raises a question about whether language models can synthesize information across languages. In this paper, we introduce a simple two-hop question answering setti...
598. Kinship Data Benchmark for Multi-hop Reasoning ​
Author: Tianda Sun, Dimitar Kazakov
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.07794v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly evaluated on their ability to perform multi-hop reasoning, i.e., to combine multiple pieces of information into a coherent inference. We introduce KinshipQA, a benchmark designed to probe this cap...
599. Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models ​
Author: Zhenghao He, Guangzhi Xiong, Bohan Liu, Sanchit Sinha, Aidong Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.08058v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that triggers this behavior remains poorly understood. Leveraging the sparse features captured by Sparse Autoen...
600. To Retrieve or To Think? Cross-Boundary Context Evolution for Multi-hop Complex Reasoning ​
Author: Rubing Chen, Jian Wang, Wenjie Li, Xiao-Yong Wei, Qing Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.08747v3 Announce Type: replace-cross Abstract: Current context augmentation methods, such as retrieval-augmented generation, play a crucial role in bridging a model's internal knowledge boundary and external evidence for multi-hop reasoning. However, they often follow a rigid policy and t...
601. AgenTRIM: Tool Risk Mitigation for Agentic AI ​
Author: Roy Betser, Amit Giloni, Shamik Bose, Sindhu Padakandla, Chiara Picardi, Lidor Erez, Roman Vainshtein
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2601.12449v2 Announce Type: replace-cross Abstract: AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risks such as indirect prompt injection and tool misuse. We charac...
602. CORE-T: COherent REtrieval of Tables for Text-to-SQL ​
Author: Hassan Soliman, Vivek Gupta, Dan Roth, Iryna Gurevych
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2601.13111v3 Announce Type: replace-cross Abstract: Realistic text-to-SQL workflows often require joining multiple tables. As a result, accurately retrieving the relevant set of tables becomes a key bottleneck for end-to-end performance. We study an open-book setting where queries must be answ...
603. POCI-Diff: 3D-Layout Guided Diffusion for Controllable Synthetic Surveillance Data Generation ​
Author: Andrea Rigo, Luca Stornaiuolo, Weijie Wang, Mauro Martino, Bruno Lepri, Nicu Sebe
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2601.14056v2 Announce Type: replace-cross Abstract: Training robust visual surveillance models requires large-scale datasets with precise spatial annotations, yet collecting real surveillance data is costly, privacy-sensitive, and often legally constrained. Synthetic data generation offers a c...
604. Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models ​
Author: Injin Kong, Hyoungjoon Lee, Yohan Jo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2601.14758v5 Announce Type: replace-cross Abstract: Post-training pretrained autoregressive models (ARMs) into masked diffusion models (MDMs) provides an efficient route to diffusion language modeling, but it remains unclear whether the resulting models reuse inherited autoregressive computati...
605. Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation ​
Author: Xinyi Wang, Grazziela Figueredo, Ruizhe Li, Xin Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.16753v2 Announce Type: replace-cross Abstract: Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is crucial for monitoring disease progression and guiding clinical decisions. Many recent automated radi...
606. Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content ​
Author: Parth Bhalerao, Ruiwen Guan, Diola Dsouza, Oana Ignat
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.17173v3 Announce Type: replace-cross Abstract: Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that provide reflection and guidance. Existing QA benchmarks rarel...
607. The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models ​
Author: Nora Graichen, Iria de-Dios-Flores, Gemma Boleda
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.19926v3 Announce Type: replace-cross Abstract: We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on over 3,000 datapoints spanning a wide range of syntactic phenomena, languages, models, and methods. We...
608. Securing Time Integrity in Energy IoT Against Clock Drift and Y2K38 Failures ​
Author: Saeid Jamshidi, Foutse Khomh, Carol Fung, Omar Abdul Wahab, Rolando Herrero
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.23147v3 Announce Type: replace-cross Abstract: Time integrity across distributed Internet of Things (IoT) devices is fundamental to reliable sensing, control, and security in energy cyber-physical systems. However, operational energy IoT systems remain vulnerable to clock-drift escalation...
609. R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation ​
Author: Zhuohong Chen, Zhengxian Wu, Zirui Liao, Shenao Jiang, Hangrui Xu, Yang Chen, Chaokui Su, Xiaoyu Liu, Haoqian Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.00104v4 Announce Type: replace-cross Abstract: Vision-centric retrieval for VQA requires retrieving images to supply missing visual cues and integrating them into the reasoning process. However, selecting the right images and integrating them effectively into the model's reasoning remains...
610. Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters ​
Author: Lana Do, Gio Jung, Juvenal Francisco Barajas, Andrew Taylor Scott, Shasta Ihorn, Alexander Mario Blum, Vassilis Athitsos, Ilmi Yoon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2602.01390v3 Announce Type: replace-cross Abstract: Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production, quality...
611. Semi-supervised CAPP Transformer Learning via Pseudo-labeling ​
Author: Dennis Gross, Helge Spieker, Arnaud Gotlieb, Emmanuel Stathatos, Panorios Benardos, George-Christopher Vosniakos
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.01419v2 Announce Type: replace-cross Abstract: High-level Computer-Aided Process Planning (CAPP) generates manufacturing process plans from part specifications. It suffers from limited dataset availability in industry, reducing model generalization. We propose a semi-supervised learning a...
612. Zero-shot Generalizable Graph Anomaly Detection with Mixture of Riemannian Experts ​
Author: Xinyu Zhao, Qingyun Sun, Jiayi Luo, Xingcheng Fu, Jianxin Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.06859v3 Announce Type: replace-cross Abstract: Graph Anomaly Detection (GAD) aims to identify irregular patterns in graph data, and recent works have explored zero-shot generalist GAD to enable generalization to unseen graph datasets. However, existing zero-shot GAD methods largely ignore...
613. Robustness of Vision Language Models Against Split-Image Harmful Input Attacks ​
Author: Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.08136v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such attacks due to extensiv...
614. Artifact Reduction in Undersampled 3D Cone-Beam CTs using a Hybrid 2D-3D CNN Framework ​
Author: Johannes Thalhammer, Tina Dorosti, Sebastian Peterhansl, Daniela Pfeiffer, Franz Pfeiffer, Florian Schaff
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.08727v2 Announce Type: replace-cross Abstract: Undersampled CT volumes minimize acquisition time and radiation exposure but introduce artifacts degrading image quality and diagnostic utility. Reducing these artifacts is critical for high-quality imaging. We propose a computationally effic...
615. Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ? ​
Author: Xu Hu, Yifan Zhang, Songtao Wei, Chen Zhao, Qiannan Li, Bingzhe Li, Feng Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.11166v2 Announce Type: replace-cross Abstract: Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assumed to improve factual correctness. However, how the parameter-efficient fine-tuning methods affect hal...
616. MAS-on-the-Fly: In-Context Structural Adaptation of LLM-Based Multi-Agent Systems ​
Author: Guangyi Liu, Haojun Lin, Huan Zeng, Heng Wang, Quanming Yao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2602.13671v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, existing works often rely on manual designs or "one-size-fits-all" automation and lack adaptability after deployment. ...
617. Reverse N-Wise Output-Oriented Testing for AI/ML and Quantum Computing Systems ​
Author: Lamine Rihani
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.14275v2 Announce Type: replace-cross Abstract: Artificial intelligence/machine learning (AI/ML) systems and emerging quantum computing software present unprecedented testing challenges characterized by high-dimensional/continuous input spaces, probabilistic/non-deterministic output distri...
618. ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies ​
Author: Xingjian Wu, Xvyuan Liu, Junkai Lu, Siyuan Wang, Xiangfei Qiu, Yang Shu, Jilin Hu, Chenjuan Guo, Bin Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2602.14681v5 Announce Type: replace-cross Abstract: LLM-powered Multi-Agent Systems (MAS) have emerged as an effective approach towards collaborative intelligence, and have attracted wide research interests. Among them, ``self-evolving'' MAS, treated as a more flexible and powerful technical r...
619. DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation ​
Author: Ziyuan Liu, Shizhao Sun, Danqing Huang, Yingdong Shi, Meisheng Zhang, Ji Li, Jingsong Yu, Jiang Bian
Published: 9/1/2026, 4:00:00 AM
Categories: cs.GR, cs.AI, cs.CV, cs.LG, cs.MM
arXiv:2602.17690v3 Announce Type: replace-cross Abstract: Graphic design generation demands a delicate balance between high visual fidelity and fine-grained structural editability. However, existing approaches typically bifurcate into either non-editable raster image synthesis or abstract layout gen...
620. Social-JEPA: Emergent Geometric Isomorphism ​
Author: Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.02263v3 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such models from distinct viewpoints of the same environment without any parameter sharing or coordinatio...
621. Toward Generalizable Deep Learning Based Peatland Fire Detection via Walsh Hadamard Transform and Domain Adaptation ​
Author: Emadeldeen Hamdan, Ahmad Faiz Tharima, Mohd Zahirasri Mohd Tohir, Dayang Nur Sakinah Musa, Erdem Koyuncu, Adam J. Watts, Ahmet Enis Cetin
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.02465v2 Announce Type: replace-cross Abstract: Machine learning-based wildfire detection has advanced significantly using deep learning models trained on large wildfire image and video datasets. However, peatland fires exhibit distinct characteristics, including smoldering combustion, low...
622. Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion ​
Author: Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Zhaoran Wang, Han Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2603.03485v4 Announce Type: replace-cross Abstract: Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time. In ...
623. MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultation ​
Author: Guoyi Li, Shihao Xu, Jiatong Ma, Zhongjiang Yao, Yunyun Han, Jianhua Chen, Yafeng Deng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.03677v2 Announce Type: replace-cross Abstract: Psychiatric consultation requires agents to elicit discriminative evidence, map uncertain narratives to diagnostic criteria, and decide when evidence suffices. Existing dialogue and retrieval-augmented systems condition policies on raw histor...
624. Scale-Plan: Scalable Language-Enabled Task Planning for Heterogeneous Multi-Robot Teams ​
Author: Piyush Gupta, Sangjae Bae, Jiachen Li, David Isele
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.ET, cs.MA
arXiv:2603.08814v2 Announce Type: replace-cross Abstract: Long-horizon task planning for heterogeneous multi-robot systems is essential for deploying collaborative teams in real-world environments; yet, it remains challenging due to the large volume of perceptual information, much of which is irrele...
625. Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment ​
Author: Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2603.10009v2 Announce Type: replace-cross Abstract: Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), opt...
626. Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services ​
Author: Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali
Published: 9/1/2026, 4:00:00 AM
Categories: q-fin.CP, cs.AI, cs.CY
arXiv:2603.10807v2 Announce Type: replace-cross Abstract: Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally...
627. Efficient Cross-View Localization in 6G Space-Air-Ground Integrated Network ​
Author: Min Hao, Yanbing Xu, Maoqiang Wu, Jinglin Huang, Chen Shang, Jiacheng Wang, Jiawen Kang, Zhu Han, Wei Ni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2603.11398v3 Announce Type: replace-cross Abstract: Recently, visual localization has become an important supplement to improve localization reliability, and cross-view approaches can greatly enhance coverage and adaptability. Meanwhile, future 6G will enable a globally covered mobile communic...
628. Grammar of the Wave: Towards Explainable Multivariate Time Series Event Detection via Neuro-Symbolic VLM Agents ​
Author: Sky Chenwei Wan, Yifei Y. Wang, Tianjun Hou, Xiqing Chang, Aymeric Jan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2603.11479v4 Announce Type: replace-cross Abstract: Time Series Event Detection (TSED) aims to localize semantically meaningful events in time series data, with critical applications in high-stakes domains. Unlike statistical anomalies, events are often defined by natural-language descriptions...
629. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection ​
Author: Shuo Jiang, Gaojia Zhang, Min Tan, Yufei Yin, Gang Pan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.11521v2 Announce Type: replace-cross Abstract: Unsupervised Camouflaged Object Detection (UCOD) remains a challenging task due to the high intrinsic similarity between target objects and their surroundings, as well as the reliance on noisy pseudo-labels that hinder fine-grained texture le...
630. The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas ​
Author: Aleksander Jarz\k{e}bowicz, Adam Przyby{\l}ek, Jacinto Estima, Yen Ying Ng, Jakub Swacha, Beata Zielosko, Lech Madeyski, Noel Carroll, Kai-Kristian Kemell, Bartosz Marcinkowski, Alberto Rodrigues da Silva, Viktoria Stray, Netta Iivari, Anh Nguyen-Duc, Jorge Melegati, Boris Deliba\v{s}i'c, Emilio Insfran
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2603.11842v2 Announce Type: replace-cross Abstract: The post-ChatGPT surge has rapidly reframed IS research and practice. As organizations and society grapple with GenAI adoption, a body of secondary studies and research agendas has emerged to synthesize early evidence and chart directions for...
631. An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process ​
Author: Nicholas M. Synovic, Karolina Ryzka, Alessandra V. Vellucci Solari, Kenny Lyons, James C. Davis, George K. Thiruvathukal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2603.13584v3 Announce Type: replace-cross Abstract: Deep learning has achieved recognition for its impact within natural sciences, yet the prohibitive financial and technical cost of training models from scratch inhibit adoption. Following software engineering community guidance, natural scien...
632. PA3: Policy-Aware Agent Alignment through Chain-of-Thought ​
Author: Shubhashis Roy Dipta, Daniel Bis, Kun Zhou, Lichao Wang, Benjamin Z. Yao, Chenlei Guo, Ruhi Sarikaya
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2603.14602v3 Announce Type: replace-cross Abstract: Conversational assistants powered by large language models (LLMs) excel at tool-use tasks but struggle with adhering to complex, business-specific rules. While models can reason over business rules provided in context, including all policies ...
633. Interpretable Predictability-Based AI Text Detection: A Replication Study ​
Author: Adam Skurla, Dominik Macko, Jakub Simko
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2603.15034v2 Announce Type: replace-cross Abstract: This paper replicates and extends the system used in the AuTexTification shared task for authorship attribution of machine-generated texts. Exact replication was not possible because of differences in data splits, model availability, and impl...
634. Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm ​
Author: Anna Babarczy, Andras Lukacs, Peter Vedres, Zeteny Bujka
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.18007v2 Announce Type: replace-cross Abstract: The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to infer others' beliefs, intentions, and emotions from text. Given that LLMs are trained on language data ...
635. PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking ​
Author: Jiacheng Bao, Haoran Yang, Yucheng Xin, Junhong Liu, Yuecheng Xu, Han Liang, Pengfei Han, Xiaoguang Ma, Dong Wang, Bin Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2603.19305v3 Announce Type: replace-cross Abstract: Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings. Existing text-to-motion generation models are predominantly trained on captured human motion datasets, whose priors assume human biomechan...
636. Mixture-Greedy for Online Generative Model Selection: Is UCB Necessary in Diversity-Aware Multi-Armed Bandits? ​
Author: Bahar Dibaei Nia, Farzan Farnia
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2603.21716v2 Announce Type: replace-cross Abstract: Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models is costly. This problem can be viewed as a multi-armed bandit (MAB) task. Under diversity-aware eval...
637. T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search ​
Author: Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2603.22341v2 Announce Type: replace-cross Abstract: While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution, particularly in ra...
638. Variable-Length Audio Fingerprinting ​
Author: Hongjie Chen, Hanyu Meng, Huimin Zeng, Ryan A. Rossi, Lie Lu, Josh Kimball
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM
arXiv:2603.23947v2 Announce Type: replace-cross Abstract: Audio fingerprinting converts audio to much lower-dimensional representations, allowing distorted recordings to still be recognized as their originals through similar fingerprints. Existing deep learning approaches rigidly fingerprint fixed-l...
639. M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification ​
Author: Mengting He, Shihao Xia, Haomin Jia, Wenfei Wu, Linhai Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.PL, cs.AI
arXiv:2603.24595v2 Announce Type: replace-cross Abstract: Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is implicit and poorly specified. Models and kernels evolve independently and often make incompatible as...
640. Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning ​
Author: Guangfu Guo, Xiaoqian Lu, Yue Feng, Mingming Sun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.26737v2 Announce Type: replace-cross Abstract: Current multimodal LLMs encode images as static visual prefixes and rely on text-based reasoning, lacking goal-driven and adaptive visual access. Inspired by human visual perception-where attention is selectively and sequentially shifted from...
641. Can We Change the Stroke Size for Easier Diffusion? ​
Author: Yunwei Bai, Ying Kiat Tan, Yao Shu, Tsuhan Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.26783v3 Announce Type: replace-cross Abstract: Diffusion models can be challenged in the low signal-to-noise regime, where they have to make pixel-level predictions despite the presence of high noise. The geometric intuition is akin to using the finest stroke for oil painting throughout, ...
642. Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing ​
Author: Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.SY, eess.SY
arXiv:2603.28900v3 Announce Type: replace-cross Abstract: We address robust separation assurance for small Unmanned Aircraft Systems (sUAS) under GPS degradation and spoofing via Multi-Agent Reinforcement Learning (MARL). In cooperative surveillance, each aircraft (or agent) broadcasts its GPS-deriv...
643. Where Does Robustness Live? Neuron-Guided Adaptation for Retrieval-Augmented Language Models ​
Author: Jae O Lee, Jaemin Kim, Sumyeong Ahn, Seo Yeon Park
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.02194v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Language Models (RALMs) have shown strong potential in knowledge-intensive tasks, yet they remain vulnerable when retrieved contexts are noisy or irrelevant. Robustness against such contexts requires two distinct capabilit...
644. Interactive Clarification for Cloud Infrastructure-as-Code Synthesis ​
Author: Zhenning Yang, Kaden Gruizenga, Tongyuan Miao, Patrick Tser Jern Kon, Hui Guan, Andrew Barto, Ang Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2604.02382v2 Announce Type: replace-cross Abstract: The scale and complexity of modern cloud infrastructure have made "Infrastructure-as-Code" (IaC) essential for managing deployments through declarative configurations. While large language models (LLMs) are increasingly used to generate IaC c...
645. Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation ​
Author: Kening Zheng, Wei-Chieh Huang, Jiahao Huo, Zhonghao Li, Henry Peng Zou, Yibo Yan, Xin Zou, Jungang Li, Junzhuo Li, Hanrong Zhang, Xuming Hu, Philip S. Yu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.03592v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the internal mechanisms driving these gaps remain poorly understood. In this work, we conduct a systematic analysis of expert routing patterns in M...
646. Do We Still Need Humans in the Loop? Human vs. LLM Annotation in Active Learning for TikTok Hate Speech Detection ​
Author: Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein, Hinrich Sch"utze
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.13899v5 Announce Type: replace-cross Abstract: Annotating data remains a costly bottleneck for supervised NLP. Active learning (AL) reduces the number of human labels needed by selecting only the most informative instances, while instruction-tuned LLMs attack the same bottleneck from the ...
647. Perturbation Sensitivity of Maximum-Likelihood Pairwise Ranking in Computational Decision Systems ​
Author: Junyi Yao, Zihao Zheng, Jiayu Long
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT
arXiv:2604.17805v3 Announce Type: replace-cross Abstract: Maximum-likelihood pairwise ranking is a com- mon computational mechanism for prioritization, reputation estimation, and comparison-driven decision support. Despite its broad use, the perturbation sensitivity of this estimator under structure...
648. KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness ​
Author: Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo, Minho Jang, Keunwoo Choi, Seungyoun Shin, Ji Won Yoon
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD, eess.AS
arXiv:2604.19782v2 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this pa...
649. Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models ​
Author: Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis, Seshu Tirupathi, John D. Kelleher
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2604.20994v2 Announce Type: replace-cross Abstract: The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions. Injection and jailbreaking attacks h...
650. SketchVLM: Vision language models can annotate images to explain thoughts and guide users ​
Author: Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.22875v3 Announce Type: replace-cross Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which can be difficult for use...
651. When Chain-of-Thought Fails, the Solution Hides in the Hidden States ​
Author: Houman Mehrafarin, Amit Parekh, Ioannis Konstas
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.23351v2 Announce Type: replace-cross Abstract: Whether intermediate reasoning is computationally useful or merely explanatory depends on whether chain-of-thought (CoT) tokens contain task-relevant information. We present a mechanistic causal analysis of CoT on GSM8K using activation patch...
652. Inverting Foundation Models of Brain Function with Simulation-Based Inference ​
Author: Niels Leif Bracher, Xavier Intes, Stefan T. Radev
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2604.23865v3 Announce Type: replace-cross Abstract: Foundation models of brain activity promise a new frontier for in silico neuroscience by emulating neural responses to complex stimuli across tasks and modalities. A natural next step is to ask whether these models can also be used in reverse...
653. CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies ​
Author: Fan Du, Feng Yan, Jianxiong Wu, Xinrun Xu, Weiye Zhang, Weinong Wang, Yu Guo, Bin Qian, Zhihai He, Fei Wang, Heng Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.24622v3 Announce Type: replace-cross Abstract: Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to recover action structure from uninformative Gaussian noise, lead...
654. Star-Fusion: A Multi-modal Transformer Architecture for Discrete Celestial Orientation via Spherical Topology ​
Author: May Hammad, Menatallh Hammad
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.26582v2 Announce Type: replace-cross Abstract: Reliable celestial attitude determination is a critical requirement for autonomous spacecraft navigation, yet traditional "Lost-in-Space" (LIS) algorithms often suffer from high computational overhead and sensitivity to sensor-induced noise. ...
655. Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors ​
Author: Zi Li, Tian Zhou, Wenze Li, Jingyu Hua, Yunlong Mao, Sheng Zhong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2604.27426v2 Announce Type: replace-cross Abstract: Local fine-tuning datasets routinely contain sensitive secrets such as API keys, personal identifiers, and financial records. Although "local offline fine-tuning" is often viewed as a privacy boundary, we reveal that compromised model code is...
656. AirFM-DDA: Air-Interface Foundation Model in the Delay-Doppler-Angle Domain for AI-Native 6G ​
Author: Kejia Bian, Meixia Tao, Jianhua Mo, Zhiyong Chen, Leyan Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, eess.SP, math.IT
arXiv:2605.00020v2 Announce Type: replace-cross Abstract: The success of large foundation models is catalyzing a new paradigm for AI-native 6G network design: wireless foundation models for physical-layer design. However, existing models often operate on channel state information (CSI) in the spatia...
657. Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations ​
Author: Pratyush Acharya, Nuraj Rimal, Habish Dhakal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.01609v2 Announce Type: replace-cross Abstract: We find that transformer concept representations systematically anti-concentrate in the spectral tail of the unembedding covariance, encoding word-level concepts in low-variance directions across a 17-model core suite and an expanded set of 2...
658. Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level ​
Author: Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, Zequn Sun
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.06387v4 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its standard advantage weighted p...
659. DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain ​
Author: Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.07699v3 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations. Despite the prevalence o...
660. Robust Multi-Agent LLMs under Byzantine Faults ​
Author: Haejoon Lee, Vincent-Daniel Yun, Dimitra Panagou, Sai Praneeth Karimireddy
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG
arXiv:2605.09076v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same interactions can also introduce vulnerability to unreliable or Byzantine agents that can propagate incorre...
661. Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning ​
Author: Swapnil Parekh, Naman Goyal
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.11467v2 Announce Type: replace-cross Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of reasoning theater: deliberative-looking steps that contribute nothing to correctness. This wastes inference tokens, pollutes interp...
662. Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety ​
Author: Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CR
arXiv:2605.12729v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly being used in network operations (NetOps) and artificial intelligence for IT operations (AIOps) for tasks ranging from telemetry retrieval and incident diagnosis to configuration planning and boun...
663. IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation ​
Author: Shijie Lian, Bin Yu, Xiaopeng Lin, Zhaolong Shen, Laurence Tianruo Yang, Yurun Jin, Haishan Liu, Changti Wu, Hang Yuan, Cong Huang, Kai Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.CV
arXiv:2605.14712v3 Announce Type: replace-cross Abstract: Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human demonstrators act with different short-horizon intents, task phases, or recent context. Existing frame-co...
664. Prefix-Adaptive Block Diffusion for Efficient Document Recognition ​
Author: Mingxu Chai, Ziyu Shen, Chenyu Liu, Jihua Kang, Tao Gui, Qi Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.16861v2 Announce Type: replace-cross Abstract: Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind denoising and cache commitment to fixed block boundaries: par...
665. Adversarial Trust Poisoning in Vehicular Collaborative Perception ​
Author: Yutong Liu, Chenyi Wang, Ming F. Li, Qingzhao Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2605.22122v2 Announce Type: replace-cross Abstract: Collaborative perception (CP) enables connected and autonomous vehicles to share sensor data and jointly reason about their environment. To defend against adversaries that fabricate or manipulate shared data, existing systems employ cross-veh...
666. Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light ​
Author: Valeria Pais, Malena Mendilaharzu, Daniele Faccio, Luis Oala, Christoph Clausen, Bruno Sanguinetti
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, physics.optics
arXiv:2605.22455v2 Announce Type: replace-cross Abstract: Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven: long-tailed or unbalanced distributions hinder generalization, and the low number of sam...
667. Measuring the Depth of LLM Unlearning via Activation Patching ​
Author: Jaeung Lee, Dohyun Kim, Jaemin Jo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.24614v2 Announce Type: replace-cross Abstract: Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target knowledge is truly erased remains challenging. Existing output-level metrics fail to detect wh...
668. When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection ​
Author: Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang, Yuetai Li, Zhihan Xiong, Yue Liu, Junhao Lin, Yao Su, Lijie Hu, Kaize Ding, Teng Xiao, Radha Poovendran
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2605.26872v5 Announce Type: replace-cross Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate student training data, impl...
669. QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents ​
Author: Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2605.27068v2 Announce Type: replace-cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and l...
670. BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving ​
Author: Tianyao Shi, Yi Ding
Published: 9/1/2026, 4:00:00 AM
Categories: q-bio.OT, cs.AI, cs.CY
arXiv:2605.27480v3 Announce Type: replace-cross Abstract: Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related pathways. We present BIRDS, a framework for Biodiversity Impact of Request-Driven LLM Serving. B...
671. Locality-Aware Redundancy Pruning for LLM Depth Compression ​
Author: Vincent-Daniel Yun, Youngrae Kim, Woosang Lim, YoungJin Heo, Minkyu Kim, Sunwoo Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.27786v3 Announce Type: replace-cross Abstract: Large language models are known to contain representational redundancy across network depth, making depth pruning an effective approach for improving inference efficiency. Existing one-shot pruning methods rely on local layer importance or fi...
672. Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses ​
Author: Kerui Peng, Feifei Li, Xingyu Fan, Wenhui Que
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.27971v2 Announce Type: replace-cross Abstract: When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a failure we term Cross-Style Collapse. We trace this collapse to the cross-entropy objective, which und...
673. KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs ​
Author: Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.27984v2 Announce Type: replace-cross Abstract: Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily centered on English, limiting reliable assessment of multilin...
674. Integrated and Cross-Architecture Interpretation of LLM Reasoning ​
Author: Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.28006v2 Announce Type: replace-cross Abstract: Understanding how LLMs reason is hindered by a practical asymmetry: while their generated outputs are observable, the underlying reasoning patterns remain opaque. Relying on single probes, such as Mutual Information Peak (MIP) or Deep-Thinkin...
675. Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts ​
Author: Liu O. Martin, Lucas Bandarkar, Nanyun Peng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.28042v2 Announce Type: replace-cross Abstract: Modern large language models (LLMs) achieve state-of-the-art machine translation performance, but they do so as broad generalists largely trained for many tasks and capabilities unrelated to translation. Thus, they are heavily overparameteriz...
676. BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law ​
Author: Sebastian Nagl, Ann-Kristin Mayrhofer, Martin Heidebach, Aleyna Ko\c{c}ak, Anne Zettelmeier, Elly Breu, Angelina Greiner, Sofija Milijas, Matthias Grabmair
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.28183v4 Announce Type: replace-cross Abstract: We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The dataset combines 596 exam-style free-text legal case tasks across multiple levels of le...
677. Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text ​
Author: Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.28740v2 Announce Type: replace-cross Abstract: As large language models are increasingly deployed for clinical text, ensuring they can reliably signal their own uncertainty becomes critical. Most existing uncertainty quantification (UQ) methods are designed for open-domain generation and ...
678. Skill-Conditioned Gated Self-Distillation for LLM Reasoning ​
Author: Jiazhen Huang, Xiao Chen, Xiao Luo, Yong Dai, Senkang Hu, Yuzhi Zhao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.28791v3 Announce Type: replace-cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision. Existing methods usually assume trusted PI, such as reference answe...
679. OISD: On-Policy Internal Self-Distillation of Language Models ​
Author: Xinyu Liu, Darryl Cherian Jacob, Yang Zhou, Jindong Wang, Pan He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2605.29089v2 Announce Type: replace-cross Abstract: Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while largely overlooking predictive signals encoded in intermediate representations. In this paper, we...
680. Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents ​
Author: Aditya Nawal, Manit Baser, Mohan Gurusamy
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR
arXiv:2605.29224v2 Announce Type: replace-cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms tha...
681. How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions ​
Author: Ningzhi Tang, Chaoran Chen, Gelei Xu, Yiyu Shi, Yu Huang, Collin McMillan, Tao Dong, Toby Jia-Jun Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC
arXiv:2605.29442v2 Announce Type: replace-cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that miss how developers actually experience misalignment. We present an observational study of 20...
682. LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training ​
Author: Minju Gwak, Minseo Kwak, Dongseok Lee, Guijin Son, Alan Ritter, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.29888v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has shown to improve reasoning in large language models (LLMs). However, there has been little exploration on the problem of data contamination in RL post-training, potentially undermining generalizat...
683. Exploring Autonomous Agentic Data Engineering for Model Specialization ​
Author: Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG
arXiv:2605.30407v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data. Existing LLM-based data curation methods primarily rely on h...
684. FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection ​
Author: Paramananda Bhaskar, Naquee Rizwan, Daksh Jogchand, Saurabh Kumar Pandey, Animesh Mukherjee
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.MM
arXiv:2605.31349v2 Announce Type: replace-cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding rhetorical hate mechanisms with target community features and preventing causal evaluation o...
685. Vision-Language Models Suppress Female Representations Under Ambiguous Input ​
Author: Arnau Marin-Llobet, Simon Henniger, Mahzarin R. Banaji
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.CY, cs.HC
arXiv:2605.31556v2 Announce Type: replace-cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in full gear, a figure seen from behind), case...
686. Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning ​
Author: Chishui Chen, Jiaye Lin, Te Sun, Yi Yang, Junxi Wang, Cong Qin, Yangen Hu, Lu Pan, Ke Zeng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.00510v2 Announce Type: replace-cross Abstract: Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks. However, existing methods mainly focus on selecting relevant skills or improving the skills themselves, while overl...
687. Linguistics-Aware Non-Distortionary LLM Watermarking ​
Author: Shinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub Han
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.00613v2 Announce Type: replace-cross Abstract: Watermarking should identify language-model output without degrading quality or limiting verification to the model provider. Multilingual deployment makes this harder because morphology, segmentation, and script change where watermark evidenc...
688. DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models ​
Author: Abdullah Al Shafi, Kazi Saeed Alam, Sk Imran Hossain, Engelbert Mephu Nguifo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2606.00798v2 Announce Type: replace-cross Abstract: Parameter compression of class-conditional diffusion models exposes a structural limitation in output-level distillation: supervising only the guided output leaves the two score branches non-identifiable, so the classifier-free guidance gap i...
689. The Value of Spike Timing: A Leakage-Resistant Benchmark of SNN Design Choices for Network Intrusion Detection ​
Author: Raj Patel, Shaswata Mitra, David Amebley, Taye Akinrele, Sayanton Dibbo, Shahram Rahimi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.NE
arXiv:2606.01442v2 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs) are increasingly studied for network intrusion detection, but comparative evidence on how neuron models and spike encodings affect performance remains limited. Evaluation choices can influence results when prepr...
690. FSA-GRPO: Teaching Auditory LLMs to Use Few-Shot Demonstrations ​
Author: Haolong Zheng, Siyin Wang, Xulin Fan, Zengrui Jin, Mark Hasegawa-Johnson
Published: 9/1/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2606.02615v2 Announce Type: replace-cross Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large language models are not explicitly trained to perform inference in ...
691. GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning ​
Author: Liyan Tan, Yequan Zhao, Yifan Yang, Ruijie Zhang, Xinling Yu, Zheng Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.02857v2 Announce Type: replace-cross Abstract: Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limited by the high variance of gradient estimation. We propose GRZO, a Group-Relative Zeroth-Ord...
692. OpenAgenet / OAN White Paper: Open Infrastructure for Trusted Agent Interconnection ​
Author: Jinliang Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2606.03161v4 Announce Type: replace-cross Abstract: OpenAgenet, abbreviated as OAN, is an open infrastructure project for trusted Agent interconnection. It addresses a problem that becomes visible when Agents move from isolated applications into open, multi-operator networks: before an Agent c...
693. OpenAgenet / OAN Yellow Paper: Technical Architecture for Trust-Governed Resource Identity and Discovery ​
Author: Jinliang Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.DC
arXiv:2606.03163v4 Announce Type: replace-cross Abstract: This yellow paper describes the technical architecture of OpenAgenet / OAN. OAN is a protocol-neutral trust layer for open Agent interconnection and discoverable AI resource products. It specifies the role architecture, \texttt{did:oan} ident...
694. The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics ​
Author: Nils Schwager, Christoph Hau, Simon M"unker, Achim Rettinger
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.03357v2 Announce Type: replace-cross Abstract: When prompting language models for psychometric assessment, researchers assume that the responses reflect the injected persona and the meaning of the survey item. We test this premise using a diagnostic design that crosses five semantically d...
695. Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning ​
Author: Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.03965v2 Announce Type: replace-cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by sho...
696. QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards ​
Author: Rongzhi Zhang, Rui Feng, Zhihan Zhang, Jingfeng Yang, Qingyu Yin, Xin Liu, Zixuan Zhang, Priyanka Nigam, Bing Yin, Tuo Zhao, Chao Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.03968v2 Announce Type: replace-cross Abstract: Rubric-based RL is a promising route for extending reinforcement learning beyond verifiable rewards, yet existing methods optimize rubrics while treating the query distribution as fixed. We identify a structural bottleneck: rubric quality is ...
697. POLARIS: Guiding Small Models to Write Long Stories ​
Author: Rishanth Rajendhran, Jenna Russell, Mohit Iyyer, John Frederick Wieting
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.04095v2 Announce Type: replace-cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models. ...
698. ITP-STDP: A Hardware-Efficient Intrinsic-Timing Power-of-Two Synaptic Learning Engine for On-Chip SNNs ​
Author: Haihang Xia, Xinyu Zhao, Xuecheng Wang, John Goodenough, Charith Abhayaratne, Panagiotis A. Panagiotou, Chunyi Song, Tiantai Deng
Published: 9/1/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.NE
arXiv:2606.06159v2 Announce Type: replace-cross Abstract: Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across a wide range of applications. However, the large number of synaptic connections in SNNs lead...
699. Twelve quick tips for designing AI-driven HPC workflows ​
Author: Jamie J. Alnasir
Published: 9/1/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.SE
arXiv:2606.07491v2 Announce Type: replace-cross Abstract: High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pipelines optimised for predictable performance. However, the pervasive integration of artifici...
700. ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding ​
Author: Zirui Wang, Yusen Hou, Shaofeng Liang, Bowen Tian, Yanlin Zhang, Wenshuo Chen, Yutao Yue
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.07524v2 Announce Type: replace-cross Abstract: The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison increasingly important for provenance auditing, security analysis, and model selection. Exist...
701. Summarization is Not Dead Yet ​
Author: Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.08000v3 Announce Type: replace-cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem. We re-examine this n...
702. Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark ​
Author: Shundong Li
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.10084v2 Announce Type: replace-cross Abstract: This submission documents the divide-and-conquer modeling strategy developed for the CTF-4-Science Lorenz Chaotic Systems Challenge at AI-DEEDS 2026. The challenge uses the CTF-4-Science Lorenz benchmark to evaluate chaotic-system prediction ...
703. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models ​
Author: Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.10829v3 Announce Type: replace-cross Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictio...
704. BioDivergence: A Benchmark and Evaluation Framework for Hidden Contextual Contradictions in Biomedical Abstracts ​
Author: Elias Hossain, Sanjeda Sara Jennifer, Sabera Akter Bushra, Niloofar Yousefi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.11208v2 Announce Type: replace-cross Abstract: Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions. Variations in cohort, geography, assay protocol, disease subtype, and clinical setting can make bo...
705. A Zero-shot Generalized Graph Anomaly Detection Framework via Node Reconstruction ​
Author: Phan Nguyen, Dat Cao, Hien Chu, Khue Hoang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.12673v2 Announce Type: replace-cross Abstract: Cross-domain graph anomaly detection (GAD) aims to identify abnormal nodes in unseen target graphs, showing strong potential in real-world applications with heterogeneous graph data. However, existing methods often depend on dataset-specific ...
706. TRACE: Trajectory-Routed Causal Memory for Delayed-Evidence Visuomotor Imitation ​
Author: Zihao Li, Ranpeng Qiu, Yincong Chen, Guoqiang Ren, Weiming Zhi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2606.14551v3 Announce Type: replace-cross Abstract: Robots under autonomous operation may require decisions based on evidence that is no longer visible. We study delayed-evidence tasks, where an early cue disappears before a later decision point, so visually similar observations can require di...
707. CARE: Context-Aware Ranking Evolution with Executable Scoring Programs for Budgeted Reaction Optimization ​
Author: Guanyu Liu, Weiyi Kong, Chao Tang, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, Tianyu Shi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.14581v5 Announce Type: replace-cross Abstract: High-throughput experimentation can evaluate many reaction conditions, yet combinatorial condition spaces still exceed the available experiment budget. This makes experiment selection a sequential decision problem: each new condition must be ...
708. MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios ​
Author: Haitian Wang, Ruoxi Sun, Quantong Qiu, Juntao Li, Junhui Li, Hua Chen, Jinxiong Chang, Min Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.14747v2 Announce Type: replace-cross Abstract: Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs). However, larger context windows do not necessarily translate into effective comprehension and representation of long-contex...
709. Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot ​
Author: Yuyang Dai, Yushun Dong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2606.15810v3 Announce Type: replace-cross Abstract: Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a defense that redirects ex...
710. MagicSim: A Unified Infrastructure for Executable Embodied Interaction ​
Author: Haoran Lu, Songling Liu, Yue Chen, Guo Ye, Mutian Shen, Shuyang Yu, Yu Xiao, Shang Wu, Jiayi Wang, Jianshu Zhang, Jihai Zhao, Xiangtian Gui, Chuye Hong, Yuran Wang, Maojiang Su, Ruihai Wu, Zhaoran Wang, Han Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2606.17511v2 Announce Type: replace-cross Abstract: Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, not only as a renderer, controller testbed, or fixed task environment. Existing pipelines split these la...
711. APT: Atomic Physical Transitions for Causal Video-Language Understanding ​
Author: Shang Wu, Haoran Lu, Songling Liu, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Zhaoran Wang, Han Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.18586v2 Announce Type: replace-cross Abstract: Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bounce" can be correct while hiding the process that makes the event physically valid, from support loss a...
712. Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance ​
Author: Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.18613v4 Announce Type: replace-cross Abstract: The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistanc...
713. Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents ​
Author: Anoushka Vyas, Aarushi Dhanuka, Sina Khoshfetrat Pakazad, Henrik Ohlsson
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.DB
arXiv:2606.19319v2 Announce Type: replace-cross Abstract: Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present Data Intelligence Agents (DIA), a system ...
714. Skills for the future software profession: beyond agentic AI! ​
Author: Abhik Roychoudhury, Sungmin Kang, Baishakhi Ray
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2606.21894v3 Announce Type: replace-cross Abstract: As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engineers? To identify where software engineering is headed and thus what skills will be needed, we summariz...
715. DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection ​
Author: Yujin Tang, Chenming Shang, Ruize Xu, Nikhil Singh
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2606.27499v2 Announce Type: replace-cross Abstract: Agent benchmarks for measuring memory largely study textual cases, in which information is deliberately extracted from the environment, written down, and then later retrieved. In other words, they assess what agents elected to record, not wha...
716. MultiHashFormer: Hash-based Generative Language Models ​
Author: Huiyin Xue, Atsuki Yamaguchi, Nikolaos Aletras
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2606.28057v2 Announce Type: replace-cross Abstract: Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter footprint, prior work proposes hashing many tokens into a single vector within encoder-only models. While...
717. Predicting Metastatic Risk from Primary Cancer Tissue Architecture via Distance-Aware Spatial Modeling ​
Author: Sandesh Pokhrel, Hamid Manoochehri, Beatrice S Knudsen, Tolga Tasdizen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.28676v2 Announce Type: replace-cross Abstract: Predicting distant metastasis from the digital H & E slides of the primary tumor is a critical yet challenging task in computational pathology. Multiple Instance Learning (MIL) approaches can attend to subdomains in whole slide images (WSIs) ...
718. Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery ​
Author: Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric Dieuleveut
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2606.29403v2 Announce Type: replace-cross Abstract: Conformal prediction guarantees marginal coverage, but a pooled calibration quantile can hide systematic undercoverage across heterogeneous regions of the feature space. We introduce Self-Organized Conformal Prediction (SOCP), a calibration s...
719. Memory-Native Non-Terrestrial Networks for Embodied Intelligence ​
Author: Chengyang Li, Yikun Wang, Jiahui He, Yujie Wan, Shuai Wang, Yuan Wu, Yik-Chung Wu, Chengzhong Xu, Huseyin Arslan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA, cs.NI
arXiv:2607.00029v2 Announce Type: replace-cross Abstract: Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in the wilderness to leverage cloud resources or report critical information to remote centers. However, the synergy is nontrivial ...
720. Scene-Conditioned PINN-GNN for Multipath RF Maps: Cross-Scene Generation and In-Scene Completion ​
Author: Lizhou Liu, Xiaohui Chen, Zihan Tang, Mengyao Ma, Wenyi Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2607.01777v2 Announce Type: replace-cross Abstract: Radio frequency (RF) maps provide a compact representation of multipath propagation characteristics and are fundamental to channel modeling, coverage analysis, and environment-aware wireless optimization. This paper proposes a unified RF map ...
721. Cultural Bias Without a Cultural Self:A Disassociation Study of LLM's Persona and Bias ​
Author: Yuan Yuan
Published: 9/1/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.DG
arXiv:2607.02368v3 Announce Type: replace-cross Abstract: Language models prompted with cultural personas increasingly stand in for human respondents in cross-cultural research. Their responses separate personas cleanly, and that separation is read as evidence of a cultural point of view. We show th...
722. Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models ​
Author: Kaiyun Yang, Ruilin Yang, Zhimin Yao, Jikai Wang, Wei Ge
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.02575v2 Announce Type: replace-cross Abstract: Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced, its decisio...
723. A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training ​
Author: Junze Ye, Jiayi Cheng, Miao Lu, Michal Mankowski, Jose Blanchet, Mohsen Bayati
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.04574v2 Announce Type: replace-cross Abstract: For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses full teacher demonstrations, creating a mismatch between teach...
724. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong ​
Author: Nima Kelidari, Mohammadsaeed Haghi, Mahdi Salmani
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT
arXiv:2607.06854v2 Announce Type: replace-cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade, since they beat a random opponent over 99 percent of the time and only tie copies of themse...
725. SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation ​
Author: Andrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan, Victor Martinez, Marcello Federico
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.07469v2 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annota...
726. PRISM Edit: One Vector for All Temporal Answers ​
Author: Chen Huang, Qi Zheng, Ruiqin Zheng, Long Zeng, Yuantong Xu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.11327v3 Announce Type: replace-cross Abstract: Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer shou...
727. Decoupled Structure-Feature Alignment via Alternating Optimization for Graph Learning ​
Author: Chengcheng Yan, Feifei Zhao, Dai Zhu, Wei Liu, Qingsong Wang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.11577v2 Announce Type: replace-cross Abstract: Conventional Graph Neural Networks (GNNs) couple feature transformation and neighborhood aggregation, which often renders them vulnerable to topological noise and heterophilous connections. To decouple this dependency, we present a constraine...
728. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning ​
Author: Zihao Yu, Xiu Yuan, Chongjie Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL
arXiv:2607.14252v2 Announce Type: replace-cross Abstract: Embodied agents accumulate experience over time. We study how accumulated experience can be formed into persistent memory for future reasoning and action. We formulate Embodied Action Memory (EAM) as the capability to form and use memory over...
729. SALT: Salience-Aware Lexical Trie for Long-Context Compression ​
Author: Oteo Mamo, Hyunjin Yi, Joydhriti Choudhury, Shangqian Gao, Weikuan Yu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.PF, cs.AI, cs.LG
arXiv:2607.17486v2 Announce Type: replace-cross Abstract: As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-level prompt compression methods address this, but rank each sen...
730. ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video ​
Author: Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.17790v2 Announce Type: replace-cross Abstract: Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reco...
731. Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft ​
Author: Zeynep Engin, Tim Gordon, Viviana Bastidas, Tom Crick, Jon Crowcroft, Jean-Martin Denis, David J. Hand, Ed Humpherson, Lauren Maffeo, Jakob M"okander, Irene Ng, Anastasija Nikiforova, Giulio Quaggiotto, David Uriel Socol de la Osa, Rhonda Syler, Philip Treleaven, Stefaan Verhulst
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.ET, cs.SI, cs.SY, eess.SY
arXiv:2607.18483v3 Announce Type: replace-cross Abstract: The digital substrate - data, algorithms, infrastructure, platforms, applications - is being governed without adequate conceptual foundations. The ability and legitimacy required to govern this substrate, and to govern with it, are simultaneo...
732. Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information ​
Author: Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Boris Vidolov, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.19313v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning...
733. Error Certificates for KV-Cache Eviction via Randomized Design ​
Author: Peng Xie Amr Alanwar
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21475v3 Announce Type: replace-cross Abstract: Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest, and after the deletion the serving system cannot know what the eviction cost it on the current query. We replace the deterministic tail w...
734. Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models ​
Author: Liangyu Li, Qingwen Liu, Mingqing Liu, Wen Fang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.23602v2 Announce Type: replace-cross Abstract: Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly...
735. SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding ​
Author: Seoyeon Kim, Minjae Kang, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.23991v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient ...
736. A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks ​
Author: Du Yin, Xiachong Lin, Yue Tan, Jinliang Deng, Estrid He, Hao Xue, Flora D. Salim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.25875v4 Announce Type: replace-cross Abstract: Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies typically assume fixed sensor graphs, overlooking the continuous evolution of real-world traffic networ...
737. Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification ​
Author: Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich, Elena Kozachok, Egor Ushakov, Oleg Samovarov
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.26765v2 Announce Type: replace-cross Abstract: Background/Objectives: Dermoscopic skin-lesion classifiers lose accuracy when images arrive from a new clinic or a new device. We asked which data augmentations reduce that loss, and measured the effect under a protocol that keeps policy sele...
738. Learning to Trace Seiberg Dualities ​
Author: Jonathan J. Heckman, Shani Meynet, Alessandro Mininno, Gary Shiu
Published: 9/1/2026, 4:00:00 AM
Categories: hep-th, cs.AI, cs.LG, hep-ph
arXiv:2607.28628v2 Announce Type: replace-cross Abstract: Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when a...
739. Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations ​
Author: Hainiu Xu, Zhaoyue Sun, Hanqi Yan, Jinhua Du, Caroline Catmur, Yulan He
Published: 9/1/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL
arXiv:2607.28648v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, whi...
740. SERUM: State Extraction and Refinement for User Modeling ​
Author: Andy J. Phu, Karin de Langis, James Mooney, Khanh Chi Le, Dongyeop Kang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.29181v2 Announce Type: replace-cross Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a mul...
741. Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy ​
Author: Kaike Ping, Buse \c{C}ar{\i}k, Caleb Wohn, Xiaohan Ding, Tongshuai Wang, Eugenia Rho
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01017v2 Announce Type: replace-cross Abstract: Large language models can answer a medical question correctly and still abandon that answer when a user pushes back. We study this failure as medical sycophancy and ask when models are most likely to give in. Across five open-weight models, 5...
742. Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity ​
Author: Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.02665v2 Announce Type: replace-cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form va...
743. Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation ​
Author: Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03044v2 Announce Type: replace-cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitiv...
744. Approximate Speculative Decoding ​
Author: Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03447v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding...
745. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution ​
Author: Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.08311v3 Announce Type: replace-cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive ...
746. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability ​
Author: Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi, Mislav Balunovic, Theophane Weber, Vahab Mirrokni
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.09538v3 Announce Type: replace-cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical comput...
747. Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics ​
Author: Qingjie Zhang, Ziqi Tang, Jie Zhang, Gelei Deng, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Tianwei Zhang, Han Qiu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10678v2 Announce Type: replace-cross Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expo...
748. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse ​
Author: Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10810v3 Announce Type: replace-cross Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotat...
749. Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction ​
Author: Yi Liu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11318v2 Announce Type: replace-cross Abstract: Many sequential construction tasks have exact terminal symmetries even though execution is directed and depends on history. Process evidence supplies order; terminal correspondence transports it between equivalent outcomes; the realized state...
750. Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs ​
Author: Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11573v2 Announce Type: replace-cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, t...
751. LookBack: Where and How to Score LVLM Responses via Visual Reference Usage ​
Author: Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.11847v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hal...
752. Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference ​
Author: Zixuan Lan, Yanhong Li, Jiawei Zhou
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.13426v2 Announce Type: replace-cross Abstract: Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inferen...
753. Writing Style Similarity Reflects Academic Genealogy ​
Author: Cameron Manzo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.14843v2 Announce Type: replace-cross Abstract: As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate authors. These systems conflate stylistic similarity with individual identity....
754. Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models ​
Author: Yang Liu, Yuming Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.15156v3 Announce Type: replace-cross Abstract: We ask whether a small, directly addressable hidden-state intervention can place a learned world model on an intended counterfactual future and then let the model's own dynamics carry that future forward. In a controlled two-object collision ...
755. PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation ​
Author: Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR
arXiv:2608.16984v2 Announce Type: replace-cross Abstract: Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and...
756. ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback ​
Author: Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.17323v2 Announce Type: replace-cross Abstract: Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performance under ideal conditions but often remain sensitive to small execution errors and distribution shifts....
757. Modeling the Structure of Human Behavior with AI Prompt Vectors ​
Author: Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
Published: 9/1/2026, 4:00:00 AM
Categories: econ.TH, cs.AI
arXiv:2608.18265v2 Announce Type: replace-cross Abstract: We introduce a general, easy-to-implement AI-based method for modeling and analyzing the structure and complexity of human behavior. We assign a large language model a "type vector" and then prompt it to choose actions across settings in whic...
758. Decomposing Wrong-Consensus Agreement in LLM Self-Consistency ​
Author: Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18795v2 Announce Type: replace-cross Abstract: Agreement among repeated samples of a language model is routinely read as evidence about answer reliability, yet wrong answers can agree just as strongly as right ones. This paper asks what information wrong-consensus agreement actually conta...
759. SPADE: Self-Play in Adaptive Synthetic Executable Environments ​
Author: Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.19197v3 Announce Type: replace-cross Abstract: Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribu...
760. What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals ​
Author: Xi Qin, Jizhou Tong
Published: 9/1/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.19269v4 Announce Type: replace-cross Abstract: Benchmarks can run without determining what their results license. We freeze a large evaluation collection and attempt to replay its historical claims. Most units stop because the evidence required for replay is not bound. Where replay is pos...
761. When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception ​
Author: Guangyuan Dong, Chuang Liu, Haoyu Wang, Yangchen Zeng, Jiaqi Zhang, Li Jiuxing, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin, Alexander Lim Han Yang, Yusen Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CV, cs.GR
arXiv:2608.20810v2 Announce Type: replace-cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve. When the scene co...
762. Atom Learning Model (ALM): how a real classroom got tokenised ​
Author: Philipp Bogdan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.21106v3 Announce Type: replace-cross Abstract: The Atom Learning Model (ALM) tokenises a school curriculum. 757 pages of GCSE and Further Mathematics material were read by machine into 1,934 atoms, each one thing a learner can do in a single step, ordered by 4,616 machine-written prerequi...
763. Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data ​
Author: Renfei Zhang, Niloofar Mireshghallah
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.21727v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally i...
764. DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation ​
Author: Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Peilin Huang, Han Fu, Zhuo Li, Xiaoxue Ren
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2608.22067v4 Announce Type: replace-cross Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action m...
765. Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules ​
Author: Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.22642v3 Announce Type: replace-cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. To address these challenges, we presen...
766. LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space ​
Author: Jinghui Zhang, Lang Gao, Ao Li, Mingzhe Li, Ruihong Zeng, Zirui Song, Kentaro Inui, Xiuying Chen
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.23124v3 Announce Type: replace-cross Abstract: Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, existing approaches to author modeling and ...
767. Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation ​
Author: Mohammad Mahdi Dehshibi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2608.23323v2 Announce Type: replace-cross Abstract: Continuous optimisation methods need to balance sharing information and maintaining alternative search directions. In this paper, we introduce Mycelial Search (Myco), a graph-structured metaheuristic designed around active tips, community-wei...
768. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology ​
Author: Marie-Lisa Eich, Kai Standvoss, Timo Milbich, Alexander M"ollers, Miriam H"agele, Philipp Anders, Lars Tharun, Hanna Kontradiuk, Sebastian Kons, Nader Aldoj, Recepcan Adig"uzel, Adam Narai, Lukas H"onig, Jonathan Striebel, Binru Yang, Mihnea P. Dragomir, Marvin Sextro, Philipp Keyl, Philipp Jurmeister, Rosemarie Krupar, Evelyn Ramberger, James Wells, Julika Ribbat-Idel, Andreas Kunft, Hussam Shuaib, Christian Groh'e, Reinhard B"uttner, David Horst, Klaus-Robert M"uller, Lukas Ruff, Maximilian Alber, Frederick Klauschen, Simon Schallenberg
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.23803v2 Announce Type: replace-cross Abstract: Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immunohistochemical, and molecular features. Yet pathological assessment remains largely visual and semi-quan...
769. The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem ​
Author: Elioth Sanabria
Published: 9/1/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.PF
arXiv:2608.23986v2 Announce Type: replace-cross Abstract: Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context. The industry's accounting says this saves money...
770. When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs ​
Author: Zhiyang Qi
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.24080v2 Announce Type: replace-cross Abstract: In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, an...
771. Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity ​
Author: Heng Zhang, Haotian Xiang, Konstantinos D. Polyzos, Tara Javidi, Qin Lu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.24721v2 Announce Type: replace-cross Abstract: Hyperparameter selection remains a key challenge in Bayesian optimization (BO) and Bayesian active learning (AL), as model misspecification can lead to suboptimal performance, while more accurate fully Bayesian treatments typically rely on co...
772. FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference ​
Author: Gongwei Lee, Ji Liu, Juncheng Jia, Ji Wu
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC
arXiv:2608.24945v2 Announce Type: replace-cross Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stan...
773. Demystifying Reinforcement Learning Post-Training of Language Models ​
Author: Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.24949v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers and practitio...
774. DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning ​
Author: Menghui Zhou, Zhipeng Yuan, Vitaveska Lanfranchi, Po Yang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.25073v2 Announce Type: replace-cross Abstract: Digital mobility outcomes (DMOs) derived from wearable sensors characterise mobility in daily life and offer a promising means of monitoring disease progression. However, existing DMO studies have typically focused on either a single disease ...
775. MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective Alignment ​
Author: Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.25200v2 Announce Type: replace-cross Abstract: We study learning a mixture of $k$ Plackett-Luce models from multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. ...
776. It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning ​
Author: Liam P. H. Mertens, Lucas N. Alegre, Florent Delgrange, Diederik M. Roijers, Ann Now'e, Peter Vamplew
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.25723v2 Announce Type: replace-cross Abstract: Time is of the essence when dealing with multiple reward signals and non-linear utility. In this paper we argue that the current main approaches in multi-objective RL (SER and ESR), and successor features, are insufficient. While each approac...
777. A Statistical Audit of Physical AI Benchmark Redundancy ​
Author: Zaruhi Navasardyan, Hrant Davtyan
Published: 9/1/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.25940v2 Announce Type: replace-cross Abstract: Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI b...
778. ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations ​
Author: Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sonja Greven, Kerstin Ritter
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ML
arXiv:2608.26083v2 Announce Type: replace-cross Abstract: Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Auditing for shortcuts requires testing many candidate concepts, such as acquisition settings or demographics. Current concept-based explainabilit...
779. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models ​
Author: Frederik Berenz
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27367v3 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-provisioned for simple tasks and under-provisioned for complex ones, with significant redundancy across ...
780. XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering ​
Author: Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27481v2 Announce Type: replace-cross Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries in...
781. A Survey on Rubric-Guided Reinforcement Learning for Language Models ​
Author: Zifei Shan, Fangning Shao
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27505v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail...
782. Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification ​
Author: Cameron Wilding, Mina Shaker, Fatemeh Ganji
Published: 9/1/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.27954v2 Announce Type: replace-cross Abstract: Post-deployment changes to large language models can alter behavior while leaving routine outputs largely unchanged, creating a challenge for AI governance when model weights are proprietary. We present a privacy-preserving zk-SNARK-based aud...
783. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss ​
Author: Niccol`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
Published: 9/1/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28308v2 Announce Type: replace-cross Abstract: We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates and batch sizes, we investigate their marginal evolution with...