arXiv cs.AI - 2026-08-28 ​
312 items collected.
1. EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction ​
Author: Yu Fu, Yongqi Kang, Yong Zhao, Rongfang Bie
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpr...
2. Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset ​
Author: Di Zhu, Chen Xie, Haoyun Zhang, Zihan Wei, Ziwei Wang, Jiazhao Shi, Ziyu Wang, Qiyang Xie
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a p...
3. Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap ​
Author: Jiale Liu, Huan Wang, Weicheng Wang, Rong Zhu, Qiqi Wang, Min Xie
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26111v1 Announce Type: new Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electric vehicles, grid storage, and consumer electronics. Conventional BPHM approaches, including physics-ba...
4. PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices ​
Author: Deepak Vungarala, Deniz Najafi, Abdulrahman Aljoudi, Zahra Ghanaatian, Navid Khoshavi, Gourav Datta, Arman Roohi, Mahdi Nikdast, Shaahin Angizi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PD...
5. CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering ​
Author: Kunjesh Parekh, Anil Kumar Tiwari, Divya Saxena
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, q-fin.CP
arXiv:2608.26114v1 Announce Type: new Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, th...
6. The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning ​
Author: Marko Cvjetko, Benedikt Hartl, Michael Levin, Cl'ement Moulin-Frier, Pierre-Yves Oudeyer
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26116v1 Announce Type: new Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop fr...
7. The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting ​
Author: Jaeik Jeong, Tai-Yeon Ku, Wan-Ki Park
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26134v1 Announce Type: new Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device forecasting for mission-critical edge environments, including military systems. However, this paper identi...
8. LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs ​
Author: Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.IR
arXiv:2608.26145v1 Announce Type: new Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in suppo...
9. Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse ​
Author: Edouard Lansiaux, Hugo Kazzi, Aur'elien Loison, Slim Hammadi, Emmanuel Chazard
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.26149v1 Announce Type: new Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality ca...
10. Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models ​
Author: Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DL, cs.IR
arXiv:2608.26150v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extrac...
11. Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration ​
Author: Sandeep Gaddamwar
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.26151v1 Announce Type: new Abstract: Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature markets eroding billions in revenue annually. Predictive models can flag at-risk customers accurately, yet they are ro...
12. EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG ​
Author: Xuan-The Tran, Le Trung Kien Nguyen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26153v1 Announce Type: new Abstract: Clinical electroencephalography (EEG) reporting remains largely manual and time-consuming, and current EEG software ecosystems do not produce the structured EEG-text supervision needed for training modern language models. Most toolboxes focus on visual...
13. Selection Bias Correction in Retail Intelligence ​
Author: Spandan Ghose Chowdhury
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.26156v1 Announce Type: new Abstract: Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "long tail" of niche items. This simulation study investigates selection bias in inflation estimation and compares c...
14. GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions ​
Author: Aravind Sasidharan Pillai
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26157v1 Announce Type: new Abstract: Natural-language analytics over enterprise data warehouses is increasingly important, but production use is limited by hallucinated metrics, invalid joins, wrong grain, unsafe data access, and unsupported explanations. Existing text-to-SQL systems ofte...
15. SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum ​
Author: Viorica Rozina Chifu, Tudor Cioara, Vasile Ofrim, Liana Toderean, Ionut Anghel, Laura Daniele, Cornelis Bouter
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26160v1 Announce Type: new Abstract: Nowadays semantic models provide limited support for representing distributed AI workflows and their execution across heterogeneous edge, fog, and cloud environments. Therefore, AI processes and resources are often described using incompatible semantic...
16. A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian ​
Author: Lei Wang, Xiao Wang, Lei Li
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26162v1 Announce Type: new Abstract: Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health sup...
17. A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving ​
Author: Ibrokhimsho Abduchaborov
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26164v1 Announce Type: new Abstract: Large language models can interpret natural-language chemistry questions, but their internal reasoning is difficult to inspect, constrain, and validate. This paper presents ChemOntoRule, a proof-of-concept symbolic core for AI-assisted school-level che...
18. Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript ​
Author: Sagnik De, Sreenija Pavuluri
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.LG, eess.AS
arXiv:2608.26167v1 Announce Type: new Abstract: Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to distinguish appropriate abstention from an unsupported prediction. Seven large language models were evaluated on the T...
19. Knowledge Cards: Structured Knowledge for AI Systems ​
Author: Liliana Ferreira
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.26176v1 Announce Type: new Abstract: AI systems whose outputs inform real decisions, and increasingly consequential ones, require something that current documentation practice does not provide: a structured, inspectable representation of the knowledge they need to ground, contextualize, a...
20. AI Revealed Preferences ​
Author: Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26178v1 Announce Type: new Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run th...
21. Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI ​
Author: Ashita Ashok, Franziska Babel, Patrick Holthaus, Rucha Khot, Karla Bransky, Fethiye Irmak Dogan, Karsten Berns, Silvia Rossi, Minha Lee, Guy Laban
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO
arXiv:2608.26182v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt design in human-robot interaction (HRI) remains underspecified. As a result, robots may present hallucinated capabilities, unclear behavioural bounda...
22. TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education ​
Author: David Barron, Xiaohang Tang, Rezky Dwisantika, Minsun Kim, David H. Smith IV, Jiaming Cui, Yan Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.26184v1 Announce Type: new Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to adapt support to learners' needs. We present TutorTrace, a dataset and behavioral abstraction pipeline that makes learners' behavioral context visib...
23. Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion ​
Author: Yue Shen, Rehema Abulikemu, Ryan P. McMahan, Yan Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.HC
arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or professional consequences, especially when raising a point requires voicing it themselves. ...
24. Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments ​
Author: Huy Nguyen, Yue Lin
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26188v1 Announce Type: new Abstract: Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe se...
25. Invocation-Level Reliability of Tool-Using Agents ​
Author: Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.26189v1 Announce Type: new Abstract: Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. We measure a correct-invocation rate that separates the two, under both a clean teache...
26. Predicting Consequences and Reinforcing Navigation Policies with Latent World Models ​
Author: Zengmao Wang, Wei Gao, Shuhan Shen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26190v1 Announce Type: new Abstract: World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approaches primarily focus on reconstructing future observations or features, which introduces unnecessary complexity a...
27. Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs ​
Author: Animesh Agarwal, Meysam Ghaffari, Nina Fatehi, Carlos Morato
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because relevant signals are multimodal, weak in isolation, and distributed across irregular patient histories. We propose structured evidence routing, a router-...
28. AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes ​
Author: Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26193v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro e...
29. Agentic AI for operating scientific instruments for nanoscale characterization ​
Author: Zahra Ayar, Marcos Penedo, Mahdi Mehdikhani, Nahid Hosseini, Prabhu Prasad Swain, Georg E. Fantner
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, physics.ins-det
arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making. A trained user defines the experimental intent, translates it into instrument commands, assesses incoming data, adjusts imaging param...
30. Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling ​
Author: Leonardo Liparulo, Francesco Pierri
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.26199v1 Announce Type: new Abstract: We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows in an industry-realistic tool-calling setting. In these environments, engineers issue repetitive, dependency-order...
31. GameWAM: A World Action Model for Video Games ​
Author: Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.26200v1 Announce Type: new Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas...
32. Same Model, Different Harness: Different Coding-Agent Results ​
Author: Sydney Lewis
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurati...
33. Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy ​
Author: Mazhar Shaikh, Anurag Rajkumar Bombarde, Harshal Pathak
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.MA, cs.SE
arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators reach for is the service mesh's: retry, timeout, and error-rate circuit breaking. We report a ...
34. LLM Agents for Time-Series: A Survey ​
Author: Yilong Chen, Xiao Qin, Chenghao Liu, Liang Wu, Noelle I. Samia, Kaize Ding
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-series problems they add...
35. The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts ​
Author: Sachin Gopal Wani, Ajay Dholakia, David Ellison
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.PF
arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the a...
36. 6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation ​
Author: Brandon Colelough, Vladimir Martirosyan, Ishan Tamrakar, William Regli, Aditya Kumar, Anh N. Nhu, Dhruv Dubey, Raj Ambavane, Haowei Deng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.26236v1 Announce Type: new Abstract: We present a six-stage framework for auditing the reproducibility of scientific claims across a research literature within the computer science domain, and instantiate our framework for the neuro-symbolic AI (NSAI) subdomain. Instantiating the framewor...
37. SKILL.state: Scalable Long-Horizon Agent Skills ​
Author: Sanket Badhe, Priyanka Tiwari, Jonghyun Chung
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an eve...
38. Assessing mentalization in humans and large language models ​
Author: Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, q-bio.NC
arXiv:2608.26291v1 Announce Type: new Abstract: Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mi...
39. Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems ​
Author: Ilai Shraga, Roei Eshel, Lior Gorelik
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26306v1 Announce Type: new Abstract: A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: ...
40. FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence ​
Author: Ziyu Wang, Qiming Dai, Yishan Wu, Zaiwen Wen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26310v1 Announce Type: new Abstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based na...
41. ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving ​
Author: Wenqian Ye, Ziwei Guan, Eric Xie, Bohan Liu, Shivani Modi, Buyun Zhang, Ellie Dingqiao Wen, Henry Kautz, Aidong Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26334v1 Announce Type: new Abstract: Automated theorem proving offers a natural foundation for recursive self-improvement in scientific discovery. However, existing neural provers do not fully preserve this recursive structure, where the learning process should be self-improving over time...
42. Fine-Tuning of Transformer models with Frames ​
Author: Harshavardhan Adepu, Li Zhang, Sanjiv Kumar, Vikas Singh
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26430v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $...
43. Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI ​
Author: Md Jueal Mia, M. Hadi Amini
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.26442v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed token budgets...
44. PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents ​
Author: Yang Xiao, Yusong Sun, Haoyi Wu, Wenyang Hui, Wen Da, Zhaokai Luo, Mu Chuan, Yao Hu, Wenjie Li, Chengyue Jiang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26530v1 Announce Type: new Abstract: Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate...
45. Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation ​
Author: Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang, Peiwei Zhao, Junchi Chen, Yunfeng Diao, He Liu, Qi Chu, Tao Gong, Nenghai Yu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26535v1 Announce Type: new Abstract: Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no...
46. DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows ​
Author: Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and mor...
47. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling ​
Author: Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge rel...
48. SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation ​
Author: Li Mingqian
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.26683v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) faces significant challenges in maintaining robust coordination under noisy observations. Although observation disturbances are often introduced independently across agents, their downstream effects...
49. Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation ​
Author: Hyeonchu Park, Bugeun Kim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26694v1 Announce Type: new Abstract: Detecting AI-generated text (AIGT) remains challenging because existing approaches rely on token-level statistical signals or independent stylometric features, causing them to overfit to specific generators and fail under distribution shift. We identif...
50. Five Primitives for Governing Autonomous AI Agents at Runtime ​
Author: Jiten Oswal, John Cadeddu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.SE
arXiv:2608.26696v1 Announce Type: new Abstract: Enterprise deployments of autonomous AI agents inherit a control model built for human users and long-lived services, and the fit fails in three specific ways: agent principals are ephemeral, appearing and vanishing faster than provisioning; their acti...
51. Accelerating Scientific Research with Gemini in the Real-World ​
Author: Samuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin Li'{e}vin, Jingyun Yang, Yuchen Zhuang, Tim Strother, Alex Bijamov, Min Woo Sun, Anil Palepu, Justin Chen, David Steiner, Jacqueline Shreibati, Wei-Hung Weng, Yilin Zhao, Xingjian Hu, Nicholas Zahn, Sadhya Garg, Julia Kirby, Yuxiang Gan, Jiaoli Li, Divy Thakkar, Shekoofeh Azizi, David Racz, Juraj Gottweis, Vivek Natarajan, Chenglin Wu, Tal Danino, Keran Rong, Haozhe Wang, Benoit Schillings, Yong Cheng, Quoc V. Le, Tao Tu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26701v1 Announce Type: new Abstract: We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving be...
52. Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing ​
Author: Hyeonchu Park, Gahye Jeong, Bugeun Kim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26710v1 Announce Type: new Abstract: AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high fa...
53. Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training ​
Author: Tingyun Li, Wenfeng Feng, Weiqing Li, Abudukelimu Wuerkaixi, Guohua Liu, Yuewei Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26730v1 Announce Type: new Abstract: Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often entails repeated post-training. Autonomous systems automate parts of this process by proposing updates, training candidates, and using ...
54. Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds ​
Author: Waqas Khan, Tabinda Sarwar, Jingyue Cong, Xun Yi, Estrid He
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26743v1 Announce Type: new Abstract: Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model utility, offering a practical alter...
55. AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design ​
Author: Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang, Shuting Jin, Zhuo Yang, Tianfan Fu, Fang Wu, Xiangxiang Zeng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26747v1 Announce Type: new Abstract: Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code changes...
56. Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case ​
Author: Ahlame Diouan (ERIC, UL2), Eric Ferey (ERIC, UL2), Sabine Loudcher (ERIC, UL2), J'er^ome Darmont (ERIC, UL2)
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26750v1 Announce Type: new Abstract: Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel, a two-stage method...
57. DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation? ​
Author: Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26757v1 Announce Type: new Abstract: Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-ready quantities, and rendering them accurately. Modern LLMs can produce visually plausible, instruction-compliant c...
58. Categorizer Automata for Discounted-Sum Payoffs ​
Author: Nathalie Bertrand, Pranav Ghorpade, Senthil Rajasekaran, Sasha Rubin, Moshe Vardi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.FL
arXiv:2608.26763v1 Announce Type: new Abstract: Categorizing continuous data into discrete bins is a fundamental operation in artificial intelligence. We introduce the categorizer automaton, a deterministic automaton that reads an infinite sequence of rewards and identifies which of finitely many bi...
59. AI Control Scientist: LLM-driven Agentic System for Automated Control Design ​
Author: Haiteng Wang, Weihao Li, Jing Zhang, Lei Ren
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26780v1 Announce Type: new Abstract: Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tuning, resulti...
60. Decoupling Planning and Control for Instructable Agents ​
Author: Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA, cs.RO
arXiv:2608.26788v1 Announce Type: new Abstract: Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to realize such plans as reliable low-latency action sequences in unfamilia...
61. SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers ​
Author: Haizhao Fan, Yuchi Xiong, Jize Wang, Xinping Guan, Xinyi Le
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.26836v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing solutions like Chain-of-Th...
62. LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems ​
Author: Jiaqi Xu, Yiran Qiao, Jing Chen, Qiwei Zhong, Xiang Ao, Xueqi Cheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.MA
arXiv:2608.26849v1 Announce Type: new Abstract: User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in so...
63. BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click ​
Author: Mesut Toruk
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26867v1 Announce Type: new Abstract: Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with...
64. C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning ​
Author: Yassir Lairgi, Ludovic Moncla, Khalid Benabdeslem, R'emy Cazabet, Pierre Cl'eau
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SI
arXiv:2608.26870v1 Announce Type: new Abstract: Weak signals are early, low-visibility indicators that precede significant changes before those changes become established. Existing detection methods, based on keyword frequency, topic modeling, or untyped graph topology, fail to capture the semantic ...
65. Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency ​
Author: Nikol Figalov'a, Lynn Huestegge, Anne B"ockler-Raettig
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.SE
arXiv:2608.26885v1 Announce Type: new Abstract: Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a pre...
66. Learning-Augmented Online Allocation under Unreliable Advice: Robustness, Exposure Fairness, and Distribution Shift ​
Author: Fredy Pokou (MRE, CRIStAL)
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26889v1 Announce Type: new Abstract: Learning-augmented algorithms improve online decisions using predictions, but unreliable advice may harm efficiency and fairness. We study an online allocation problem with finite candidate sets, irreversible decisions, and exposure constraints. We pro...
67. AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion ​
Author: Jakub Seredy'nski, Georgios Tsaousoglou
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.MA, cs.SY, eess.SY
arXiv:2608.26896v1 Announce Type: new Abstract: As electricity market participants increasingly adopt learning-based agents for their bidding strategies, electricity markets are becoming algorithmic. Evidence from algorithmic markets in other domains shows that tacit collusion can arise purely throu...
68. Counterfactual Bias Testing for Application Tracking System ​
Author: Sai Yashwant, Shruti Bansal, Anurag Dubey, Samaroha Chatterjee, Satyam Kumar, Shreyash Gupta, Gantala Thulsiram
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26899v1 Announce Type: new Abstract: Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demographic bias is expensive: classical correspondence-audit studies require hand-crafted resumes and manual submissi...
69. A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering ​
Author: I~nigo Alonso, Mirella Lapata
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26949v1 Announce Type: new Abstract: Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression, which represents context as images, promises to reduce token cost, but its effect on table understanding remain...
70. From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities ​
Author: Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.26950v1 Announce Type: new Abstract: Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagn...
71. GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory ​
Author: Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26983v1 Announce Type: new Abstract: Organizing long-term memory for multimodal agents remains challenging because existing methods either suffer from expensive question-agnostic offline summaries or naive embedding similarity matching that introduces incomplete and redundant context. To ...
72. DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research ​
Author: Linsen Zhu, Yi Shi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.26990v1 Announce Type: new Abstract: Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evidence, expose unavailable data and model capabilities, and control how generated opinions affect a final report. We...
73. ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions ​
Author: Rui Xie, Lu Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26991v1 Announce Type: new Abstract: Powerful code agents can execute scripts, call tools, and manage files, yet many important applications remain accessible primarily through graphical user interfaces. We argue that screenshot-and-click is an inefficient interface for software-operating...
74. A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems ​
Author: Varvara Mama, Eleni Veroni, Nikolaos Kapsalis, Christos D. Nikolopoulos, Anargyros T. Baklezos
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27010v1 Announce Type: new Abstract: In the present work an efficient border control management procedure is proposed. Compared to operational queue management systems, whose operations are based on mostly static data, the proposed work takes into account dynamic traffic conditions, thus ...
75. Omni-Interactive Universal Embedder ​
Author: Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.27044v1 Announce Type: new Abstract: Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, existing approaches primarily foc...
76. A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes ​
Author: Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Zhenghe Hou, Jiaxing Song
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE
arXiv:2608.27086v1 Announce Type: new Abstract: Enterprise AI deployment is a coordination problem across business units, application and AI teams, testing, platform engineering, infrastructure, security, operations, and data governance. Use-case benchmarks show whether one agent completes one task,...
77. pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning ​
Author: Shivam Mishra, Dhannu Ram Meena, Muneendra Ojha, Krishna Pratap Singh, Kuldeep Singh
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27101v1 Announce Type: new Abstract: Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain terms, produce inconsistent formats, and favor hierarchical over associative relations. In the LLMs4OL 2026 Chall...
78. LAAF: A Layered Accountability Architecture Framework for LLM Applications ​
Author: Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi, Muhammad Waqas, George Loukas, Alireza Jolfaei, Lucas Cordeiro, Pierre Dantas
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.27102v1 Announce Type: new Abstract: Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, ...
79. TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation ​
Author: Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27127v1 Announce Type: new Abstract: Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes across cultures and languages is a central challenge fo...
80. GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL ​
Author: Zike Yuan, Han Zhang, Jianzhi Yan, Le Liu, Cai Ke, Huozhi Zhou, Jian Xie, Jiran Yin, Yukun Cao, Yue Yu, Hui Wang, Ming Liu, Bing Qin
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27142v1 Announce Type: new Abstract: Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topological structu...
81. Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection ​
Author: Xiang Wang, Zhijun Cheng, Zhenyu Meng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27144v1 Announce Type: new Abstract: In recent years, graph anomaly detection (GAD) based on frequency-domain filtering have achieved promising results. However, existing approaches still face three major challenges: First, they use static basic function to constructed graph filter which ...
82. When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents ​
Author: Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang, Dongjin Yu, Yu Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.27146v1 Announce Type: new Abstract: Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer merely provide data but begin to specify concrete actions, they effectively become ``commands'' that can drive real...
83. Thomson: Continual Learning of Frontier Models for SovereignAI ​
Author: Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofr`e, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27147v1 Announce Type: new Abstract: The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent publ...
84. BPMN4CAI: A BPMN Extension for Modeling Dynamic Conversational AI ​
Author: Bj"orn-Lennart Eger, Daniel Rose, Barbara Dinter
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27149v1 Announce Type: new Abstract: Conversational AI systems, such as chatbots and virtual assistants, are becoming increasingly important to digital business processes. However, the established Business Process Model and Notation (BPMN) standard faces challenges when representing dynam...
85. Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable ​
Author: Pranav Aggarwal
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27167v1 Announce Type: new Abstract: An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is esc...
86. What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents ​
Author: Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, Weiwen Liu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27260v1 Announce Type: new Abstract: LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience th...
87. Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search ​
Author: Yuan Chang, Xiaoqi Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27266v1 Announce Type: new Abstract: Efficiently improving autonomous agents across diverse tasks is central to accelerating recursive self-improvement (RSI) in agentic AI, with prompt optimization emerging as a promising approach capable of delivering performance gains comparable to thos...
88. BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models ​
Author: Jinghan Zhang, Fengran Mo, Zhiyu Chen, Xiaoyan Han, Kunpeng Liu, Chang-Tien Lu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2608.27268v1 Announce Type: new Abstract: Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way. However, it is unclear whether existing AI systems are inclusive enough for blind and d...
89. LLMs Can Design Near-Optimal OR Algorithms ​
Author: Jackie Baek
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.27296v1 Announce Type: new Abstract: We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. We study inventory control, queueing network control, and assortment optimization. We evaluate two levels of LLM use: at l...
90. Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification ​
Author: Jinghan Xu, Yikai Zhang, Aili Chen, Weiyuan Li, Jiaqing Liang, Deqing Yang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27311v1 Announce Type: new Abstract: Agent harnesses shape how language-model agents use instructions, tools, and runtime components, but adapting these harnesses requires costly verification. Existing propose-and-verify methods typically score every candidate on a fixed task set, wasting...
91. Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance ​
Author: Allison Zhuang, Santiago Aranguri
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27340v1 Announce Type: new Abstract: Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized e...
92. Sophistication in GenAI Use: Field Evidence from a Large Firm ​
Author: Nicholas J. Hallman, Zachary T. Kowaleski, Anu Puvvada, Jaime J. Schmidt
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.27364v1 Announce Type: new Abstract: We study how sophistication in generative AI (genAI) use varies among the back-office workforce of a large firm. Using proprietary data, we observe 713,564 employee prompts and their corresponding large language model responses from nearly 4,000 back-o...
93. CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases ​
Author: Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG
arXiv:2608.27391v1 Announce Type: new Abstract: LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench...
94. Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study ​
Author: Kevin Zhu, Ryan Zhang, Baraa Abed, Tilendra Choudhary, Malvern Madondo, Mehak Arora, Yixuan Yang, Alasdair Gent, Aditya Nagori, Omer T. Inan, Krista L. Haines, Patrick Georgoff, Suresh M. Agarwal, Vijay Krishnamoorthy, Tetsu Ohnuma, Mihai V. Podgoreanu, Michael R. Pinsky, Gilles Clermont, Craig M. Coopersmith, Craig S. Jabaley, Rishikesan Kamaleswaran
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.27421v1 Announce Type: new Abstract: Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. No alternative learned directly from patie...
95. Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation ​
Author: Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27429v1 Announce Type: new Abstract: Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecules or through heuristic graph edits that operate directly on molecular ...
96. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution ​
Author: Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, Tu Vu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27454v1 Announce Type: new Abstract: Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. ...
97. Exploring the Role of LLMs in HPC Programming: A Survey ​
Author: Strahinja Ljaljevic, Josep Jorba, Sergio Iserte
Published: 8/28/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.26110v1 Announce Type: cross Abstract: Large Language Models (LLMs) are emerging as promising assistants in High-Performance Computing (HPC), where programming remains complex and expertise-intensive. This survey systematically reviews their application across five categories: code genera...
98. From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement ​
Author: Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy
Published: 8/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.26117v1 Announce Type: cross Abstract: RDBMS (Relational Database Management System) databases face several limitations, including slow execution with multi-hop queries and a lack of explainability by graphical interpretations. In contrast, Graph database offers a more intuitive and effic...
99. Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales ​
Author: Muhammad Deedahwar Mazhar Qureshi, Sannaan Khan, Muhammad Atif Qureshi, Wael Rashwan
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26125v1 Announce Type: cross Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when de...
100. FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes ​
Author: Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.26129v1 Announce Type: cross Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist que...
101. Syntax vs. Semantics: How Transformers Learn Deep Dependencies ​
Author: Jiangrui Zhao, Xiaoting Du
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26139v1 Announce Type: cross Abstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as ...
102. Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation ​
Author: Yuhan Liu, Yixiong Zou, Yuhua Li, Ruixuan Li
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26142v1 Announce Type: cross Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES performance, their proh...
103. Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes ​
Author: Muhammad Jawad Chowdhury, Adiba Hasan, Ishrak Hossain, Shahriar Ivan, Sabbir Ahmed
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26143v1 Announce Type: cross Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It r...
104. Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search ​
Author: Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.26152v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or e...
105. Evaluating AI Generated Summaries for Cancer Patients ​
Author: Muhammad Aurangzeb Ahmad, Kim Shyu, Leon Oliver, Fergus Sleight, Paul Landau
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26154v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns abou...
106. VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation ​
Author: Yixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26155v1 Announce Type: cross Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text supervision. Although multi...
107. Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation ​
Author: Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26159v1 Announce Type: cross Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors. Specifically, an LLM may recognize outputs from other copies of the same model and ma...
108. Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models ​
Author: Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26161v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently exhibit systematic bi...
109. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents ​
Author: Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.MA, cs.SD
arXiv:2608.26163v1 Announce Type: cross Abstract: Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming ...
110. Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment ​
Author: Sam Grouchnikov, Phillip Gregory, Jiho Noh
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26165v1 Announce Type: cross Abstract: Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking practical accuracy. We introduce a novel approach by using Poly-Encoder for computationally efficient and accurate ...
111. Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning ​
Author: Philipp Schr"oppel
Published: 8/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.26166v1 Announce Type: cross Abstract: Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - intermediate steps toward solutions - open up high-stakes applications by enabling human inspection of AI decision-m...
112. Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention ​
Author: Naveen Lamba, Sanju Tiwari, Manas Gaur
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26168v1 Announce Type: cross Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-stakes environments, including health, legal, and scientific research. Although previous surveys have primarily ...
113. DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling ​
Author: Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
Published: 8/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.26172v1 Announce Type: cross Abstract: Deploying natural-language interfaces over enterprise OLTP catalogs fails at scale because semantic parsers collapse under schema-graph scaling, inflating context beyond stable LLM attention budgets. We present DRL (Deterministic Relational Middlewar...
114. ClassVision: AI-Powered Classroom Attendance System ​
Author: Ankit Kumar Aggarwal, Veerabhadra Rao Marellapudi, Ovadia Sutton, Youshan Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CV, cs.LG
arXiv:2608.26173v1 Announce Type: cross Abstract: Students and working professionals have to go through the attendance process every day. Traditional methods of marking attendance using pen and paper or online platforms are human-intensive and time-consuming. To address the challenges in manual atte...
115. Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors ​
Author: Mantas Lukauskas
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26175v1 Announce Type: cross Abstract: Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token premium: the same ...
116. A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs ​
Author: Yifan Song
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26177v1 Announce Type: cross Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic of the ``lost-in-th...
117. PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants ​
Author: Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26180v1 Announce Type: cross Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of...
118. Investigating the Influence of Prompt and Response Languages on LLM Content Generation ​
Author: Thi Thanh Nhan Nguyen, Mai Khoi Tieu, Michael A. Riegler, P{\aa}l Halvorsen, Thu Nguyen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26186v1 Announce Type: cross Abstract: This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answers to 68 non translation questions across four language conditions: English to English, English to Norwegian, N...
119. When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models ​
Author: Dai Shi, Xiaoyu Li, Jos'e Miguel Hern'andez-Lobato
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.LO
arXiv:2608.26187v1 Announce Type: cross Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable o...
120. Comparing Chunking and Embedding Strategies for Turkish RAG Systems ​
Author: Mustafa Serta\c{c} T"urkel, Fatma Nur Korkmaz, Ahmet Tu\u{g}rul Bayrak
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26192v1 Announce Type: cross Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG) quality, yet neither has been systematically studied for morphologically rich languages such as Turkish. We com...
121. A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers ​
Author: Inho Kim, Sumyeong Ahn
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.26194v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Language Models (LLMs). Although knowledge databases for RAG are increasingly diversifying to include various modal...
122. ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices ​
Author: Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2608.26204v1 Announce Type: cross Abstract: Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguous instructi...
123. Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs ​
Author: Ranit Debnath Akash, Ashish Kumar, Gang Tan, Saeid Tizpaz-Niari
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.26209v1 Announce Type: cross Abstract: Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from criminal justice to financial lending. However, these systems often exhibit individual discrimination---unjustified disparities in which a program yiel...
124. Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model ​
Author: Ross Williams, Niyousha Hosseinichimeh
Published: 8/28/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, cs.LG, cs.MA
arXiv:2608.26221v1 Announce Type: cross Abstract: As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, prod...
125. NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation ​
Author: Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.SE
arXiv:2608.26222v1 Announce Type: cross Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically re...
126. How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation ​
Author: Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Con...
127. On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study ​
Author: Aditya Pratap Singh
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.26292v1 Announce Type: cross Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report that current knowledge-editing benchmarks cannot measure this decision at all. Using INLAY, a gradient-free edito...
128. MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models ​
Author: Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov, Alexander Panchenko, Ilseyar Alimova
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA, cs.SE
arXiv:2608.26295v1 Announce Type: cross Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce MemToC, a controlled b...
129. Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models ​
Author: Rohit Patel, Dieuwke Hupkes, Sloan Strader
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2608.26317v1 Announce Type: cross Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We...
130. How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models ​
Author: Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26327v1 Announce Type: cross Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-m...
131. Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds ​
Author: JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.26346v1 Announce Type: cross Abstract: We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group delay analysis. Our central finding is that AI-generated impulsive sounds show near-identical onset-region group-delay distributions but exhibit me...
132. Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives ​
Author: Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26372v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would p...
133. CG4AI: A Column Generation Framework for Training AI Models Under Constraints ​
Author: Youcef Magnouche, Abderrahmane Driouch, S'ebastien Martin, Pierre Bauguion
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DM
arXiv:2608.26375v1 Announce Type: cross Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined rules or constraints on its outputs. In many real-world applications, ranging from autonomous systems...
134. Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries ​
Author: Alden Do Rosario, Hussein Younes, Felipe Pires
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26385v1 Announce Type: cross Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one that declines when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kala...
135. Co-Evolving Structured Knowledge and Reasoning in Language Models ​
Author: Ryan Thomas Noonan, Linxi Zhao, Menghan Xu, Akanksha Sarkar, Mihir Mishra, Dongyoung Go, Kilian Q. Weinberger, Yoav Artzi, Jennifer J. Sun
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.26386v1 Announce Type: cross Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured text often introduces irrelevant context and offers limited control over the retrieved information. Structured k...
136. Simultaneous Envy and Equitability Guarantees ​
Author: Hadi Hosseini, Shraddha Pathak, Lirong Xia, Chengkai Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, econ.TH
arXiv:2608.26410v1 Announce Type: cross Abstract: Recent work in fair division has focused on either simultaneously satisfying closely related fairness notions or achieving a single notion across the ex-ante and ex-post worlds. We study the compatibility of two fundamentally different fairness notio...
137. Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI ​
Author: Architect Labs
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.26418v1 Announce Type: cross Abstract: Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and...
138. The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection ​
Author: Jaturong Kongmanee, Smile Thanapattheerakul
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR
arXiv:2608.26423v1 Announce Type: cross Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxo...
139. SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning ​
Author: Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL
arXiv:2608.26432v1 Announce Type: cross Abstract: Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call ...
140. Diff Mining: Logit Differences Reveal Finetuning Objectives ​
Author: Greg Kocher, Robert West, Cl'ement Dumas, Julian Minder
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.26462v1 Announce Type: cross Abstract: Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains unclear exactly which behaviors emerge during this process. As models grow ever more capable, understanding finetun...
141. Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance ​
Author: Victor Gao (Sang Won), Vida Khosrowshahi (Sang Won), Ali Khosrowshahi (Sang Won), Xihao Sun (Sang Won), Juhyun Lee (Sang Won), Simon (Sang Won), Lee
Published: 8/28/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL, cs.SE
arXiv:2608.26480v1 Announce Type: cross Abstract: Multi-agent large language model systems are widely reported to beat single-model baselines, but the evidence is mixed, and comparisons are usually confounded: pipelines change token budgets, tool calls, and prompts simultaneously, so an aggregate ga...
142. RTNav: Towards Real-Time Zero-Shot Object Navigation ​
Author: Easop Lee, Lingyu Zhang, Boyuan Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.26496v1 Announce Type: cross Abstract: Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation models. However, these models also introduce non-negligible inference latency, which becomes an important conce...
143. Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations ​
Author: Yuehao Song (School of Automation, Central South University, Changsha, China), Zhong Chen (School of Automation, Central South University, Changsha, China), Lihui Cen (School of Automation, Central South University, Changsha, China), Liang Wu (Johns Hopkins University, Baltimore, USA), Kai Zhang (State Key Laboratory of Simulation and Regulation of Water Cycle in River Basin, China Institute of Water Resources and Hydropower Research, Beijing, China)
Published: 8/28/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.LG, cs.NA, cs.SY, eess.SY
arXiv:2608.26549v1 Announce Type: cross Abstract: While Physics-Informed Neural Networks (PINNs) have emerged as a transformative paradigm for solving complex differential equations, their reliance on backpropagation-based gradient descent and automatic differentiation (AD) imposes significant compu...
144. J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data ​
Author: Gyouk Chu, Myeongho Jeon, Eunho Yang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.26582v1 Announce Type: cross Abstract: Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverif...
145. Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries ​
Author: Alistair Reid, Simon O'Callaghan, Dustin Venini, Liam Carroll, Tiberio Caetano
Published: 8/28/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.26626v1 Announce Type: cross Abstract: This report presents a framework to help organisations, policymakers and researchers reason about the risks that emerge when AI agents interact with each other, how those risks change as interactions cross organisational boundaries, and the controls ...
146. CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes ​
Author: Yuanxiang Ni, Xianliang Huang, Chenhang Ma, Chen Xiao, Yuewen Ma, Ruxin Wang, Hao Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26656v1 Announce Type: cross Abstract: Multi-object removal in 3D scenes is challenging due to severe occlusions, semantic entanglement, and the difficulty of maintaining geometric and multi-view consistency. Existing 3D Gaussian Splatting (3DGS) methods perform well for single-object edi...
147. PailitaoGR: Latent Think-with-Images for Generative Image Retrieval ​
Author: Xiaomeng Fan, Yueran Liu, Shengyu Zhou, Chenghan Fu, Wanxian Guan, Feng Li, Chuan Yu, Jian Xu, Bo Zheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR
arXiv:2608.26658v1 Announce Type: cross Abstract: Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain diverse information, includin...
148. Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference ​
Author: Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26674v1 Announce Type: cross Abstract: As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent's behavior consistently reflects the psychological and stylistic characteristics of a target p...
149. FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models ​
Author: Junyoung Lee, Sehyeon Park, Shinhyoung Jang, Seonha Ryu, Hojeong Kim, Hyunsei Lee, Il Hong Suh, Yeseong Kim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.26676v1 Announce Type: cross Abstract: Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy remain largely unchanged. In this work, we present a token-level analy...
150. AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability ​
Author: Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26713v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain...
151. LiveVVT: High-Fidelity Video Virtual Try-On in Real Time ​
Author: Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26714v1 Announce Type: cross Abstract: Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively en...
152. Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning ​
Author: Jintang Li, Yuhong Chen, Ruofan Wu, Binli Luo, Jiayi Ji, Hui Li, Rongrong Ji
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.26732v1 Announce Type: cross Abstract: Graph neural networks (GNNs) are typically conceptualized as message-passing neural networks, yet it remains unclear why neighborhood aggregation reliably outperforms node-wise multilayer perceptrons (MLPs). Despite its empirical success, this paradi...
153. Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction ​
Author: Yu-Lin Tsai, Yu-An Lu, Ci-Yang Tsai, Muxi Lyu, Raluca Ada Popa, Chia-Mu Yu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.26733v1 Announce Type: cross Abstract: Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted providers can keep these files secret while selling access to task results, making the skill itself a valuable target....
154. FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs ​
Author: Zeming Liu, Hang Lyu, Jingtao Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.26746v1 Announce Type: cross Abstract: Generated operational programs are often validated with either a few hand-written examples or exhaustive regression suites. The former can miss sparse boundary and interaction faults, while the latter can be unnecessarily expensive. We introduce Faul...
155. Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research ​
Author: Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.26753v1 Announce Type: cross Abstract: LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show tha...
156. Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory ​
Author: Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26807v1 Announce Type: cross Abstract: Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or engage in multi-turn clarification to elicit user preferences. However,...
157. Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay ​
Author: Prateek Chhikara
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26846v1 Announce Type: cross Abstract: Interactive language-model agents use confidence signals to decide whether to answer immediately, retrieve additional evidence (from memory or external knowledge), or defer. Yet confidence is usually evaluated in isolation, without measuring the traj...
158. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA ​
Author: Haowen Gu, Gensheng Pei, Zeren Sun, Mingwu Ren, Xiangbo Shu, Yazhou Yao, Fumin Shen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26848v1 Announce Type: cross Abstract: Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited annotated data and the high computational demands of existing large vision-language models. We propose MedFG-VQA,...
159. From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation ​
Author: Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26856v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features often lacks precise pixel-level grounding, thereby limiting clinical tr...
160. Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic ​
Author: Biao Yin, Abderrahmane Kasmi, Nadir Farhi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2608.26860v1 Announce Type: cross Abstract: Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety and traffic capacity. However, platoon control in real-world traffic is challenging due to uncertainty and heterogeneous driving behaviors. Reinforc...
161. PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact? ​
Author: Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.26882v1 Announce Type: cross Abstract: Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an autonomous agent convert ...
162. When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems ​
Author: Hanchong Chen, Xing Tang, Lingjie Li, Xiongfeng Shan, Xiuqiang He
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.26895v1 Announce Type: cross Abstract: Agentic recommender systems ground each decision of a large language model (LLM) in a persistent memory of the user, and in existing agents that memory is text: a narrative written and maintained by further LLM calls. Text limits this memory in two w...
163. Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS ​
Author: Kangmin Seo, Jae-Pil Heo
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26951v1 Announce Type: cross Abstract: Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images in one network execution, making 3D reconstruction increasingly accessible for casual captures. However, such captures frequently contain t...
164. Magnon-induced phononic Chern insulator ​
Author: Rui-Chang Shen, Yihao Yang, Haoran Xue
Published: 8/28/2026, 4:00:00 AM
Categories: cond-mat.mes-hall, cs.AI, physics.app-ph
arXiv:2608.27011v1 Announce Type: cross Abstract: High-frequency artificial phononic crystals offer a low-loss platform compatible with on-chip integration, yet realizing Chern phononic phases at GHz frequencies remains challenging. Here, we propose a magnon-induced phononic Chern insulator in a hon...
165. FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets ​
Author: Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne
Published: 8/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2608.27021v1 Announce Type: cross Abstract: LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench, a benchmark of 200...
166. Multi-Person Human Motion Forecasting in Complex Scenes ​
Author: Serdar Ozsoy, Lars Doorenbos, Juergen Gall
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27039v1 Announce Type: cross Abstract: Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framewo...
167. Performance Foundations of Parallel & Distributed Reasoning Language Models ​
Author: Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.PF
arXiv:2608.27046v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. The resulting recent Reasoning Language Models (RLMs) such as DeepSeek...
168. Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations ​
Author: Leon Ranke, Wolfgang H"ubner, Ronny Hug, Michael Arens, J"urgen Beyerer
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.27066v1 Announce Type: cross Abstract: Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection. This trade-off is commonly evaluate...
169. Emotional Preferences as Goal-Priority Regulation ​
Author: Shiqi Liu, Yihua Tan, Hu Fu, Guanyu Qi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27072v1 Announce Type: cross Abstract: A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. U...
170. Learning Transverse Momentum Distributions from Raw Scattering Events via Conditional Diffusion ​
Author: Jitao Xu, Christopher Cocuzza, Kevin Braga, Daniel Lersch, Nobuo Sato, Yaohang Li
Published: 8/28/2026, 4:00:00 AM
Categories: hep-ph, cs.AI
arXiv:2608.27077v1 Announce Type: cross Abstract: Extracting transverse momentum dependent parton distribution functions (TMD PDFs) from semi-inclusive deep inelastic scattering (SIDIS) data is a central goal of the nucleon structure program at Jefferson Lab and the future Electron-Ion Collider. Tra...
171. Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors ​
Author: Jitao Xu, Nobuo Sato, Yaohang Li
Published: 8/28/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2608.27080v1 Announce Type: cross Abstract: Many scientific and engineering applications require estimating unknown parameters from experimentally observable data -- an inverse problem that is inherently challenging due to nonlinearity, noise, and ill-posedness. In this paper, we propose an ac...
172. Active sensing to characterize the heterogeneity of plant stress ​
Author: Ayman Laaroussi, Peter Hanappe, David Colliaux
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.27088v1 Announce Type: cross Abstract: While most phenotyping platforms rely primarily on image-based measurements, advanced plant characterization requires the integration of active physiological sensing modali- ties such as chlorophyll fluorescence. We present an autonomous robotic plat...
173. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents ​
Author: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27141v1 Announce Type: cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The ...
174. ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification ​
Author: M. Middleton, H. Kayan, B. Sen Bhattacharya, T. Ali, E. Baikas, M. Vousden, C. Perera, O. Rhodes, E. Gheorghiu, M. A. Trefzer
Published: 8/28/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV
arXiv:2608.27150v1 Announce Type: cross Abstract: Object classification in event-based computer vision is a task that is attracting considerable research attention. Event-based object classification is a fundamental task in the fields of security and applied computer vision, which typically use sync...
175. When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue ​
Author: Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, eess.AS
arXiv:2608.27176v1 Announce Type: cross Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality sol...
176. LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration ​
Author: Matthew Youngman, Cristian Sestito, Themis Prodromakis
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.SE
arXiv:2608.27184v1 Announce Type: cross Abstract: Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively automate synthesis, optimisation, and verification. Large language models (LLMs) extend this trajectory by enabling ...
177. PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference ​
Author: Junjie Liu, Shengyuan Ye, Xu Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27206v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two fundamental limitations. First,...
178. STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration ​
Author: Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.27225v1 Announce Type: cross Abstract: Effective human-robot collaboration in industrial settings requires robots to understand human intentions and assist with task planning, reducing workload. Recent works have explored the use of Multi-modal Large Language Models (MM-LLMs) for task pla...
179. Compositional Online Learning for Semantic Data Processing Systems ​
Author: Pawe\l{} Liskowski, Fuheng Zhao, Benjamin Han, Anupam Datta, Dimitris Tsirogiannis
Published: 8/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.27244v1 Announce Type: cross Abstract: An LLM call in a semantic data processing system is expensive enough to dominate query cost, yet slow enough to hide a CPU-side learner's update behind its round-trip. In production, LLM compute accounts for $80-90%$ of query cost, and each call cos...
180. TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection ​
Author: Su Wang, Yaochen Li, Min Yang, Jiaohao Nie, Chang Liu, Yuehu Liu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO, cs.SY, eess.SY
arXiv:2608.27282v1 Announce Type: cross Abstract: Most single-stage 3D object detectors complete different tasks with the same extracted features. Nevertheless, it is impossible to project features into a common space that is adaptive for all the tasks. We present a novel task-aware deformable predi...
181. Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit ​
Author: Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.27309v1 Announce Type: cross Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating sca...
182. PAWBench: How Far Are We from Probabilistically Aligned World Modeling? ​
Author: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi'c, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27345v1 Announce Type: cross Abstract: Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible be...
183. RCMN: Understanding Misleadingness in Influential Public Discourse ​
Author: Peiling Yi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27358v1 Announce Type: cross Abstract: Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated. Yet less research has focused on how such misleadingness ...
184. KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations ​
Author: Chenchen Ge, Hanwen Shen, Bowen Jing, Jiyuan Cai, Xiaofeng Wang, Hongsen Lei, Weitao Zhou, Dandan Zhang, Haibao Yu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27365v1 Announce Type: cross Abstract: Physics-integrated 3D Gaussian representations now allow reconstructed deformable objects to be simulated and rendered under explicit material models. Existing pipelines, however, assume that material parameters are known or manually specified, limit...
185. Stageboost: Recommending Signals Based on Counterfactual Estimation ​
Author: Darpan Singhal, Matan Mandelbrod, Tal Franji, Manasa Kolla, Vipul Gaba, Yuri Brovman
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.27366v1 Announce Type: cross Abstract: Signals are short textual or visual snippets displayed on the eBay View-Item (VI) page, providing additional, contextual information for users about the viewed item. The aim of displaying these signals is to facilitate intelligent purchase and to inc...
186. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models ​
Author: Frederik Berenz
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27367v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-provisioned for simple tasks and under-provisioned for complex ones, with significant redundancy across attentio...
187. Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions ​
Author: Timothy Oladunni, Farouk Ganiyu-Adewumi
Published: 8/28/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.27392v1 Announce Type: cross Abstract: Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint performance does not establish whether other physiological properties of source contact photoplethysmography (PPG) remain preserved record...
188. LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics ​
Author: Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27395v1 Announce Type: cross Abstract: Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coup...
189. Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction ​
Author: Jin Mu, Guanhua Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27397v1 Announce Type: cross Abstract: Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guid...
190. How Language Models Organize and Structure Moral Knowledge ​
Author: Orion Reblitz-Richardson
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.27402v1 Announce Type: cross Abstract: How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between ...
191. CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators ​
Author: Kechen Liu, Ola Shorinwa
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.27406v1 Announce Type: cross Abstract: State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To brid...
192. Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners ​
Author: Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27424v1 Announce Type: cross Abstract: Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelS...
193. Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit ​
Author: Yisen Xi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.27427v1 Announce Type: cross Abstract: Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. W...
194. RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution ​
Author: Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27439v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming meth...
195. From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench ​
Author: Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2608.27442v1 Announce Type: cross Abstract: In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models ...
196. SWE-Prime: Fewer Trajectories, Better Performance ​
Author: Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2608.27449v1 Announce Type: cross Abstract: To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task succes...
197. Designing Cellular Manufacturing Systems in the Presence of Alternative Process Plans ​
Author: Md. Kutub Uddin, Md. Saiful Islam, Md Abrar Jahin, Md. Tanjid Hossen Irfan, Md. Saiful Islam Seam, M. F. Mridha
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2411.15361v4 Announce Type: replace Abstract: In the design of cellular manufacturing systems (CMS), numerous technological and managerial decisions must be made at both the design and operational stages. The first step in designing a CMS involves grouping parts and machines. In this paper, fo...
198. LLM-Powered Swarms: A New Frontier or a Conceptual Stretch? ​
Author: Muhammad Atta Ur Rahman, Melanie Schranz, Samira Hayat
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2506.14496v3 Announce Type: replace Abstract: Swarm intelligence describes how simple, decentralized agents can collectively produce complex behaviors. Recently, the concept of swarming has been extended to large language model (LLM)-powered systems, such as OpenAI's Swarm (OAS) framework, whe...
199. Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models ​
Author: Evangelos Georganas, Dhiraj Kalamkar, Alexander Heinecke, Pradeep Dubey
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.PF
arXiv:2508.06753v3 Announce Type: replace Abstract: The advent of ultra-low-bit LLM models, approaching the perplexity and task accuracy of their full precision counterparts, is ushering in a new era of LLM inference. While these advances promise models that are cost-effective regarding latency, mem...
200. Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning ​
Author: Yunxin Sun, Abulhair Saparov
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2509.03345v3 Announce Type: replace Abstract: Non-deductive reasoning, encompassing inductive and abductive reasoning, is essential in addressing complex real-world questions. One key feature of inductive and abductive reasoning is that there are many valid hypotheses; the simplest ones (those...
201. Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate ​
Author: Pratik S. Sachdeva, Tom van Nuenen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2510.10002v4 Announce Type: replace Abstract: As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes essential. Multi-agent debate has been studied as a mechanism to improve accuracy, but less is known about how de...
202. DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping ​
Author: Wei Fan, Wenlin Yao, Zheng Li, Feng Yao, Xin Liu, Liang Qiu, Qingyu Yin, Yangqiu Song, Bing Yin
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2510.12979v2 Announce Type: replace Abstract: Large language models (LLMs) augmented with multi-step reasoning and action generation abilities have shown promise in leveraging external tools to tackle complex tasks that require long-horizon planning. However, existing approaches either rely on...
203. Beyond Linearization: Attributed Table Graphs for Table Reasoning ​
Author: Yuxiang Wang, Junhao Gan, Shengxiang Gao, Shenghao Ye, Zhengyi Yang, Jianzhong Qi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.08444v2 Announce Type: replace Abstract: Table reasoning, a task to answer questions by reasoning over data presented in tables, is an important topic due to the prevalence of knowledge stored in tabular formats. Recent solutions use Large Language Models (LLMs) for their semantic underst...
204. DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning ​
Author: Yiming Yang, Zhuoyuan Li, Fanxiang Zeng, Hao Fu, Yue Liu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.08586v4 Announce Type: replace Abstract: Multi-agent LLM systems consistently outperform single-agent baselines, yet practitioners still cannot predict which design works for a new task or diagnose why one fails. We argue this gap persists largely because the field lacks a diagnostic fram...
205. Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences ​
Author: Hugo Math
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2603.16313v3 Announce Type: replace Abstract: Electronic control units (ECUs) embedded within modern vehicles generate a large number of asynchronous events known as diagnostic trouble codes (DTCs). These discrete events form complex temporal sequences that reflect the evolving health of the v...
206. Nomad: Autonomous Exploration and Discovery ​
Author: Bokang Jia, Samta Kamboj, Satheesh Katipomu, Seung Hun Han, Neha Sengupta, Andrew Jackson
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.29353v3 Announce Type: replace Abstract: We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of documents, databases, or other data sources, users rarely know the full set of questions, hypotheses, or connections that could be explored. As a ...
207. From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems ​
Author: Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du, Yi Zheng, Shouxi Ren, Deepayan Chakrabarti, Dengdu Jiang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG, cs.SI, q-fin.CP
arXiv:2605.23955v4 Announce Type: replace Abstract: Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical challenges ...
208. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation ​
Author: Kailin Lyu, Di Wu, Pengwei Zhang, Yuhang Zheng, Yingxin Lai, Long Xiao, Kangyi Wu, Pengna Li, Chen Gao, Lianyu Hu, Xiaobin Hu, Jie Hao, Ce Hao, Weihao Yuan, Shuicheng Yan
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.11637v4 Announce Type: replace Abstract: Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remai...
209. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents ​
Author: Ander Alvarez, Santhiya Rajan, Alessandro Genuardi, Oliver Wirjadi, Samuel Mugel, Rom'an Or'us
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA
arXiv:2606.18037v3 Announce Type: replace Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an ans...
210. Learning the ARTS of Search for Automated Discovery ​
Author: Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani, William Yang Wang, Xin Eric Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.21891v2 Announce Type: replace Abstract: Scientific discovery can be formulated as an iterative search process over the space of hypotheses and experiments. Contemporary methods navigate this space using heuristics such as MCTS. These algorithms conflate the merit of a hypothesis with the...
211. Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models ​
Author: MY Pitsane, Hope Mogale
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2607.04562v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism f...
212. Rethinking the Evaluation of Harness Evolution for Agents ​
Author: Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.12227v2 Announce Type: replace Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raise...
213. Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? ​
Author: Sergey Rodionov
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15439v2 Announce Type: replace Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, prompted simplification, and exact replay verification, leaving their individual contributions unclear. An executable world model is a persistent, agent-authored environment hypothesis...
214. Are the High-weight Neurons the Important Ones in Image Classification Neural Networks? ​
Author: Qitao Chen, Dongfu Yin, Xirui Yang, Zhaoye Li, Liang Xiao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.25529v2 Announce Type: replace Abstract: As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability. Yet existing work lacks clarity on the weight-importance relationship. We address this with a neuron importan...
215. Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations ​
Author: Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu, Chun Ouyang, Zhongxue Gan
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MM
arXiv:2608.03611v2 Announce Type: replace Abstract: Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-b...
216. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs ​
Author: Aditya Katkar, Om Karkele, Kartik Mandhane, Manisha More, Yash Kashid
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07167v2 Announce Type: replace Abstract: Autonomous LLM agents with tool execution capabilities introduce severe security risks through prompt injection, goal hijacking, and unauthorized action invocation. Existing guardrails rely on unverified, host local software filters system prompts,...
217. Blast Radius ​
Author: MY Pitsane, Hope Mogale
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07440v3 Announce Type: replace Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables rev...
218. FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training ​
Author: Josef Chen (Independent Researcher), Erim Hayretci (Imperial College London)
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.SE
arXiv:2608.20574v2 Announce Type: replace Abstract: Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment. Each task asks for...
219. Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization ​
Author: Praphul Singh, Shanu Kumar, Akshat Agarwal
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20768v3 Announce Type: replace Abstract: Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexami...
220. SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing ​
Author: Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.22018v3 Announce Type: replace Abstract: Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, supporting arguments, and target-level labels. Documents with multiple targets, conflicting local readings, or cul...
221. ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation ​
Author: Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2608.22559v2 Announce Type: replace Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require black-box LLM judges, and typically assume criteria aggregate inde...
222. Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents ​
Author: Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.22963v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated text. Over long trajectories, this text can domi...
223. From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation ​
Author: Xiangxin Zhang, Zhanwei Zhang, Zhihang Fu, Binbin Lin, Wenxiao Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.23045v2 Announce Type: replace Abstract: Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: once an agent has produced a query, plan, or intermediate conclusion, it becomes less objective when later judgi...
224. Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement ​
Author: Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li, Xin Cong, Yuzhuo Bai, Kangyang Luo, Maosong Sun
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.23098v2 Announce Type: replace Abstract: Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry systems follow a one-shot generation paradigm, which reduces users to prompt providers and weaken...
225. Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping ​
Author: Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui, Zujie Wen, Zhiqiang Zhang, Jun Zhou
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.24135v2 Announce Type: replace Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) is pivotal for enhancing LLM code generation, yet its efficacy is often hindered by insufficient test case coverage, leading to reward hacking and policy degradation. To address this, we propose...
226. RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards ​
Author: Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.24275v2 Announce Type: replace Abstract: Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to u...
227. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use ​
Author: Rongfeng Guo, Yinxuan Huang, Yusen Wu, Maoqing Zhong, Yunlu Chen, Meng Tang, Teng Long, Vincent Tao Hu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.24368v2 Announce Type: replace Abstract: Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each action remains consistent with it. However, direct function-calling and ReAct-style policies learn state tracking and action generation within th...
228. Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2 ​
Author: Xuchen Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.24893v2 Announce Type: replace Abstract: In competitive first-person shooter (FPS) games such as Counter-Strike 2 (CS2), account-integrity review often asks whether an account's recent behavior remains consistent with its historical operator. This consistency question arises in cases such...
229. Recurrent Reinforcement Learning with Memoroids ​
Author: Steven Morad, Chris Lu, Ryan Kortvelesy, Stephan Liwicki, Jakob Foerster, Amanda Prorok
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2402.09900v4 Announce Type: replace-cross Abstract: Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. Neither model scales particularly well to long sequences,...
230. CollaFuse: Collaborative Diffusion Models ​
Author: Simeon Allmendinger, Domenique Zipperling, Lukas Struppek, Niklas K"uhl
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2406.14429v5 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic images. However, the application of diffusion models poses numerous challenges, particularly concerning...
231. The BS-meter: Detecting Politics and Labour through ChatGPT's Language ​
Author: Alessandro Trevisan, Harry Giddens, Sarah Dillon, Alan F. Blackwell
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2411.15129v3 Announce Type: replace-cross Abstract: What can we learn about language from studying how it is used by ChatGPT and other large language model (LLM)-based chatbots? In this paper, we analyse the distinctive character of language generated by ChatGPT, in relation to questions raise...
232. Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling ​
Author: Hengran Zhang, Keping Bi, Jiafeng Guo, Xiaojie Sun, Shihao Liu, Daiting Shi, Dawei Yin, Xueqi Cheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2504.05216v5 Announce Type: replace-cross Abstract: Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for downstream tasks such as re-ranking and augmenting generation. Recently, large language models (LLMs) have demonstrated impressive semantic understandin...
233. Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study ​
Author: Jiawei Zhou, Kritika Venkatachalam, Minje Choi, Koustuv Saha, Munmun De Choudhury
Published: 8/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2505.08143v2 Announce Type: replace-cross Abstract: With the wide adoption of large language models (LLMs) in information assistance, it is essential to examine their alignment with human communication styles and values. We situate this study within health fact-checking, where effective commun...
234. Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models ​
Author: Keunwoo Peter Yu, Joyce Chai
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2505.11326v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-time interactive environments impose new demands on VLMs, requiring them to generate utterances...
235. HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement ​
Author: Jilin Hu, Jianyu Zhang, Yongwang Zhao, Talia Ringer
Published: 8/28/2026, 4:00:00 AM
Categories: cs.FL, cs.AI, cs.SE
arXiv:2505.15740v2 Announce Type: replace-cross Abstract: Formal methods play a crucial role in ensuring the reliability of critical systems through rigorous mathematical verification. However, their adoption remains limited due to the labor-intensive nature of manual proof construction. Recent adva...
236. From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning ​
Author: Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2505.22203v3 Announce Type: replace-cross Abstract: Trustworthy verifiers are essential for the success of reinforcement learning with verifiable reward (RLVR), which is the core methodology behind various large reasoning models such as DeepSeek-R1. In complex domains like mathematical reasoni...
237. Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation ​
Author: Peibo Li, Shuang Ao, Hao Xue, Yang Song, Maarten de Rijke, Johan Barth'elemy, Tomasz Bednarz, Flora D. Salim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2506.21599v5 Announce Type: replace-cross Abstract: Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task faces two fundamental challenges: (i) although existing methods produce semantic IDs that incorporate semantic information, their topology-blind i...
238. Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics ​
Author: Chenyang Cao, Miguel Rogel-Garc'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO
arXiv:2507.00611v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback...
239. Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling ​
Author: Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Got=o, Tomonori Hosokawa, Makoto Nishimura, Y=osuke Sat=o, Izumi Sezai, Tomohiro Inoue
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2507.02950v4 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark study compared four configured LLM evaluator systems (GPT-5.5...
240. AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air ​
Author: Shiyi Yang, Xiaoxue Yu, Rongpeng Li, Jianhang Zhu, Zhifeng Zhao, Honggang Zhang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2507.11515v2 Announce Type: replace-cross Abstract: Operating Large Language Models (LLMs) on edge devices is increasingly challenged by limited communication bandwidth and strained computational and memory costs. Thus, cloud-assisted remote fine-tuning becomes indispensable. Nevertheless, exi...
241. Toward a New Science of AI as Cognitive Infrastructure ​
Author: Giuseppe Riva
Published: 8/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2507.22893v3 Announce Type: replace-cross Abstract: Contemporary human-AI interaction research overlooks how AI systems fundamentally reshape human cognition pre-consciously, a critical blind spot for understanding distributed cognition. This paper introduces "Cognitive Infrastructure Studies"...
242. Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics ​
Author: Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.11017v3 Announce Type: replace-cross Abstract: Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to stu...
243. Recurrence Meets Transformers for Universal Multimodal Retrieval ​
Author: Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.MM
arXiv:2509.08897v3 Announce Type: replace-cross Abstract: With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks have emerged. Existing methods predominantly rely on task-specific fine-tuning of vision-language models ...
244. GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts ​
Author: Fan Yuan, Yuchen Yan, Yifan Jiang, Haoran Zhao, Tao Feng, Jinyan Chen, Yanwei Lou, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2509.25160v2 Announce Type: replace-cross Abstract: Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicitly symbolic visual inputs. It remains unclear whether VLMs can reason mathematically when information m...
245. Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras ​
Author: Han Zhang, Lalithkumar Seenivasan, Jose L. Porras, Roger D. Soberanis-Mukul, Hao Ding, Hongchao Shu, Benjamin D. Killeen, Ankita Ghosh, Lonny Yarmus, Jeffrey K. Jopling, Masaru Ishii, Angela C. Argento, Mathias Unberath
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2510.04802v2 Announce Type: replace-cross Abstract: Observing surgical practice has historically relied on fixed vantage points or recollections, leaving the egocentric perspectives that shape clinical decisions undocumented. Ambient fixed cameras capture the operating room (OR) at room scale ...
246. MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning ​
Author: Nian Ran, Zhongzheng Li, Yue Wang, Qingsong Ran, Xiaoyuan Zhang, Shikun Feng, Richard Allmendinger, Xiaoguang Zhao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.06270v2 Announce Type: replace-cross Abstract: Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. Traditional evolutionary algorithms often get trapped in local optima, while expert...
247. LLM-Specific Utility for Retrieval-Augmented Generation ​
Author: Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2510.11358v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved passages are useful for a large language model (LLM) to generate correct and complete answers. We argue...
248. MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ​
Author: ChangSu Choi, Hoyun Song, Dongyeon Kim, Minkyung Cho, WooHyeon Jung, Sunjin Park, NohHyeob Bae, Seona Yu, KyungTae Lim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.18383v4 Announce Type: replace-cross Abstract: Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical application. The predominant approach, supervised fine-tuning (SFT), is an off-policy distillation method ...
249. The Principles of Diffusion Models ​
Author: Chieh-Hsin Lai, Yang Song, Dongjun Kim, Yuki Mitsufuji, Stefano Ermon
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GR
arXiv:2510.21890v3 Announce Type: replace-cross Abstract: This book presents the core principles that have guided the development of diffusion models, tracing their origins and showing how diverse formulations arise from shared mathematical ideas. Diffusion modeling starts by defining a forward proc...
250. What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs ​
Author: Aakriti Shah, Yifan Hu, Thai Le
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.25732v2 Announce Type: replace-cross Abstract: Unlearning in large language models (LLMs) is usually evaluated as whether an "unlearned" fact can be recovered. We instead ask whether a fact's structural entanglement with the rest of a model's knowledge predicts whether it leaks after unle...
251. Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation ​
Author: Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou, Haikuo Peng, Xueqi Li, Chun Yu, Junliang Xing
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.12631v3 Announce Type: replace-cross Abstract: While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to sub...
252. Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study ​
Author: Chorok Lee
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2601.00908v2 Announce Type: replace-cross Abstract: Conformal prediction provides distribution-free coverage guarantees, but these degrade under distribution shift - and practitioners lack tools to anticipate which deployed models will fail before observing test data. We propose SHapley Additi...
253. CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models ​
Author: Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.MM
arXiv:2601.04778v2 Announce Type: replace-cross Abstract: Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filtering or random video per...
254. Subspace Alignment for Vision-Language Model Test-time Adaptation ​
Author: Zhichen Zeng, Wenxuan Bao, Xiao Lin, Ruizhong Qiu, Tianxin Wei, Xuying Ning, Yuchen Yan, Chen Luo, Monica Xiao Cheng, Jingrui He, Hanghang Tong
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2601.08139v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptation (TTA) emerges as a predominant strategy to adapt VLMs to unlabeled test data on the fly. However, e...
255. LoRA as Oracle ​
Author: Marco Arazzi, Antonino Nocera
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2601.11207v2 Announce Type: replace-cross Abstract: Practitioners increasingly deploy neural networks they did not train, and must audit them after the fact for hidden backdoors, without the training pipeline, the poisoned data, or knowledge of any trigger. We introduce a low-rank auditing len...
256. Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content ​
Author: Parth Bhalerao, Diola Dsouza, Ruiwen Guan, Oana Ignat
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.17173v2 Announce Type: replace-cross Abstract: Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career guidance-require mentorship: responses that provide reflection and guidance. Existing QA benchmarks rarel...
257. A Very Big Video Reasoning Suite ​
Author: Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thadd"aus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Jiachen Li, Hanwen Xing, Tianqi Zhao, Fengyuan Yu, Weihang Xiao, Yizheng Jiao, Jianheng Hou, Danyang Zhang, Pengcheng Xu, Boyang Zhong, Zehong Zhao, Gaoyun Fang, John Kitaoka, Yile Xu, Hua Xu, Kenton Blacutt, Tin Nguyen, Siyuan Song, Haoran Sun, Shaoyue Wen, Linyang He, Runming Wang, Yanzhi Wang, Mengyue Yang, Ziqiao Ma, Rapha"el Milli`ere, Freda Shi, Nuno Vasconcelos, Daniel Khashabi, Alan Yuille, Yilun Du, Ziming Liu, Bo Li, Dahua Lin, Ziwei Liu, Vikash Kumar, Yijiang Li, Lei Yang, Zhongang Cai, Hokin Deng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM, cs.RO
arXiv:2602.20159v3 Announce Type: replace-cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can nat...
258. SynthCharge: An Electric Vehicle Routing Instance Generator with Feasibility Screening to Enable Learning-Based Optimization and Benchmarking ​
Author: Mertcan Daysalilar, Fuat Uyguroglu, Gabriel Nicolosi, Adam Meyers
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.03230v2 Announce Type: replace-cross Abstract: The electric vehicle routing problem with time windows (EVRPTW) extends the classical VRPTW by introducing battery capacity constraints and charging station decisions. Existing benchmark datasets are often static and lack verifiable feasibili...
259. Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization ​
Author: Francesco Pio Monaco, Elia Cunegatti, Flavio Vella, Giovanni Iacca
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.16105v4 Announce Type: replace-cross Abstract: Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance. While several compression approaches have been proposed, less emphasis has been placed on selecting...
260. How LLMs Distort Our Written Language ​
Author: Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.18161v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate that LLMs not only alter the voice and tone of human writing but also consistently alter the intended mea...
261. High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking ​
Author: Peipeng Yu, Jinfeng Xie, Chengfu Ou, Xiaoyu Zhou, Jianwei Fei, Yunshu Dai, Zhihua Xia, Chip Hong Chang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.23940v2 Announce Type: replace-cross Abstract: The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Existing versatile watermarking systems typically rely on embedding explicit localization payloads...
262. Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation ​
Author: Daiwei Chen, Zhoutong Fu, Chengming Jiang, Haichao Zhang, Ran Zhou, Tan Wang, Chunnan Yao, Guoyao Li, Rui Cai, Yihan Cao, Ruijie Jiang, Fedor Borisyuk, Jianqiang Shen, Jingwei Wu, Ramya Korlakai Vinayak
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.02324v2 Announce Type: replace-cross Abstract: Language models (LMs) are increasingly extended with new learnable vocabulary tokens for domain-specific tasks, such as Semantic-ID tokens in generative recommendation. The standard practice initializes these new tokens as the mean of existin...
263. A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting ​
Author: Junlin Li, Xinhao Song, Siqi Wang, Haibin Huang, Yili Zhao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.GR, cs.AI, cs.CV
arXiv:2604.13427v2 Announce Type: replace-cross Abstract: Text-driven motion editing and intra-structural retargeting, where skeletons share topology but may differ in bone lengths and rest pose, are traditionally handled by fragmented pipelines with incompatible inputs and representations: editing ...
264. CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations ​
Author: Xiping Li, Aier Yang, Jianghong Ma, Kangzhe Liu, Shanshan Feng, Haijun Zhang, Yi Zhao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2604.14586v4 Announce Type: replace-cross Abstract: The rapid expansion of gaming industry requires advanced recommender systems tailored to its dynamic landscape. Existing Graph Neural Network (GNN)-based methods primarily prioritize accuracy over diversity, overlooking their inherent trade-o...
265. Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning? ​
Author: Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.14892v4 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM Jury, composed of three frontier AI models, for scoring 3...
266. MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation ​
Author: Markus Knauer, Edoardo Fiorini, Maximilian M"uhlbauer, Stefan Schneyer, Promwat Angsuratanawech, Florian Samuel Lay, Timo Bachmann, Samuel Bustamante, Korbinian Nottensteiner, Freek Stulp, Alin Albu-Sch"affer, Jo~ao Silv'erio, Thomas Eiband
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.HC, cs.LG
arXiv:2604.20468v3 Announce Type: replace-cross Abstract: Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different adaptations benefit from different interaction modalities. We present an interac...
267. MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction ​
Author: Aladin Djuhera, Haris Gacanin, Holger Boche
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, eess.SP, math.IT
arXiv:2604.21957v2 Announce Type: replace-cross Abstract: Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel state prediction (CSP) performance by capturing long-range temporal dependencies across channel state info...
268. Cartan flow matching ​
Author: Francesco Ruscelli, Ferdinando Zanchetta, Rita Fioresi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.03588v2 Announce Type: replace-cross Abstract: We introduce Cartan flow matching, a general framework for training flow matching models on Riemannian symmetric spaces, i.e. Riemannian manifolds with the property that at any point there exists a geodesic symmetry. This is a large class of ...
269. MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection ​
Author: Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.04180v2 Announce Type: replace-cross Abstract: Large language models fabricate in medicine, producing fluent statements that are factually wrong, so reliable fabrication detection is a prerequisite for clinical deployment. Reported progress on this task is inflated by two evaluation artif...
270. No Plan, Yet Human: A Reactive Robotics Model Predicts Human Planning Failures on a Clinical Task ​
Author: Michael Migacev, Vito Mengers, Antonia K"ongeter, Oliver Brock
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2605.16514v2 Announce Type: replace-cross Abstract: Understanding why some sequential planning problems are harder than others requires models that go beyond average performance. They should capture the specific pattern of which problems are hard, and ideally fail in the same way people do whe...
271. HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents ​
Author: Woongyeong Yeo, Yumin Choi, Taekyung Ki, Sung Ju Hwang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2605.17873v2 Announce Type: replace-cross Abstract: Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods a...
272. GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets ​
Author: Zhangyang Yao, Haiyan Zhao, Haoyu Wang, Xu Han
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.18475v2 Announce Type: replace-cross Abstract: Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive modules. However, automating this allocation at LLM scale faces a unique combination of constraints: le...
273. A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery ​
Author: Sarmad Khan, Basim Azam, Arslan Shaukat
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.20445v2 Announce Type: replace-cross Abstract: COVID-19 was a significant challenge that led to the loss of numerous lives daily. Not only a certain country was involved in this outbreak, but even the world has suffered because of the coronavirus. Imaging techniques using computed tomogra...
274. Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated Image Segmentation Architectures ​
Author: Sarmad Khan, Basim Azam, Arslan Shaukat
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.20459v2 Announce Type: replace-cross Abstract: In recent years, there has been a notable increase in the level of attention that is given to algorithms based on deep learning in the context of medical image segmentation. Nevertheless, the reliability of the field has been hindered due to ...
275. MIMO: Multilingual Information Retrieval via Monolingual Objectives ​
Author: Youngjoon Jang, Seongtae Hong, Heuiseok Lim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2605.31171v2 Announce Type: replace-cross Abstract: Multilingual Information Retrieval (MLIR) reflects real-world search environments in which queries and relevant documents may appear in different languages within a mixed-language corpus. However, existing embedding models are primarily optim...
276. Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains ​
Author: Jiawei Guo, Donglei Yu, Yu Chen, Xiang Wang, Shuai Li, Xinpei Zhao, Huaxing Liu, Qinghao Wang, Minpeng Liao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.02357v2 Announce Type: replace-cross Abstract: Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a tool-call trace alone does not show whether the tool supplied a...
277. LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling ​
Author: Wenkai Chen, Tianshu Li, Wenyong Huang, Yichun Yin, Lifeng Shang, Chengwei Qin
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.04438v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. However, mainstream looped architectures rely on dense backbones that couple parameter count with per-tok...
278. Summarization is Not Dead Yet ​
Author: Dongqi Liu, Chenxi Whitehouse, Zheng Zhao, Zhuchen Cao, Jian Li, Yabiao Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.08000v2 Announce Type: replace-cross Abstract: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem. We re-examine this n...
279. Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries ​
Author: Federico Bianchi, Yongchan Kwon, Aneesh Pappu, James Zou
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.10402v2 Announce Type: replace-cross Abstract: Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's ideas over long time horizons. Recent AI systems have shown that language-model-based agents can make me...
280. Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot ​
Author: Yuyang Dai, Yushun Dong
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2606.15810v2 Announce Type: replace-cross Abstract: Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a defense that redirects ex...
281. SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval ​
Author: Youngjoon Jang, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2606.18801v2 Announce Type: replace-cross Abstract: With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant documents from multil...
282. PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement ​
Author: Xianren Zhang, Delvin Ce Zhang, Dongwon Lee, Suhang Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV
arXiv:2607.02897v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns. Machine unlearning offers a way to remove such private knowledge without retraining fr...
283. Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift ​
Author: Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.10358v3 Announce Type: replace-cross Abstract: Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and can...
284. Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data ​
Author: Nursultan Askarbekuly, Mohamad Al Mdfaa, Ahmed Helaly, Gonzalo Ferrer, Manuel Mazzara
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.18064v2 Announce Type: replace-cross Abstract: Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent receives a dataset, an evaluation script, and one editable file, and iterates without supervision: mod...
285. Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating ​
Author: Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR, q-bio.QM
arXiv:2607.19020v3 Announce Type: replace-cross Abstract: Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as accuracy: once retraining touches every parameter, no one can say afterwards where the update acted. We p...
286. ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour ​
Author: Jianhang Xie, Sicheng Tan, Vishnu Naresh Boddeti, Zhichao Lu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2607.23478v2 Announce Type: replace-cross Abstract: Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. Its non-linear operations, such as softmax, normalization, and activation, must...
287. TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation ​
Author: Susil Kumar Mohanty, Rohit Patel, Kosuru Yuvaraj, Jeenal Chaudhary, Disha Singhania
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2607.23838v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) grounds LLM answers in query-time retrieved documents, so reliability depends on what the retriever returns. PoisonedRAG (Zou et al., USENIX Security'25) showed five crafted documents mislead an undefended...
288. REPREC: Representation Driven Parameter-Efficient Recommendation System ​
Author: Harshini Kavuru, Dwipam Katariya, Giri Iyengar, Pranab Mohanty, Kalanand Mishra, Raghu Machiraju
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2607.24845v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been applied to sequential recommendation by formulating it as a natural language task. Previous work has improved personalization by incorporating collaborative and sequential signals through input condition...
289. Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities ​
Author: Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.27747v2 Announce Type: replace-cross Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluatio...
290. When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs ​
Author: Jiaming Cheng, Subhransu Das, Rajiv Ramnath
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.04893v2 Announce Type: replace-cross Abstract: Multi-agent LLM systems relay key-value caches instead of text and credit their gains to exchanged "latent thoughts". That credit is a claim about which example's cache is relayed, not merely that one is. We audit it causally in released syst...
291. UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on ​
Author: Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.05745v2 Announce Type: replace-cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for huma...
292. ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB ​
Author: Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.DC
arXiv:2608.07945v2 Announce Type: replace-cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our...
293. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes ​
Author: Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10954v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on ...
294. REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ​
Author: Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11698v3 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imit...
295. A 12-CNOT Double Qubit Excitation Gate ​
Author: Irfansha Shaik
Published: 8/28/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2608.11733v3 Announce Type: replace-cross Abstract: Effective implementation of high-level quantum gates is essential for practical quantum computing. In this work, we presented, to the best of our knowledge, the first reported 12-CNOT decomposition of the double qubit excitation operator. We ...
296. M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation ​
Author: Jing Zhu, Ye Wang, Fumin Wang
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12196v2 Announce Type: replace-cross Abstract: Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathe...
297. MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification ​
Author: Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2608.13463v2 Announce Type: replace-cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image Classification with LLM...
298. Pre-training Visual Dexterity in Simulation ​
Author: Sarthak Kamat, Adam Rashid, Satvik Sharma, Aseem Doriwala, Chelsea Finn, Phillip Isola, C. Karen Liu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.15917v2 Announce Type: replace-cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers. Dexterous, multi-fingered hands remain co...
299. X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization ​
Author: Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2608.16658v2 Announce Type: replace-cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented loc...
300. Complexity Induction: Compositional Generalization via Structured Training Distortion ​
Author: Aleksandr V. Abramov
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.21464v2 Announce Type: replace-cross Abstract: We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geome...
301. Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules ​
Author: Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.22642v2 Announce Type: replace-cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. To address these challenges, we presen...
302. Language Chain in Alignment: Cross-lingual Ranking Preference Optimization ​
Author: Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.23149v2 Announce Type: replace-cross Abstract: The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-lingual Ranking Preference Optimization~(CR...
303. The Limits of Automatic Evaluation of Creativity in Large Language Models ​
Author: Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.23705v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investiga...
304. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$ ​
Author: Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen
Published: 8/28/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.23758v2 Announce Type: replace-cross Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typically focus on a single au...
305. When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk ​
Author: Liliana Santos-Deonizio, James Malamut, Ram'on Antonio Mart'inez, Dorottya Demszky
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.23780v2 Announce Type: replace-cross Abstract: LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include...
306. CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery ​
Author: Mahir Shahriar Tamim, Sharjil Khan, Md. Samiul Alim, Tanvir Ahmed Khan, Shafin Rahman, Nabeel Mohammed
Published: 8/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.24947v2 Announce Type: replace-cross Abstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii)...
307. Unsupervised Post-Training of Foundation Models: A Survey ​
Author: Yijie Xu, Qianyi Cai, Huizai Yao, Yili Wang, Tianfu Wang, Cehao Yang, Xingbo Yao, Zhiyu Guo, Aiwei Liu, Xuming Hu, Weiyu Guo, Hui Xiong
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.MM
arXiv:2608.24982v2 Announce Type: replace-cross Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is deri...
308. DataKernelBench: Can LLMs Optimize Database Queries on GPUs? ​
Author: Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie, Andrea Giovannini, Katja Hose
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB, cs.LG, cs.PL
arXiv:2608.25061v2 Announce Type: replace-cross Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement...
309. MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration ​
Author: Miseon Yu, Jaehoon Choi, Younghan Lee, Yunheung Paek
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA
arXiv:2608.25457v2 Announce Type: replace-cross Abstract: Despite their strong ability to generate code, large language models often fail to produce secure code, as their outputs frequently contain security vulnerabilities. Secure code generation is inherently challenging because it requires solving...
310. 4DStreamCtrl: Interactive Video Generation with Online 4D Control ​
Author: Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.25479v2 Announce Type: replace-cross Abstract: Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only par...
311. When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory ​
Author: Kazuki Nakayashiki
Published: 8/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.25553v2 Announce Type: replace-cross Abstract: An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer authoritative record. Under a scarce verification budget, does the agent recover the withdrawal, and if n...
312. Learning New Facts with QLoRA: An Acquisition-Retention Frontier ​
Author: Estelle Zheng, S'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara
Published: 8/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.25677v2 Announce Type: replace-cross Abstract: Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters. We show that this assumption depends strongly on adapter capacity. We study factual acquisition in a co...