arXiv cs.AI - 2026-08-20 ​
267 items collected.
1. Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions ​
Author: Matthew Riemer, Tommaso Tosato, Amin Memarian, Maximilian Puelma Touzel, Glen Berseth, Irina Rish, Guillaume Dumas
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is...
2. Position: Profiling Game Worlds by Transition Complexity ​
Author: Lele Cao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18079v1 Announce Type: new Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying transition prediction problem is at the declared interface (pixels/tokens/latents with finite history). ...
3. Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges ​
Author: Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18080v1 Announce Type: new Abstract: We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate fin...
4. Position: Behavioral Systems Require Behavioral Tests ​
Author: Manuel Cherep, Nikhil Singh, Pattie Maes
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18081v1 Announce Type: new Abstract: Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time. Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral...
5. Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models ​
Author: Sungwon Chae, Keonwoo Kim, Hoki Kim, Jaeyeon Ju, Sangchul Park
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18086v1 Announce Type: new Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. Although model cards have been widely adopted as transparency artifacts in model repositories, existing fra...
6. A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring ​
Author: Seyma Yaman Kayadibi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.18088v1 Announce Type: new Abstract: Drone propeller faults can create safety and reliability risks when their effects are distributed across multiple flight-log channels rather than appearing as a single diagnostic signal. This paper proposes a Metamorphic Artificial Age Score (AAS) deci...
7. Position: Multi-Agent Systems Should Prioritize Concurrency Control ​
Author: Xin Yang, Letian Li, Zimo Ji, Terry Jingchen Zhang, Wenyuan Jiang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability. This position paper argues that many MAS failures are fundamentally concurrency control problems: agents concurrently read and write shared...
8. FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management ​
Author: Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, q-fin.PM
arXiv:2608.18099v1 Announce Type: new Abstract: Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce auditable structu...
9. Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective ​
Author: Yuanyuan Xu, Wenjie Zhang, Yin Chen, Xuemin Lin, Ying Zhang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18104v1 Announce Type: new Abstract: Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist across interactions, maintain memories, use tools, acquire skills, refine workflows, and coordinate with other agents. These capabilities make agent st...
10. Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions ​
Author: AKM Bahalul Haque, Al Amin Islam Ridoy, Mohammad Rayhan, Ivan Porres
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18110v1 Announce Type: new Abstract: Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapid transformation across various domains.This rapid advancement and the potential to revolutionize various domains...
11. Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry ​
Author: Hsien Xin Peng, Anthony Kim, Alvin Li, Calvin Supasanya, Shivank Garg, Kevin Zhu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18111v1 Announce Type: new Abstract: Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry problem and drawing...
12. Position: AI Leaderboards Are Underserving the Global South: A Case Study from India ​
Author: Sourav Banerjee, Saikat Saha
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.18117v1 Announce Type: new Abstract: This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-...
13. Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs ​
Author: Namya Bhatnagar
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY
arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may pr...
14. Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis ​
Author: Kim-Anh Nguyen, Huy Hoang Le, Ba Tu Phung
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18133v1 Announce Type: new Abstract: Reliable transformer fault diagnosis is essential for maintaining power system stability. The IEEE Key Gas Method (KGM), a widely utilized approach in Dissolved Gas Analysis (DGA), exhibits limitations in addressing ambiguous data and ensuring high dia...
15. Improving Rural Medication Safety with AI: A Scoping Review ​
Author: Jeong-ah Kim, Muhammad Ashad Kabir, Daniel Terry, Maryam Rouhi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18135v1 Announce Type: new Abstract: Introduction: Medication errors (MEs) represent a significant threat to global healthcare systems, contributing to patient harm. Introducing artificial intelligence (AI) in rural healthcare enhances patient safety. The aim is to explore the application...
16. FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud ​
Author: Dheeraj Mohandas Pai, Lu Xian
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18136v1 Announce Type: new Abstract: Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can ...
17. Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu ​
Author: Toneema Zubair, Muhammad Junaid Asif, Faisal Kamiran, Hafiz Hassan Saeed, Rana Fayyaz Ahmad
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.18142v1 Announce Type: new Abstract: It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a challenge is Roman Urdu which i...
18. RDFdL: Integrating RDF with Differential Dynamic Logic ​
Author: Yuyang Li, Lukas Kubelka, Julia Butte, Tobias K"afer
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.FL, cs.LO
arXiv:2608.18165v1 Announce Type: new Abstract: Knowledge graphs modeled in RDF are powerful for describing static knowledge, but they cannot capture or reason about the dynamic behavior of physical systems, e.g., systems described by differential equations, which is a critical gap for AI-driven cyb...
19. Adversarial Review: Structured Disagreement for Grounded Agentic Code Review ​
Author: Eric S. Qiu, Joyce Gill
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.18167v1 Announce Type: new Abstract: Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level coding tasks. Recent alternatives treat agents as passive tools (subagents), yet this removes the benefits of agent in...
20. Looped Language Models Improve Compositional Tool Calling ​
Author: Andrei Cristian Popescu, Haitz S'aez de Oc'ariz Borde, Pietro Li`o
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18171v1 Announce Type: new Abstract: Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. We study this question in compositional tool-calling settings, where models must coordinate multiple API c...
21. On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices ​
Author: Costin B\u{a}dic\u{a}, Amelia B\u{a}dic\u{a}
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.DM, math.CO
arXiv:2608.18194v1 Announce Type: new Abstract: This paper presents new theoretical results on generalizing the Jaccard distance for lattices and real valuations. We demonstrate that when the valuation is strictly positive, monotone, and modular, the Jaccard distance satisfies the triangle inequalit...
22. GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks ​
Author: Arefin Amin, Labiba Faiza Karim, M. Monir Uddin
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, q-bio.QM
arXiv:2608.18238v1 Announce Type: new Abstract: Genomic analysis on viruses such as SARS-CoV-2 variants: Beta, Gamma, Delta, and Omicron is heavily dominated by classical bioinformatics methods, including Sequence Alignment, Phylogenetic Analysis, and Mutation Frequency Statistics. These approaches ...
23. Redakto - The Incognito Tab for LLMs ​
Author: Saurav Kumar Saha, Tom R"ohr, Felix Bie{\ss}mann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.IR, cs.LG
arXiv:2608.18260v1 Announce Type: new Abstract: Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in the context of LLMs or Artificial Intelligence (AI) in general is to ensure privacy when using them, meaning that personally identifiable informatio...
24. Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study ​
Author: Shriniwas Ramesh Suram
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts si...
25. Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application ​
Author: Elias Schubert, Felix Bie{\ss}mann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.IR, cs.LG
arXiv:2608.18289v1 Announce Type: new Abstract: The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Opt...
26. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations ​
Author: Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18300v1 Announce Type: new Abstract: LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. However, most work treats...
27. SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition ​
Author: Dae Lee, Mihai Delgeanu, Adel Youssef
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose...
28. ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents ​
Author: Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2608.18307v1 Announce Type: new Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a button set) that are s...
29. Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair ​
Author: Jesus Salas
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18324v1 Announce Type: new Abstract: Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether these records can supervise bounded models, consolidating occasional or expensive cap...
30. Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme ​
Author: Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.18336v1 Announce Type: new Abstract: When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an exam...
31. A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations ​
Author: Hasan Najib Mahmud (Colorado State University), Shreya Gupta (Microsoft), Isha Chaudhary (University of Illinois Urbana-Champaign), Nathaniel Enis (Colorado State University), Ravi Mangal (Colorado State University), Gagandeep Singh (University of Illinois Urbana-Champaign), Corina Pasareanu (Carnegie Mellon University)
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18389v1 Announce Type: new Abstract: AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether coding agents that repair repository-level issues remain reliable when the ...
32. When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification ​
Author: Saba A. Farahani, Hung Cao, Amir M. Rahmani
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18397v1 Announce Type: new Abstract: Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual. On WESAD, a Random Forest reaches 93.0% mean accuracy yet yields F1 = 0 for Subject 14, whose cross-signal coupling weakens near st...
33. Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions ​
Author: Shrenil Shaun Sharma, Avi Sharma
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a significant challenge for language models, requiring them to identify feasible solutions within exponentially large search spaces while satisfying complex constraints. This challenge is especially pronounced in resource...
34. FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents ​
Author: Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18423v1 Announce Type: new Abstract: Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Be...
35. UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval ​
Author: Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18504v1 Announce Type: new Abstract: Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidd...
36. Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval ​
Author: Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18521v1 Announce Type: new Abstract: Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimizatio...
37. Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson ​
Author: Tanay Chowdhury, Saeideh Shahrokh Esfahani
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18531v1 Announce Type: new Abstract: Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds and cost that scales linearly with traffic. We separate generation from ...
38. FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems ​
Author: Pratik Ghawate
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.18534v1 Announce Type: new Abstract: Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed a...
39. Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement ​
Author: Mandar Kulkarni, Pooja A., Samir Shah
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18543v1 Announce Type: new Abstract: Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems independently, limiting opportunities for proactive customer re-engagement. This is particularly challenging for exploratory intents such as best smartph...
40. FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis ​
Author: Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, Feng Zhao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.PL
arXiv:2608.18580v1 Announce Type: new Abstract: Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if thes...
41. Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference ​
Author: Zishan Ahmad, Vishal Vaddina
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.18591v1 Announce Type: new Abstract: Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark prov...
42. CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence ​
Author: Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.18613v1 Announce Type: new Abstract: Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly (planning loops, tool protocols, context managemen...
43. Preference Reasoning under Indeterminacy in Large Language Models ​
Author: Hadi Hosseini, Samarth Khanna, Xiyuan Wang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG
arXiv:2608.18631v1 Announce Type: new Abstract: As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inheren...
44. Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search ​
Author: Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18665v1 Announce Type: new Abstract: Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search useful for reducing manual design cost. However, existing automated machine/deep learning (AutoML/AutoDL) reports typi...
45. Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction ​
Author: Zhaoxi Wei, Hongye Yang, Shuyuan Tian
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.18677v1 Announce Type: new Abstract: Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural art-historical narrative construction. We present Sanyu Studio, a multi-agent dialogue system that models 321 Sany...
46. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training ​
Author: Yugu Li, Jimmy Cao, Jianglin Qiao, Siyi Hu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18682v1 Announce Type: new Abstract: Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly u...
47. Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization ​
Author: Chenle Chen, Yangbo Wei, Chao Yao, Shaoqiang Lu, Junhong Qian, Chen Wu, Lei He
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18719v1 Announce Type: new Abstract: Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifie...
48. A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation ​
Author: Manoj N M, Vijayakrishna S, Manjunath Srinivas, Rohit Pahan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18740v1 Announce Type: new Abstract: This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operate in a sequential pipeline to process natural language queries, retrieve and analyze data, generate visualization...
49. Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots ​
Author: Xing Zhang, Yanwei Cui, Guanghui Wang, Zhihao Lin, Peiyang He
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE
arXiv:2608.18744v1 Announce Type: new Abstract: Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer g...
50. Pairwise Logical Selection of Enthymeme Completions under Semantic-Link Uncertainty ​
Author: Xuyao Feng, Antonis Bikakis
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18820v1 Announce Type: new Abstract: Arguments often omit premises or claims, forming enthymemes. We study pairwise logical selection between two candidates for the omitted component. Existing natural language methods can identify or generate candidates but often do not expose how the sel...
51. Verifiable abstention makes AI leak diagnosis accountable in water distribution networks ​
Author: Tianwei Mu, Yue Wang, Mingzhe Yuan, Manhong Huang, Wenhong Wang, Xuerui Yin, Qing Luo, Min Xiao, Hui Yang, Jun Li, Dan Xue
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18836v1 Announce Type: new Abstract: Utilities lose a substantial share of treated water to leakage, yet rarely trust artificial-intelligence localizers to dispatch crews: guessing everywhere cannot justify excavation. The gap is accountability, not accuracy: no method proves when it shou...
52. ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery ​
Author: Mingzhao Li, Chenxi Liu, Yan Zhao, Hao Miao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18846v1 Announce Type: new Abstract: Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recent learning-based methods model spatiotemporal dependencies among orders to predict courier service sequences, but...
53. SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents ​
Author: Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18852v1 Announce Type: new Abstract: Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of...
54. DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning ​
Author: Zijie Meng, Xiwei Dai, Yixuan Tang, Jin Hao, Yang Feng, Fudong Zhu, Xiaoqiang Liu, Shaosheng Cao, Zuozhu Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.18878v1 Announce Type: new Abstract: Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most exist...
55. Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models ​
Author: Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18884v1 Announce Type: new Abstract: Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We present EvoResearcher, a training-free, inference-time protocol that adds...
56. Syntactic Simplification of OWL Class Expressions ​
Author: Alkid Baci, N'Dah Jean Kouagou, Caglar Demir, Axel-Cyrille Ngonga Ngomo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18899v1 Announce Type: new Abstract: Class expression learning often produces complex OWL class expressions that are difficult to interpret and reason over. However, by following theoretically grounded simplification principles, this complexity can be reduced. In this paper, we propose Cl...
57. \textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems ​
Author: Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18900v1 Announce Type: new Abstract: As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. Deep learning models underlying modern AI systems, therefore, must undergo thorough testing to ensure their correc...
58. Breaking the weakest link to evade vision language models ​
Author: Ilan Zini, Boussad Addad, Katarzyna Kapusta
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.18938v1 Announce Type: new Abstract: Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. Despite their growing deployment, the robustne...
59. A Theory of Post-hoc Debate Judgement ​
Author: Xiang Yin, Adam Dejl, Antonio Rago, Lihu Chen, Francesca Toni
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.19002v1 Announce Type: new Abstract: Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other ag...
60. Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models ​
Author: Valentin Romanov, Monique Bax, Steven Niederer
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2608.19025v1 Announce Type: new Abstract: Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We dem...
61. Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering ​
Author: Pradeep Murugesan, Luoxiao Yang, Xueli Chen, Xinqi Fan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA
arXiv:2608.19029v1 Announce Type: new Abstract: Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuanced reasoning. Existing medical QA systems, typically based on single-agent architectures and static retrieval, o...
62. Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery ​
Author: Alizer Wong, Heng Cui, Yi Tan, Xiongchao Zhan, Liang Lin, Yuxiang Guo, Zhaorong Dai, Zixin Zeng, Wenyuan Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, math.NT
arXiv:2608.19047v1 Announce Type: new Abstract: We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics. During execution, Eureka forms Macro-Agents with specialized state, memory, operators, too...
63. What is Missing from AI Post-Training AI: An Empirical Analysis ​
Author: Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.19072v1 Announce Type: new Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct...
64. Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk ​
Author: Deep Kumar Ganguly, Jan K\v{r}et'insk'y
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2608.19073v1 Announce Type: new Abstract: An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alter...
65. Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering ​
Author: George Andrikopoulos
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.19125v1 Announce Type: new Abstract: When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class returns. I argue this is an operations problem, not a tooling problem: mechanisms for persisting corrections exist and are shipping, but...
66. Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems ​
Author: George Andrikopoulos
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.SE
arXiv:2608.19140v1 Announce Type: new Abstract: Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now s...
67. Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication ​
Author: Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.19161v1 Announce Type: new Abstract: Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for ...
68. SuTRA : Structurally-Unified Tokenization with Root Awareness ​
Author: Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Surekha, Neha Bhargava
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18087v1 Announce Type: cross Abstract: Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthogra...
69. Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining ​
Author: Godwin Abuh Faruna
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18089v1 Announce Type: cross Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource i...
70. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities ​
Author: Yousef Radwan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18090v1 Announce Type: cross Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion --...
71. Self- and Other-Labels Induce Bidirectional Bias in LLM Judges ​
Author: Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where ...
72. Abliteration Mitigation via Refusal Aliases ​
Author: Nathan Truong
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR
arXiv:2608.18093v1 Announce Type: cross Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment u...
73. NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages ​
Author: Badal Nyalang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18094v1 Announce Type: cross Abstract: Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual encoder model trained ...
74. Backdoor Learning in Language Models and Vision-Language Models ​
Author: Weimin Lyu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18095v1 Announce Type: cross Abstract: Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks th...
75. Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems ​
Author: Sukanta Ganguly
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18098v1 Announce Type: cross Abstract: Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog systems, yet existing strategies treat all cached entries uniformly or apply coarse eviction heuristics that fail to adapt as dialog topics evolve....
76. Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) ​
Author: Maha Shahid
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.18100v1 Announce Type: cross Abstract: AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks...
77. DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models ​
Author: Wenxin Duan, Hanwei Wang, Zhongying Peng, Zhonghua Lu, Jiayi An, Fan Song, Yong Liang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18103v1 Announce Type: cross Abstract: Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achievi...
78. StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data ​
Author: Akshat Parmar, Vikranth Udandarao, Abhay Shakya, Tanmay Hire, Avinash Anand, Rajiv Ratn Shah, Daniel Wang Zhengkui
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.18105v1 Announce Type: cross Abstract: StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. The system combines streaming speech recognition, retrieval-augment...
79. Different Facets of Verbalised Overconfidence: an Interpretability Study ​
Author: Davide Mazzaccara, Leonardo Bertolazzi, Raffaella Bernardi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18106v1 Announce Type: cross Abstract: Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, acros...
80. Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals ​
Author: Maikel Leyva-Vazquez, Florentin Smarandache
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18107v1 Announce Type: cross Abstract: We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API call...
81. Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation ​
Author: Spencer Gibson, Tyler Crosse, Magnus Saebo, Achyutha Menon, Eyon Jang, Diogo Cruz
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.MA
arXiv:2608.18108v1 Announce Type: cross Abstract: Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable...
82. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings ​
Author: Mingfang Zhang, Jarod L'evy, Cedric Rommel, J'er'emy Rapin, Corentin Bel, Julie Bonnaire, Daniel Nieto, Pierre Bourdillon, Svetlana Pinet, St'ephane d'Ascoli, Thomas Moreau, Jean-R'emi King
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, eess.SP, q-bio.NC
arXiv:2608.18114v1 Announce Type: cross Abstract: Restoring communication for people who have lost the ability to speak or move after a brain injury is a major challenge. While intracranial implants now enable high-performing brain-computer-interfaces, non-invasive alternatives are still lagging beh...
83. Temporal Multi-Signal Fusion for Token-Level Hallucination Detection ​
Author: Igor Itkin
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.18115v1 Announce Type: cross Abstract: Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence...
84. Global Index on Responsible AI 2026 : Conceptual Framework and Methodology ​
Author: Fola Adeleke, Rachel Adams, Ayantola Alayande, Daniela Benavente, Ana Florido, Nicol'as Grossman, Leah Junck
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.18122v1 Announce Type: cross Abstract: This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by strengthening the distinction between framework existence and implementation, restructuring dimensions from three...
85. Language Models for Portuguese: A Systematic Mapping Study ​
Author: Jhessica Silva, Carlos Caetano, Helena Maia, Breno Bernard Nicolau de Fran\c{c}a, Sandra Avila, Helio Pedrini
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18138v1 Announce Type: cross Abstract: In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In...
86. The Deontic Gap: Large Language Models and the Modal Language of Obligation ​
Author: Daniel Hart, Sarah Allred, Joseph Abbas, Morenike Alugo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18144v1 Announce Type: cross Abstract: Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic moda...
87. Entropy-Constrained Adaptive Stochastic Quantization ​
Author: Ran Ben Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay Vargaftik
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, cs.IT, math.IT
arXiv:2608.18147v1 Announce Type: cross Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness. It is designed to alleviate the communication and memory bottlenecks of...
88. TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving ​
Author: Chenxu Niu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.DC, cs.LG, cs.PF
arXiv:2608.18149v1 Announce Type: cross Abstract: Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly and a cheap predictor can be dangerously confident outside its measured scope. We present TokenPowerSandbox, an e...
89. How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection ​
Author: Syeda Anshrah Gillani, Mirza Samad Ahmed Baig, Shahid Munir Shah, Asher Ali, Hamzah Siddiqui
Published: 8/20/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET, cs.LG
arXiv:2608.18155v1 Announce Type: cross Abstract: Quantum machine learning (QML) for network intrusion detection (NIDS) is routinely reported to reach near-perfect accuracy, yet the most rigorous studies find that well-tuned classical models remain competitive, and that apparent quantum gains may be...
90. When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators ​
Author: Praphulla Lal Shrestha
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18158v1 Announce Type: cross Abstract: LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching and brand mislabelin...
91. Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation ​
Author: M P V S Gopinadh
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR
arXiv:2608.18164v1 Announce Type: cross Abstract: Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented prompts as a test case...
92. What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems ​
Author: Emanuele Ratti, Lena Zuchowski
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2608.18186v1 Announce Type: cross Abstract: In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. I...
93. A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities ​
Author: Rafael Mu~noz-Terol, Jes'us Peral, Sandra Amador, David Gil
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18188v1 Announce Type: cross Abstract: Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. As the causes of ASD remain unclear, identifying relevant features and hidden correlations is crucial for early diagnosi...
94. Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift ​
Author: Souraj Adhikary, Negar Chabi, Andre Mastmeyer
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18193v1 Announce Type: cross Abstract: Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-...
95. GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction ​
Author: Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.18234v1 Announce Type: cross Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existi...
96. Bidirectional representational alignment between biological and artificial neural networks ​
Author: Samuel Kostousov, Abhinn Kaushik, Brokoslaw Laschowski
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18244v1 Announce Type: cross Abstract: Recent work has shown that representational alignment between biological and artificial neural networks is asymmetric: model representations predict neural responses much better than neural responses predict model representations. This asymmetry rais...
97. Visual-Prompt Guided Wildlife Instance-Level Recognition ​
Author: Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl, Fredrik Gustafsson
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18246v1 Announce Type: cross Abstract: Fine-grained wildlife re-identification remains a challenging area in research. Current state-of-the-art approaches apply a detection and re-identification pipeline. We propose a one-stage end-to-end detection and re-identification model that perform...
98. How AI Prompts Can Teach Us About the Structure of Human Behavior ​
Author: Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
Published: 8/20/2026, 4:00:00 AM
Categories: econ.TH, cs.AI
arXiv:2608.18265v1 Announce Type: cross Abstract: We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large language model a ``type vector'' and then prompt it to choose actions across settings in which we observe human c...
99. SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents ​
Author: Yingjie Xu, Siwei Yu, Jianwei Ma
Published: 8/20/2026, 4:00:00 AM
Categories: physics.geo-ph, cs.AI, cs.NE, eess.SP
arXiv:2608.18272v1 Announce Type: cross Abstract: Classical seismic data reconstruction relies on manually designed structural priors and iterative operators, whose coupled design space is far larger than manual trial and error can explore systematically. Deep-learning methods encode the reconstruct...
100. What Makes Software Issue Resolution Tasks Difficult for Agents? ​
Author: Ebtesam Al-Haque, Brittany Johnson
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG
arXiv:2608.18280v1 Announce Type: cross Abstract: Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and characterization of task difficulty....
101. Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements ​
Author: Naoki Egami, Sooahn Shin
Published: 8/20/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.CL, cs.LG, stat.ML
arXiv:2608.18294v1 Announce Type: cross Abstract: An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads t...
102. FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation ​
Author: Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.18296v1 Announce Type: cross Abstract: As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demo...
103. FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning ​
Author: Holger R. Roth, Ziyue Xu, Peter Cnudde
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18311v1 Announce Type: cross Abstract: Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and ECGs. We study this setting on a MIMIC-derived respiratory deterioration task with simulated FL clients and in...
104. From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model ​
Author: Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2608.18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant efforts are devoted to adapting VLMs at test time, they rely heavily on no...
105. Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring ​
Author: Steven C. Nesbit (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA), Victor M. Vergara (AeroVironment Inc., Albuquerque, USA), Michael A. Felix (University of New Mexico COSMIAC Research Center, Albuquerque, USA), Evan T. Kain (Air Force Research Laboratory, Kirtland AFB, USA), Luis R. Garc'ia Carrillo (Air Force Research Laboratory, Kirtland AFB, USA), Gerd J. Kunde (Nuclear and Particle Physics and Applications, P-3, Los Alamos National Laboratory, Los Alamos, USA), Andrew T. Sornborger (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA)
Published: 8/20/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.ET, cs.LG, cs.SD, eess.AS
arXiv:2608.18341v1 Announce Type: cross Abstract: Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is constrained by power, latency, and deployment complexity. We demonstrate autoencoder-based acoustic anomaly detection on an Intel Loihi 2 ne...
106. Coupled-cluster molecular properties across the main group that extrapolate beyond training size ​
Author: Wenhao He, Xu Chen, Noah Song, Haowei Xu, Tim S. Hindges, Bohan Li, Zihan Lin, Yu Yao, Avetik R. Harutyunyan, Fang Liu, Yao Wang, Hao Tang, Ju Li
Published: 8/20/2026, 4:00:00 AM
Categories: physics.chem-ph, cond-mat.mtrl-sci, cs.AI, cs.LG, physics.comp-ph
arXiv:2608.18346v1 Announce Type: cross Abstract: Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off wi...
107. Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents ​
Author: Alexander Tu, Michael Tu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SY, eess.SY
arXiv:2608.18351v1 Announce Type: cross Abstract: Tool-using large language-model agents can complete a task while exercising authority that the user did not grant or the task does not need, causing excess-authority errors. Traditional permission gating systems alone for validating agent environment...
108. One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI ​
Author: Gaston Besanson
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY
arXiv:2608.18360v1 Announce Type: cross Abstract: Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates that can admit, degrade, or remediate an action before it executes. This paper's central object is remediation...
109. Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation ​
Author: Guiv Farmanfarmaian
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.18379v1 Announce Type: cross Abstract: When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. A correct aggregate answer may therefore reflect recombination, fresh solving, or both. For efficient test-time re...
110. TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs ​
Author: Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad, Eric Granger
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18386v1 Announce Type: cross Abstract: Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the wild. ER is inherently multimodal, requiring a joint understanding of facial expressions, vocalizations, languag...
111. LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents ​
Author: Daehong Kim, Haichao Miao, Shusen Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.18398v1 Announce Type: cross Abstract: Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing out...
112. Vector Symbolic Policy Gradient ​
Author: Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong, Hyunwoo Oh, Raheeb Hassan, Pietro Mercati, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SC
arXiv:2608.18404v1 Announce Type: cross Abstract: We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. Under the standard softmax policy-gradient surrogate...
113. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B ​
Author: Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18419v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with which they accomplish...
114. Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage ​
Author: Shreeya Sharma, Ravish Gupta, Saket Kumar, Abhishek Aggarwal
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.18438v1 Announce Type: cross Abstract: Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a "supervision gap" where novice therapists manage high-stakes risks with delayed professional feedback. This paper proposes a new framework utilizing a fi...
115. Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas ​
Author: Dmitry V. Alexandrov
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CC, cs.PL
arXiv:2608.18445v1 Announce Type: cross Abstract: We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a triplet-based combinatorial framework for reasoning about compatible paths through layered triplet structures, called Compact Triplet...
116. ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems ​
Author: Ergan Shang, Flavio Sales Truzzi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18469v1 Announce Type: cross Abstract: Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often leave modern accelerators underutilized. Conventional training compounds this inefficiency by scheduling the forwa...
117. Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement ​
Author: Zhiyuan Yan, Ziyue Zheng, Hongce Zhang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.18482v1 Announce Type: cross Abstract: Hardware functional verification relies on high-quality assertions to expose design bugs and establish confidence in Register Transfer Level (RTL) designs. Yet existing assertion mining methods still struggle to produce complete and reliable assertio...
118. Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models ​
Author: Pardis Taghavi, Reza Langari, Gaurav Pandey
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18484v1 Announce Type: cross Abstract: Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained ...
119. Physics-Unrolled Neural Operator for Wireless Field Modeling ​
Author: Rafid Umayer Murshed, Saif Ur Rahman, Mingyue Tang, Elahe Soltanaghai
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18495v1 Announce Type: cross Abstract: Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. Machine...
120. Science Done on a Machine by a Machine: AI Agents in Computational Chemistry ​
Author: Pavlo O. Dral, Hassan Nawaz, Arif Ullah
Published: 8/20/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI, physics.comp-ph
arXiv:2608.18508v1 Announce Type: cross Abstract: We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, and the current number approaches fifty, surveyed in this Perspective as of 8 August 2026. The capabilities of th...
121. OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation ​
Author: Soumili Ghosh, Debapriya Roy, Aryan Das, Bikash Santra
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18516v1 Announce Type: cross Abstract: Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistently high performance across datasets while maintaining low computational requirements remains a significant chal...
122. GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels ​
Author: Rafid Umayer Murshed, Shahab Hamidi-Rad, Elahe Soltanaghai, Akshay Malhotra
Published: 8/20/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, math.IT
arXiv:2608.18522v1 Announce Type: cross Abstract: Large antenna arrays allow wireless systems to serve more users and achieve higher data rates, but they also make channel feedback expensive: the receiving device must repeatedly report a large complex-valued channel matrix to the base station. Most ...
123. Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection ​
Author: Shashank Kotyan, Makoto Shing, Yuki Imajuku, Rujikorn Charakorn, Tarin Clanuwat
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18523v1 Announce Type: cross Abstract: Diffusion-based generators have made synthetic images ubiquitous, but detectors often fail under simultaneous shifts in generator, prompt/style, and source-domain. We study AI-generated image detection as a transfer system described by training prior...
124. DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents ​
Author: Hangrui Xu, Jiarui Wang, Yang Yang, Chuanbo Zhu, Fangda Chen, Ziqi Wu, Jingming Cai, Yan Song
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2608.18524v1 Announce Type: cross Abstract: Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks involving multiple...
125. Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions ​
Author: Ruiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei, Wen Shen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.18539v1 Announce Type: cross Abstract: The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelevant changes in prompts can cause dramatic fluctuations in performance, a phenomenon known as prompt sensitivity...
126. CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks ​
Author: Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, Abhishek Nagaraj
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.MA, econ.GN, q-fin.EC
arXiv:2608.18554v1 Announce Type: cross Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best o...
127. Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments ​
Author: Deepak Kanneganti, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Erik Elmroth, Aneesh Krishna, Monowar Bhuyan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18555v1 Announce Type: cross Abstract: Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT) environments, widely adopted across healthcare, smart homes, and industry due to its cost-effectiveness. Howe...
128. MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence ​
Author: Xi Wu, Yanqing Wei, Hang Yin, Pengze Li, Hongshuai Qi, Xi Chen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.geo-ph
arXiv:2608.18558v1 Announce Type: cross Abstract: The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, informing shoreline protection strategies and managing coastal ecosystems under changing environmental conditions. Ho...
129. The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations ​
Author: Alexander Johnson, Obadah Ghizawi, Ali A. Minai
Published: 8/20/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.RO, cs.SY, eess.SY, q-bio.NC
arXiv:2608.18569v1 Announce Type: cross Abstract: Spatial aliasing occurs when two or more distinct locations produce highly similar place-cell representations, primarily due to environmental symmetry or repetitive structures. This issue is most pronounced when place representations are constructed ...
130. MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment ​
Author: Yuan li, Youyuan Lin, Chenhui Chu, Shin'ya Nishida
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18579v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between quality ratings and their underlying reasoning. However, most approaches supervise reasoning through human-provid...
131. From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning ​
Author: Zuocheng Ying, Yang Yang, Yumou Wu, Chuanbo Zhu, Jiarui Wang, Ziqi Wu, Jingming Cai, Junqing Yu, Zikai Song
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.18581v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key bottleneck in factual question answering. Existing end-to-end methods entangle knowledge elicitatio...
132. OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios ​
Author: Zinuo Guo, Min Zhang, Bo Jiang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored. Existing OCR benchmarks focus largely on printe...
133. Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings ​
Author: Yubo Wang, Shujie Cui, James Bailey, Hongzhi Yin, Wenyu Liang, Min Tang, Shiyue Qin, Weiqing Wang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2608.18610v1 Announce Type: cross Abstract: Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can expose substantial in...
134. Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation ​
Author: Yunwon Tae, Minje Park, Gyunho Rho, Dongjoon Yoo, Sunghoon Joo
Published: 8/20/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2608.18639v1 Announce Type: cross Abstract: Non-invasive continuous blood pressure (BP) monitoring using photoplethysmography (PPG) is a promising alternative to cuff-based measurements. However, existing PPG-based BP estimation studies predominantly rely on aggregated performance metrics (e.g...
135. Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics ​
Author: Masafumi Endo, Kohei Honda, Yuu Jinnai, Ryo Yonetani
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.18672v1 Announce Type: cross Abstract: We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While most existing OP formulations assume rewards to be known in advance, practical applications involve uncertain a...
136. Aslema at NADI 2026: Augmentation through Fewshot for SLU ​
Author: Tajwaar Shafiq, Hunzalah Hassan Bhatti, Shammur Absar Chowdhury, Firoj Alam
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18689v1 Announce Type: cross Abstract: We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate four omni LLMs in a zero-shot setting and compare them with fine-tuned models. Our results show that fine-tuni...
137. Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections ​
Author: Jacopo Ghirri, Carlos Rodriguez-Pardo, Lara Aleluia Reis, Massimo Tavoni
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, econ.GN, q-fin.EC
arXiv:2608.18690v1 Announce Type: cross Abstract: The European Union has committed to reducing greenhouse gas emissions 55% below 1990 levels by 2030, but whether current trends are compatible with this ambition remains uncertain. We apply deep learning to high-resolution socioeconomic and sectoral ...
138. Composed Historical Image Retrieval by Modeling Temporal Representations ​
Author: Adri`a Molina Rodr'iguez, Oriol Ramos Terrades, Josep Llad'os Canet
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR
arXiv:2608.18694v1 Announce Type: cross Abstract: While time evolves linearly, the geometry of neural embedding spaces is inherently multi-dimensional, often chaotic, and difficult to interpret. In principle, one could constrain an embedding space to a single temporal dimension; however, such a redu...
139. Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts ​
Author: Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18696v1 Announce Type: cross Abstract: Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exhibit complex heterogeneous lay...
140. MemFuse: Multi-Source Memory Fusion from Fragmented Observations ​
Author: Chao Li, Yuanfa Li, Wenhao Wu, Xule Liu, Zhi Wang, Kun Shao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18704v1 Announce Type: cross Abstract: Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantly focus on single-source textual histories. In realistic settings, however, relevant information is often fragm...
141. A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation ​
Author: Steven Landgraf, Joceline Hinz, Markus Ulrich
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18709v1 Announce Type: cross Abstract: Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain generalization. Yet their lack of interpretability, tendency to be overconfident, and sensitivity to real-world ...
142. The Impact of CutMix on Reliability and Robustness in Semantic Segmentation ​
Author: Steven Landgraf, Markus Ulrich
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.18715v1 Announce Type: cross Abstract: Ensuring not only high accuracy but also reliable and robust predictions is critical for the deployment of semantic segmentation models in safety-critical applications such as autonomous driving. Despite the widespread use of CutMix - a simple yet po...
143. Budget-First Tariff Recommendation (BFTR): A Complete Algorithmic Framework for Telecom Plan Recommendation without Overcharging ​
Author: Ghislain Dorian Tchuente Mondjo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18723v1 Announce Type: cross Abstract: Telecom operators traditionally offer predefined tariff grids, forcing users to choose from a limited set of plans. This paper proposes BFTR (Budget-First Tariff Recommendation), a complete algorithmic framework integrating eight Budget-First strateg...
144. A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3 ​
Author: Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18731v1 Announce Type: cross Abstract: Medical image segmentation is essential for clinical workflows such as treatment planning and disease assessment. While specialist tools like TotalSegmentator and MRSegmentator achieve strong performance, they require large annotated datasets for tra...
145. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services ​
Author: Jos'e A. Perdiguero L'opez, Miguel A. Dur'an-Olivencia
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.18733v1 Announce Type: cross Abstract: We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications. Built on the Asynchronous Server Gateway Interface (ASGI), Flama offers ...
146. Epistemic Subordination: Generative AI and the Infrastructure of Knowledge ​
Author: Gilad Abiri, Emanuel V. Towfigh
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.18758v1 Announce Type: cross Abstract: Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expressi...
147. Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening ​
Author: Kerol Djoumessi, Philipp Berens
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18759v1 Announce Type: cross Abstract: Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), ...
148. SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation ​
Author: Jiandong Ding, Huijie Qin, Tiandeng Wu, Yi Cao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.18779v1 Announce Type: cross Abstract: Semantic-ID mappings are reusable interfaces between item tokenizers and generative recommenders, yet released mappings rarely state whether they are coherent, what structure they expose, how generated paths resolve, or what must be revalidated after...
149. Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study ​
Author: Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18795v1 Announce Type: cross Abstract: Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic agreement index Gamma i...
150. Forgetting, plasticity, and co-observation: a third facet of continual learning ​
Author: Timm Hess, Abhishek Jha, Gido M. van de Ven, Tinne Tuytelaars
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.18803v1 Announce Type: cross Abstract: Efficient continual learning remains a fundamental challenge for deep neural networks. While catastrophic forgetting and loss of plasticity are widely considered the primary obstacles to overcome, we show that these two issues cannot fully explain th...
151. A strengthening of the MCFL-ness of $O_2$ ​
Author: Marco B. Caminati
Published: 8/20/2026, 4:00:00 AM
Categories: cs.FL, cs.AI, cs.LO, math.LO
arXiv:2608.18813v1 Announce Type: cross Abstract: In the last years, a number of proofs of the fact that $O_2$ is a multiple context-free grammar (MCFG) were given. Such results can be exploited in the fields of both computational linguistics and of computational algebra. Here, we focus on a recent ...
152. Do Large Language Models Hallucinate Electric Fata Morganas? ​
Author: Kristina \v{S}ekrst
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18816v1 Announce Type: cross Abstract: AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical significance when it...
153. Identifying Implicit Premises for Logical Reconstruction of Argument Graphs ​
Author: Xuyao Feng, Anthony Hunter
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18821v1 Announce Type: cross Abstract: The logical reconstruction of argument graphs from natural language text is challenging because of the prevalence of enthymemes (i.e., arguments with implicit premises). There are natural language processing methods for identifying enthymemes in text...
154. Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis ​
Author: Souranil Kahali, Rituparna Bose, Abner Hernandez, Tomas Arias-Vergara, Andreas Maier, Ning Ma, Paula Andrea Perez-Toro
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD
arXiv:2608.18825v1 Announce Type: cross Abstract: Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilingual use cases. Although large-scale pretrained ASR models such as Whisper achieve strong generalisation, their...
155. MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models ​
Author: Chenglin Liu, Xun Wang, Ruishuo Chen, Zhuoran Li, Longbo Huang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.18827v1 Announce Type: cross Abstract: Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to r...
156. Learning-State-Aware Dynamic Generative Data Augmentation on Small-Scale Datasets ​
Author: Ting Xiang, Chenxi Deng, Jinhui Zhao, Bingting Jiang, Ke Zhang, Changjian Chen, Zhuo Tang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18907v1 Announce Type: cross Abstract: Small-scale image classification is often limited by the scarcity of training data. Generative data augmentation (GDA) based on pretrained generative models has emerged as an effective solution. However, existing methods rely on task-agnostic augment...
157. SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance ​
Author: Jian Yang, Zhenqi Feng, Zhaoyang Yu, Zhaoxin Fan, Kejian Wu, Xiaofeng Wang, Zheng Zhu, Jianjun Huang, Wei You, Bin Liang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18921v1 Announce Type: cross Abstract: Existing LRM-DoS methods rely heavily on model feedback to synthesize attack queries, requiring either repeated queries to the target model or training a dedicated attack model. These expensive operations severely weaken attack leverage. In this pape...
158. Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck ​
Author: Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofr`e
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18931v1 Announce Type: cross Abstract: Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and...
159. SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution ​
Author: Silin Chen, Han Li, Xiaodong Gu, Yuling Shi, Haibing Guan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.18933v1 Announce Type: cross Abstract: Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often struggle to resolve issues in a specific repository because they lack project-specific knowledge. Existing self-ev...
160. Graphical Design of Interpretable Architectures ​
Author: Pietro Barbiero
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2608.18936v1 Announce Type: cross Abstract: Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representations fall short in one of two ways. Symbolic equations give no global view of an architecture at a glance. Pro...
161. MedUAG: Unified Understanding and Generation for Medical Multimodal Models ​
Author: Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18937v1 Announce Type: cross Abstract: Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training an...
162. Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis ​
Author: Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.CL
arXiv:2608.18940v1 Announce Type: cross Abstract: Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting ...
163. AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL ​
Author: Daniele Lizzio Bosco, Jacopo Cossio, Carla Piazza, Giuseppe Serra
Published: 8/20/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2608.18946v1 Announce Type: cross Abstract: Clifford circuits play a foundational role in quantum computing, particularly due to their importance in quantum error correction and fault-tolerant logical synthesis. While these circuits can be efficiently simulated and represented as symplectic ma...
164. rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation ​
Author: Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.18952v1 Announce Type: cross Abstract: Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extracting a user's preferences or explaining why one item fits better than another - rather than mapping history di...
165. DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering ​
Author: Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18988v1 Announce Type: cross Abstract: Retrieve-then-generate pipelines are commonly used to produce deep-research answers for open-ended questions, but retrieval alone is insufficient: LLMs must organize noisy and fragmented evidence into comprehensive, well-cited answers. We refer to th...
166. GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery ​
Author: Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu, Jiaxu Tian, Yuheng Li, Guangqian Guo, Shan Gao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.18996v1 Announce Type: cross Abstract: Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distributed, and visually similar objec...
167. From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation ​
Author: Sepehr Ghaffarzadegan, Boubakr Nour, Makan Pourzandi, Mourad Debbabi, Chadi Assi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.19011v1 Announce Type: cross Abstract: Mechanisms for dynamically converting cyber threat intelligence (CTI) into actionable detection capabilities are necessary due to the rapid evolution of Advanced Persistent Threats (APTs). Sigma rules are an essential part of contemporary threat dete...
168. Harness Continual Learning: Continual Adaptation Beyond Model Parameters ​
Author: Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.19013v1 Announce Type: cross Abstract: Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these co...
169. One-Stage Object Detectors in Autonomous Driving ​
Author: Jonel Roman, Ryan Sirjue, Peter Nguyen, Daniel Krutky, Juan Jesus, Sudip Dhakal
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.19014v1 Announce Type: cross Abstract: Autonomous vehicles depend on fast and reliable perception systems to detect surrounding vehicles, pedestrians, cyclists, traffic signs, and other road objects in real time. This paper presents a comprehensive survey and analysis of one-stage object ...
170. Counterfactual Contrastive Analysis ​
Author: Yunlong He, Pietro Gori
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.19032v1 Announce Type: cross Abstract: Visual Counterfactual Explanations (VCEs) aim to explain image classifiers by generating minimally edited and realistic versions of an input image that change the classifier's prediction. Existing VCE methods are inherently classifier-dependent and t...
171. Bernstein-Vazirani Networks: Quantum Machine Learning by Interference ​
Author: Natacha Kuete Meli, Tolga Birdal, Prayag Tiwari, Vladislav Golyanik, Michael Moeller
Published: 8/20/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.CV, cs.LG
arXiv:2608.19043v1 Announce Type: cross Abstract: We introduce Bernstein-Vazirani Networks (BVNs), a non-variational quantum machine learning framework that leverages quantum interference for supervised learning, demonstrated on vision and representation learning tasks. In their standard form, BVNs ...
172. GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting ​
Author: Yechan Park, HyunJin Kim
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.19066v1 Announce Type: cross Abstract: This paper proposes a lightweight, plug-and-play framework that improves robustness to viewpoint shifts in Vision-Language-Action (VLA) policies without policy retraining. To our knowledge, this is the first approach to directly leverage 3D Gaussian-...
173. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models ​
Author: Jihae Jeong, Junha Choi, Hwanjo Yu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.19075v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content during decoding calls for a candidate-specific measure of how strongly the image supports the token under consid...
174. DA-WAM: Decision-Aligned Future Latents for Driving World Models ​
Author: Ruiguo Zhong, Benshan Ma, Xiaolong Chen, Lang Zhang, Mingyue Feng, Yaonong Wang, Pei Liu, Jun Ma
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.19085v1 Announce Type: cross Abstract: Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for decision-making remains unrealized. The critical challenge lies in ensuring that future modeling is not merely pred...
175. Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift ​
Author: Longtian Wang, Zhengyu Zhao, Chenhao Lin, Le Yang, Shiwei Wang, Yuhan Zhi, Xiaofei Xie, Chao Shen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.19088v1 Announce Type: cross Abstract: Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors when a hidden trigger is present. Existing detection methods either rely on trigger inversion or exploit architect...
176. Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation ​
Author: Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.19098v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practica...
177. Discretizing Continuous Time Series for Imputation with Masked Diffusion Training ​
Author: Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.19119v1 Announce Type: cross Abstract: Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations: missing and observed valu...
178. PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints ​
Author: Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.19121v1 Announce Type: cross Abstract: Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. ...
179. Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers ​
Author: Francesco Cordella, Mauro Cappelli
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.19124v1 Announce Type: cross Abstract: Astrolinguistics -- communication with minds that categorize reality differently from ours -- has been purely speculative since Freudenthal's Lincos (1960). We make it experimental. Two language models with deliberately incompatible constructed lexic...
180. Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles ​
Author: Emanuele Luzio
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2608.19127v1 Announce Type: cross Abstract: A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coo...
181. Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets ​
Author: Tate Berenbaum, Muthaiah Venkatachalam
Published: 8/20/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.SE
arXiv:2608.19147v1 Announce Type: cross Abstract: Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B-parameter LLM. We show that a handful of AIPCs, working toget...
182. Interpretable AI predicts a 2026 summer dry anomaly in central China ​
Author: Anran Wang, Wen Shi, Yong Luo, Jianbin Huang, Lijuan Chen, Junhu Zhao, Weixin Jin, Huihui Yuan
Published: 8/20/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI
arXiv:2608.19163v1 Announce Type: cross Abstract: Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability than precipitation itself. Here, we employ a deep learning model that translates dynamical circulation predicti...
183. Finetuning Strategies for Querying Sounds by Vocal Imitation ​
Author: Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.IR
arXiv:2608.19174v1 Announce Type: cross Abstract: This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, ...
184. Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning ​
Author: Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.19181v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distribu...
185. ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning ​
Author: Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.19182v1 Announce Type: cross Abstract: We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks ...
186. SPADE: Self-Play in Adaptive Synthetic Executable Environments ​
Author: Bo Liu, Simon Yu, Yiding Jiang, Ao Qu, Andrew Zhao, Zichen Liu, Junsu Kim, Zijian Zhou, Seungone Kim, Tongzheng Ren, Mickel Liu, Hanfei Yu, Zhaorun Chen, Weiyan Shi, Paul Pu Liang, Luke Zettlemoyer, Yejin Choi, Natasha Jaques
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.19197v1 Announce Type: cross Abstract: Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fix...
187. Hybrid Reinforcement Learning and Search for Flight Trajectory Planning ​
Author: Alberto Luise, Michele Lombardi
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2509.04100v3 Announce Type: replace Abstract: This paper explores the combination of Reinforcement Learning (RL) and search-based path planners to speed up the optimization of flight paths for airliners, where in case of emergency a fast route re-calculation can be crucial. The fundamental ide...
188. Conformal Policy Control ​
Author: Drew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho, Anqi Liu, Suchi Saria, Samuel Stanton
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.ST, stat.ML, stat.TH
arXiv:2603.02196v4 Announce Type: replace Abstract: An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause harm and must be taken offline, curtailing any future interaction. Imitating old behavior is safe, but excessiv...
189. SkillNet: Create, Evaluate, and Connect AI Skills ​
Author: Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Yida Xue, Xin Xu, Tongtong Wu, Kun Wang, Yang Liu, Zhen Bi, Jungang Lou, Yuchen Eleanor Jiang, Hangcheng Zhu, Gang Yu, Haiwen Hong, Longtao Huang, Hui Xue, Chenxi Wang, Yijun Wang, Zifei Shan, Xi Chen, Zhaopeng Tu, Feiyu Xiong, Xin Xie, Peng Zhang, Zhengke Gui, Lei Liang, Jun Zhou, Chiyu Wu, Jin Shang, Yu Gong, Junyu Lin, Changliang Xu, Hongjie Deng, Wen Zhang, Keyan Ding, Qiang Zhang, Fei Huang, Ningyu Zhang, Jeff Z. Pan, Guilin Qi, Haofen Wang, Huajun Chen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG, cs.MA
arXiv:2603.04448v2 Announce Type: replace Abstract: Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systematic accumulation and transfer of skills. Without a unified mechanism for skill consolidation, agents frequently...
190. From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial? ​
Author: Binyan Xu, Dong Fang, Haitao Li, Kehuan Zhang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.01608v5 Announce Type: replace Abstract: Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through workflows spanning stages, tools, shared state, verification, and repair. Distilling such workflows into a single-agent skill can reduce orchestratio...
191. Interval POMDP Shielding for Imperfect-Perception Agents ​
Author: William Scarbro, Ravi Mangal
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY
arXiv:2604.20728v2 Announce Type: replace Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for this setting: given a proposed action, a shield blocks actions that could violate safety. We consider the co...
192. When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition ​
Author: Pehu'en Moure, Niclas Pokel, Bilal Bounajma, Yingqiang Gao, Roman Boehringer, Longbiao Cheng, Shih-Chii Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, eess.AS
arXiv:2605.02782v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by conditioning on additional clinical context at inference time, but it ...
193. Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios ​
Author: Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2605.06185v2 Announce Type: replace Abstract: Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event memory and recover long-range relationships in ultra-long videos. End-to-end methods are limited by visual-token...
194. MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance ​
Author: Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Patrick Shea, Siri Du, Haoyang Liu, Daniel Guetta, Hongseok Namkoong
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.22664v5 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs have developed agents that can construct entire spreadsheets from scr...
195. RULER: Representation-Level Verification of Machine Unlearning ​
Author: Georgina Cosma, Axel Finke
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.27569v3 Announce Type: replace Abstract: Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols verify this at the output level through membership inference, retain accuracy, and forget-set accu...
196. A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice ​
Author: Ranjan Mishra, Jakob Schoeffer
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2606.06081v2 Announce Type: replace Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous interv...
197. Teaching agentic AI to learn expert reasoning for rare disease diagnosis ​
Author: Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh F. Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.16149v4 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be ...
198. ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks ​
Author: Vincent Siu, Manasi Sharma, Dawn Song, Daniel Yue Zhang, Chenguang Wang, Ying Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.21654v2 Announce Type: replace Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap with ChainWorld, which composes atomic OSWorld tasks into long horizon ...
199. ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair ​
Author: Chiwang Luk, Matin Mohammad Najafi, Zhifeng Jia, Wei Yang, Xiuchang Li, Jinwei Zhu, Yang Ren, Lei Chen, Gao Cong
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.01916v5 Announce Type: replace Abstract: Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presen...
200. ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System ​
Author: Yutong He, Daibo Li, Guohong Li, Jiahe Geng, Zhengyang Huang, Can Ren, Zekun Zhang, Yifan Liu, Shuchen Zhu, Hengrui Zhang, Boao Kong, Ming Sun, Shu Li, Chenyi Li, Jiang Hu, Kun Yuan, Zaiwen Wen, Pingwen Zhang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2607.14178v3 Announce Type: replace Abstract: Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly focused on empirically driven domains with quantitative benchmarks,...
201. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations ​
Author: Hiskias Dingeto
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.20379v2 Announce Type: replace Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims. If flipping a...
202. Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting ​
Author: Hongqiang Lin, Chao Liu, Xiaofan Bai, Xuan Jin, Yuhong Li, Nenggan Zheng, Xipeng Cao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.26643v2 Announce Type: replace Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same w...
203. Fragility of Value under Imperfect Alignment ​
Author: Winter Cross
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28881v3 Announce Type: replace Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imper...
204. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution ​
Author: Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01324v3 Announce Type: replace Abstract: Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing approaches typically rely on linear sequential reasoning for both trajectory generation and inference, making i...
205. Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks ​
Author: Christophe D. Hounwanou, John Emeka Eze, Ya'e Ulrich Gaba
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.03502v2 Announce Type: replace Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise ...
206. BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding ​
Author: Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan, Sha Zhao
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04156v2 Announce Type: replace Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this ca...
207. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence ​
Author: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC, cs.LG, cs.MA
arXiv:2608.12036v2 Announce Type: replace Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic explo...
208. S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices ​
Author: Haochen Huang, Shengxuan Qiu, Meng Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.15018v2 Announce Type: replace Abstract: Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively ...
209. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? ​
Author: Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.15265v2 Announce Type: replace Abstract: Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand use...
210. Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling ​
Author: Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.15565v3 Announce Type: replace Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label-free alternatives are unre...
211. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies ​
Author: Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA
arXiv:2608.16645v2 Announce Type: replace Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous o...
212. GRIP: Grounded Reasoning via Information-Restricted Premises ​
Author: Lirui Teng
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.16776v2 Announce Type: replace Abstract: High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call this failure mode query dominance. To address it, we introduce \textbf{GRIP} (Gr...
213. Accuracy and Robustness of Model Cascades Under Data Perturbations ​
Author: Pallavi Mitra, Jai Kushwaha, Felix Biessmann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.17711v2 Announce Type: replace Abstract: Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cas...
214. The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting ​
Author: Nazl{\i} Nur Karabulut, Tanya Braun
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.17749v2 Announce Type: replace Abstract: Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of ag...
215. D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory ​
Author: Xule Liu, Yijun Liu, Chao Li, Shao Kun
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.17756v2 Announce Type: replace Abstract: Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to locali...
216. Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization ​
Author: Pallavi Mitra, Felix Biessmann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2407.05788v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) is an efficient framework for optimization of black-box objectives when function evaluations are costly and gradient information is not easily accessible. BO has been successfully applied to automate the task of hyp...
217. `From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs ​
Author: Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2502.00735v4 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vulnerability to audio-based adversarial attacks. These systems typically follow two architectural paradi...
218. Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling ​
Author: Eldad Haber, Shadab Ahamed, Md. Shahriar Rahim Siddiqui, Niloufar Zakariaei, Moshe Eliasof
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ML
arXiv:2502.16445v4 Announce Type: replace-cross Abstract: Generative models for image generation are now commonly used for a wide variety of applications, ranging from guided image generation for entertainment to solving inverse problems. Nonetheless, training a generator is a non-trivial feat that ...
219. Sleeping Kelly ​
Author: Ben Abramowitz
Published: 8/20/2026, 4:00:00 AM
Categories: q-fin.GN, cs.AI
arXiv:2510.15911v4 Announce Type: replace-cross Abstract: The Sleeping Beauty problem is a problem of imperfect recall that has received considerable attention. One approach to resolving the Sleeping Beauty problem has been to allow Sleeping Beauty to make decisions based on her beliefs, and then ch...
220. Jailbreaking in the Haystack ​
Author: Rishi Rajesh Shah, Chen Henry Wu, Shashwat Saxena, Ziqian Zhong, Alexander Robey, Aditi Raghunathan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2511.04707v2 Announce Type: replace-cross Abstract: Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks like computer-use agents. Yet, the safety implications of these extended contexts remain unclear. To br...
221. CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning ​
Author: Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, "Ozg"ur \c{S}im\c{s}ek
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.22842v3 Announce Type: replace-cross Abstract: Causal machine learning (Causal ML) aims to answer "what if" questions using machine learning algorithms, making it a promising tool for high-stakes decision-making. Yet, empirical evaluation practices in Causal ML remain limited. Existing be...
222. Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead ​
Author: Guang Yang, Wei Zheng, Xiang Chen, Dong Liang, Peng Hu, Yukui Yang, Shaohang Peng, Zhenghan Li, Jiahui Feng, Xiao Wei, Kexin Sun, Deyuan Ma, Haotian Cheng, Yiheng Shen, Xing Hu, Terry Yue Zhuo, David Lo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2512.00020v3 Announce Type: replace-cross Abstract: Code generation has emerged as a critical research area at the intersection of Software Engineering (SE) and Artificial Intelligence (AI), attracting significant attention from both academia and industry. Within this broader landscape, Verilo...
223. Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 ​
Author: Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC
arXiv:2512.14012v2 Announce Type: replace-cross Abstract: The rise of AI agents is transforming how software can be built. The promise of agents is that developers might write code quicker, delegate multiple tasks to different agents, and even write a full piece of software purely out of natural lan...
224. Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems ​
Author: Yash Vishe, Eric Xue, Xunyi Jiang, Zachary Novack, Junda Wu, Julian McAuley, Xin Xu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2512.14629v2 Announce Type: replace-cross Abstract: Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances in music editing systems have enabled diverse editing tasks such as timbre transfer, instrument substi...
225. TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion ​
Author: Saideep Sreekumar, Zeng Wang, Akashdeep Saha, Weihua Xiao, Minghao Shao, Muhammad Shafique, Ozgur Sinanoglu, Ramesh Karri, Johann Knechtel
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.AR
arXiv:2601.17178v3 Announce Type: replace-cross Abstract: Hardware Trojans (HTs) remain a critical threat because learning-based detectors often overfit to narrow trigger/payload patterns and small, stylized benchmarks. We introduce TrojanGYM, an agentic, LLM-driven framework that automatically cura...
226. FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance ​
Author: Hyunsuk Chung, Soyeon Caren Han, Seungyeon Ji, Jinwoo Kim, Eun-Jung Holden, Kyungreem Han
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.02060v2 Announce Type: replace-cross Abstract: Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains unclear whether their predictions can be controlled by explicitly modulating reliance on different internal feature pathways. Existing approaches t...
227. Structure-Informed Estimation for Pilot-Limited MIMO Channels via Tensor Decomposition ​
Author: Alexandre Barbosa de Lima
Published: 8/20/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2602.04083v3 Announce Type: replace-cross Abstract: Accurate channel state information in wideband MIMO systems is constrained by pilot overhead, a challenge intensifying as bandwidths scale toward 6G. This paper proposes a structure-informed hybrid estimator formulating pilot-limited MIMO cha...
228. Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence ​
Author: Yungang Yi, Weihua Li, Matthew Kuo, Catherine Shi, Quan Bai
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG
arXiv:2602.19816v3 Announce Type: replace-cross Abstract: For computational efficiency, modern language models are typically trained on independently sampled fixed-length sequences. Symbolic music language models largely inherit this paradigm, despite musical structure naturally unfolding over compl...
229. Making Implicit Premises Explicit in Logical Understanding of Enthymemes ​
Author: Xuyao Feng, Anthony Hunter
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.06114v3 Announce Type: replace-cross Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language processing (NLP) methods for handling enthymemes can potentially identify enthymemes in text but the...
230. A Framework and Prototype for a Navigable Map of Datasets in Engineering Design and Systems Engineering ​
Author: H. Sinan Bank, Daniel R. Herber
Published: 8/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CE, cs.DB, cs.DL
arXiv:2603.15722v3 Announce Type: replace-cross Abstract: The proliferation of data across the system lifecycle presents both a significant opportunity and a challenge for Engineering Design and Systems Engineering (EDSE). While this "digital thread" has the potential to drive innovation, the fragme...
231. Wildfire Suppression: Complexity, Models, and Instances ​
Author: Gustavo Delazeri, Marcus Ritt
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2603.29865v2 Announce Type: replace-cross Abstract: Wildfires cause major losses worldwide, and the frequency of fire-weather conditions is likely to increase in many regions. We study the allocation of suppression resources over time on a graph-based representation of a landscape to slow down...
232. When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't ​
Author: Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2604.06422v2 Announce Type: replace-cross Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To stu...
233. AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems ​
Author: Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.16804v3 Announce Type: replace-cross Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations...
234. MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports ​
Author: Yingyun Li, Yu Wang, Haiyang Qian
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.03103v2 Announce Type: replace-cross Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical histories. In practice, this scenario commonly involves three tasks: (i) field-header (key)...
235. Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports ​
Author: Yu Wang, Yingyun Li, Ying Qin, Haiyang Qian
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.09440v2 Announce Type: replace-cross Abstract: Clinical reports are often fragmented across healthcare institutions because privacy regulations and data silos limit direct information sharing. When patients seek care at a different hospital, they often carry paper or scanned reports from ...
236. EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding ​
Author: Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2605.09874v2 Announce Type: replace-cross Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long videos, relevant information is sparsely dist...
237. ICICLE: Expanding Retrieval with In-Context Documents ​
Author: Yu-Chen Den, Yung-Yu Shih, Zhi Rui Tam, Kuan-Yu Chen, Pu-Jen Cheng, Yun-Nung Chen, Eugene Yang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2605.26902v3 Announce Type: replace-cross Abstract: Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes corpus expansion costly: adding new documents requires updating model parameters to encode new document-do...
238. DELOS: Contrastive Deep Learning for Low-SNR Blind Transit Searches in Kepler Photometry ​
Author: Qingtian Liu, Jian Ge, XingChen Yan, Kevin Willis, Xinyu Yao, QuanQuan Hu, Jiapeng Zhu
Published: 8/20/2026, 4:00:00 AM
Categories: astro-ph.EP, astro-ph.IM, cs.AI
arXiv:2605.29428v3 Announce Type: replace-cross Abstract: We present DEtection in phase-folded Light curves with cOntrastive Scoring (DELOS), a deep-learning framework that uses contrastive scoring to perform blind searches for shallow transits in Kepler photometry. DELOS combines GPU-accelerated ph...
239. Planning-aligned Token Compression for Long-Context Autonomous Driving ​
Author: Zhixuan Liang, Yuxiao Chen, Yurong You, Peter Karkus, Wenhao Ding, Boyi Li, Alexander Popov, Yan Wang, Maximilian Igl, Yiming Li, Danfei Xu, Nikolai Smolyanskiy, Boris Ivanovic, Ping Luo, Marco Pavone
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2606.07464v3 Announce Type: replace-cross Abstract: Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed real-time computational budgets when encoding extended temporal context for complex ...
240. Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis ​
Author: Vaibhav Prakash, Jayasri Dontabhaktuni
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, quant-ph
arXiv:2606.07559v3 Announce Type: replace-cross Abstract: Language models fine-tuned where the correct completion must outrank a near-synonym competitor often fail silently. The cross-entropy loss falls monotonically while the correct token never overtakes the competitor in the model's ranking. We s...
241. Sensory Restoration via Brain-Computer Interfaces: A Scoping Review ​
Author: Xuan-The Tran
Published: 8/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2606.15091v3 Announce Type: replace-cross Abstract: Brain-computer interfaces (BCIs) can restore sensory and motor function in individuals with severe neurological impairment, but the literature is fragmented between invasive neuroprosthetics and non-invasive electrophysiological decoders, wit...
242. Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining ​
Author: Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2606.16246v3 Announce Type: replace-cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch tra...
243. Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems ​
Author: Pyuyi Chufeng Huang, Zikang Song
Published: 8/20/2026, 4:00:00 AM
Categories: math.OC, cs.AI
arXiv:2606.17762v3 Announce Type: replace-cross Abstract: We study local stationary solutions of finite-horizon discrete-time Pontryagin systems near a steady extremal. Suppose that the stationarity equation for the control is regular, the reduced state--costate map is hyperbolic, and the endpoint c...
244. Hybrid ANN-SNN Pipeline with Local Plasticity ​
Author: Denis Larionov, Khairutin Shtanchaev, Mikhail Kiselev, Mikhail Korovin, Ivan Tugoy
Published: 8/20/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2606.20151v2 Announce Type: replace-cross Abstract: This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs). The architecture couples a pretrained Effic...
245. First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers ​
Author: Arjun Pillai, Christian Hoang, Anjelo Jann Laroza
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.22361v2 Announce Type: replace-cross Abstract: Why do multilingual language models sometimes generate in the wrong language, and why is this so hard to fix? We introduce Language Identity Head Ablation (LIHA), a causal intervention that zeros each attention head individually and measures ...
246. Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models ​
Author: Riccardo O. Feingold, Davide Liconti, Chenyu Yang, Robert K. Katzschmann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2607.04546v2 Announce Type: replace-cross Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation. We present Mask2Real-WM, a two-stage ...
247. Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO) ​
Author: Fabio Hellmann, Alexander Hustinx, Benjamin D. Solomon, GestaltMatcher Database Consortium, Tzung-Chien Hsieh, Peter Krawitz, Elisabeth Andr'e
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.05585v2 Announce Type: replace-cross Abstract: FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with n...
248. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 ​
Author: Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.15509v3 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independ...
249. RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce ​
Author: Xianling Zeng, Zihan Yu, Sichen Zhao, Yalun Qi, Zhiming Xue
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16230v2 Announce Type: replace-cross Abstract: Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. In practice, shipping cost is shaped not only by distance but also by destination demand mix, billa...
250. Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting ​
Author: Xingsheng Chen, Deyu Yi, Siu-Ming Yiu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.19404v2 Announce Type: replace-cross Abstract: Existing patching and multi-scale methods advance multivariate time series forecasting but treat learned representations as transient byproducts of prediction, lacking explicit mechanisms that enforce structural consistency across temporal sc...
251. SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD ​
Author: Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhicheng Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zhang, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.20145v3 Announce Type: replace-cross Abstract: Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient ...
252. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​
Author: Jie Zhang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21636v4 Announce Type: replace-cross Abstract: Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency. Yet the most commonly reported certification score, a linear (logistic-regression) classifier two-sample test (C2ST), is largely blin...
253. Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations ​
Author: Marzieh Zare
Published: 8/20/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.ET, cs.LG
arXiv:2607.24834v3 Announce Type: replace-cross Abstract: Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Appr...
254. Untrainable elements determine what physical learning remembers ​
Author: Bijaya Dangol
Published: 8/20/2026, 4:00:00 AM
Categories: cond-mat.soft, cond-mat.dis-nn, cs.AI, cs.LG
arXiv:2608.00097v2 Announce Type: replace-cross Abstract: Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive networks through local measurements. The learned function is decided by where on the solution manifold trai...
255. The Epistemic Politics of AI Anthropomorphism ​
Author: Donna M. Bye, Levin Kuhlmann
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.00961v3 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to de...
256. Approximate Speculative Decoding ​
Author: Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03447v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding...
257. Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity ​
Author: William Zhang, Tzvika Geft, Jingjin Yu, Kostas Bekris
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA
arXiv:2608.07734v2 Announce Type: replace-cross Abstract: Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-based storage (PBS) architectures increase capacity by eliminating aisles, coordinating multiple robots...
258. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol ​
Author: Christoph Trattner
Published: 8/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.08882v3 Announce Type: replace-cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own? I call this epistemic transfer. It refers t...
259. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory ​
Author: Le Zhang, Hao Chen, Vlad Roznyatovskiy, Jianzhong Zhang, Ke Sun
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.HC
arXiv:2608.12627v3 Announce Type: replace-cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for ag...
260. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving ​
Author: Bing Zhan, Shuyao Shang, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Jiahao Gu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.12854v2 Announce Type: replace-cross Abstract: Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit ...
261. Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models ​
Author: Yang Liu, Yuming Chen
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.15156v2 Announce Type: replace-cross Abstract: World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether a small, directly addressable hidden-state change can place a learned world model on the intended coun...
262. From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents ​
Author: Zhengzhao Ma, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.16002v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or ...
263. Neurosymbolic Embodied Agents ​
Author: Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL
arXiv:2608.16794v2 Announce Type: replace-cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors...
264. Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents ​
Author: Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu
Published: 8/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.16806v2 Announce Type: replace-cross Abstract: Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing att...
265. Cross-Model Memory Transfer via Target-Side Reader Adaptation ​
Author: Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
Published: 8/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.17050v2 Announce Type: replace-cross Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration wi...
266. Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL ​
Author: Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.17253v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annota...
267. MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure ​
Author: Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Published: 8/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC
arXiv:2608.17823v2 Announce Type: replace-cross Abstract: Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. We address this gap by introducing a comprehe...