arXiv cs.AI - 2026-09-03 ​
288 items collected.
1. EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models ​
Author: Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.01611v1 Announce Type: new Abstract: Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are ...
2. Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI ​
Author: Shang Lu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2609.01685v1 Announce Type: new Abstract: With the development of artificial intelligence (AI), the landscape of meta-ethics, which has largely centred on human ethics, faces pressures that may significantly reconfigure it. In particular, if future AI systems were to exhibit sufficiently integ...
3. When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic ​
Author: Surya Saka
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.01741v1 Announce Type: new Abstract: Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri's statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43. We ask what formal logi...
4. When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection ​
Author: Yohei Nakajima
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.GT
arXiv:2609.01814v1 Announce Type: new Abstract: Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper separates those effects in exact finite discovery models. A centralized action-budget profile shows that equal one-person accuracy can coexist wi...
5. Induction and Inquiry via Probabilistic Reasoning over Language and Code ​
Author: Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01815v1 Announce Type: new Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and co...
6. Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern ​
Author: Joseph Axisa
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2609.01834v1 Announce Type: new Abstract: As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap. While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the e...
7. SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval ​
Author: Przemys{\l}aw Stok{\l}osa, Janusz A. Starzyk, Pawe{\l} Raif
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01849v1 Announce Type: new Abstract: This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SSAKGs). An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of g...
8. The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents ​
Author: Jundong Hu, Shekar Ramachandran
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.01852v1 Announce Type: new Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark w...
9. Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization ​
Author: Yuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01861v1 Announce Type: new Abstract: The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidat...
10. Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence ​
Author: Marc Bara
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2609.01873v1 Announce Type: new Abstract: Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly...
11. The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction ​
Author: Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01909v1 Announce Type: new Abstract: Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emp...
12. Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence? ​
Author: Wenlong Wang, Fergal Reid
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01924v1 Announce Type: new Abstract: Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard feedforward transformer --- a functional analogue of a global workspace. Whether the same workspace functionality emerges when depth is implemented t...
13. Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment ​
Author: Anirudh Malik, M Sparsh Mehra, Poojith Devan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2609.01962v1 Announce Type: new Abstract: Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, ...
14. Benchmarking Language Models for Statistical Problem Formulation ​
Author: Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.01982v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous da...
15. When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor ​
Author: Phanindra Reddy Madduru
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2609.01985v1 Announce Type: new Abstract: As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-off...
16. ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations ​
Author: Peiying Zhu, Sidi Chang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.MA
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked trans...
17. HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models ​
Author: Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers...
18. Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision ​
Author: Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng, Qianwen Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, e...
19. DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents ​
Author: Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving undere...
20. MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity ​
Author: Yiran Zhang, Jinwen Liu, Daniel Su, Yisu Chen, Qiang Sun, Chris Gonzalez, Eun-Jung Holden, Marco Fiorentini, Wei Liu, Yihao Ding
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded ex...
21. ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction ​
Author: Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MS, cs.SE
arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they ...
22. CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning ​
Author: Yongshi Ye, Tian Lan, Feihu Jiang, Muyang Ye, Bin Zhu, Qianghuai Jia, Longyue Wang, Zhao Xu, Weihua Luo, Xiaodong Shi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02074v1 Announce Type: new Abstract: Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evo...
23. Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems ​
Author: Yiran Zhao, Lu Zhou, Liming Fang, Yufei Chen, Jiafei Wu, Zhe Liu, Xiaogang Xu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis p...
24. MASkills: Continual Skills Optimization for Multi-Agent LLM Systems ​
Author: Huaiyuan Yao, Xiaoou Liu, Charles Fleming, Tianlong Chen, Hua Wei
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mostly hard to invoke,...
25. READY or Not: Reliable Enterprise Agent Deployment ​
Author: Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02095v1 Announce Type: new Abstract: An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent...
26. Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics ​
Author: Jiani He, Dingyan Shang, Yihua Xu, Shiqi Huang, Yan Lyu, Jize Li, Shangjing Tang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02116v1 Announce Type: new Abstract: Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition fa...
27. Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents ​
Author: Jalal Mahmud
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2609.02129v1 Announce Type: new Abstract: Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data objects relevant to a task. Yet successful discovery outcomes are typically discarded rather than reused. We introduce persistent discovery conte...
28. EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision ​
Author: Ziyuan Jin, Yuxuan Ge, Zheng Tian
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02133v1 Announce Type: new Abstract: Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distrib...
29. FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs ​
Author: Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02168v1 Announce Type: new Abstract: Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol...
30. Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning ​
Author: Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang, Abhiram Gudimella, Kevin Zhu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02191v1 Announce Type: new Abstract: Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the Me...
31. ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models ​
Author: Da Cheng Gu, Yifei Dong, Xinghao Yang, Yongshun Gong, Wei Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02215v1 Announce Type: new Abstract: Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefor...
32. PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion ​
Author: Yunchi Yang, Longlong Li, Cunquan Qu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02216v1 Announce Type: new Abstract: Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during training, requiring models to learn transferable relational and structural patterns. Existing subgraph- and path-based approaches often encode re...
33. SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams ​
Author: Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02217v1 Announce Type: new Abstract: LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon worklo...
34. PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment ​
Author: Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02231v1 Announce Type: new Abstract: Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. ...
35. PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks ​
Author: Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02236v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit a...
36. Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality ​
Author: Yifan Zhu, Sammie Katt, Samuel Kaski
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.MA
arXiv:2609.02242v1 Announce Type: new Abstract: AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any pr...
37. Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training ​
Author: Jian Gao, Xiao Zhang, Xun Zhu, Miao Li, Ji Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02244v1 Announce Type: new Abstract: Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input contex...
38. LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails ​
Author: Vansh Wahi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2609.02246v1 Announce Type: new Abstract: Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position...
39. APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering ​
Author: Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02253v1 Announce Type: new Abstract: Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve ver...
40. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems ​
Author: Jinxi Yu, Yubei Li, Eric Hanchen Jiang, Zhi Zhang, Dong Liu, Wenxiao Zhao, Levina Li, Kai-Wei Chang, Ying Nian Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N ...
41. CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging ​
Author: Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02273v1 Announce Type: new Abstract: Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individ...
42. SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology ​
Author: Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02292v1 Announce Type: new Abstract: The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task leve...
43. Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds ​
Author: Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment...
44. SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning ​
Author: Zhao Ji, Wenqing Chen, Zhixuan Chu, Jianxing Yu, Jingping Liu, Shanhe Zhao, Zibin Zheng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address t...
45. Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions ​
Author: Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, Mao Yang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02371v1 Announce Type: new Abstract: With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding t...
46. Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks ​
Author: Xiang Yin, Nico Potyka, Antonio Rago, Francesca Toni
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In ...
47. UTP-Bench: Uncertainty-aware Travel Planning Benchmark ​
Author: Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02421v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic del...
48. CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI ​
Author: Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large...
49. Collective creativity in hybrid societies ​
Author: Mason Youngblood, Katie Mudd, Manuel Anglada-Tort, Cameron Jones, Elena Miu, Diana Omigie, Margaret Schedel
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.MA
arXiv:2609.02620v1 Announce Type: new Abstract: Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement co...
50. Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting ​
Author: Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2609.02649v1 Announce Type: new Abstract: Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tas...
51. Door-in-the-Face Requests and Refusal Behaviour in Large Language Models ​
Author: Til Jordan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02707v1 Announce Type: new Abstract: Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large...
52. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills ​
Author: Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen, Defu Lian, Zhicheng Dou, Chaozhuo Li, Qiwei Ye, Zheng Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.02749v1 Announce Type: new Abstract: Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how ...
53. Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems ​
Author: Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang, Weilin Luo, Jun Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02750v1 Announce Type: new Abstract: Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and ...
54. Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents ​
Author: Vasileios Rizeakos, Georgios Paisios, Alexandros Machairas, Michael Birbas, Athanasios Bachoumis
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02760v1 Announce Type: new Abstract: On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no ...
55. SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment ​
Author: Qinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan, Qingyu Liu, Yu Li, Guanxu Chen, Yanwei Fu, Xi Lin, Xia Hu, Dongrui Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2609.02786v1 Announce Type: new Abstract: The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety...
56. Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis ​
Author: Hao Zhou (Jianzhong), Mandar Kulkarni (Jianzhong), Hao Chen (Jianzhong), Yan Xin (Jianzhong), Charlie (Jianzhong), Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.02805v1 Announce Type: new Abstract: Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) of...
57. AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application ​
Author: Wenxin Jiang, Xuyang Wang, Yuxiao Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2609.02821v1 Announce Type: new Abstract: Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating w...
58. Discriminative World Models for Web Agents ​
Author: Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2609.02885v1 Announce Type: new Abstract: Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via sup...
59. Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus ​
Author: Mohammad Omar Khursheed, Mandira Sawkar, Ashiqur R. KhudaBukhsh
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG
arXiv:2608.30485v1 Announce Type: cross Abstract: The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 180...
60. WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling ​
Author: Zhongzheng Li, Qingsong Ran, Shikun Feng, Nian Ran, Wenhao Li, Xiaoyuan Zhang, Yue Wang, Xiaoguang Zhao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01608v1 Announce Type: cross Abstract: Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error ...
61. Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search ​
Author: Harish Saragadam, Sudhanshu Sharma, Meghana Pujari
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2609.01617v1 Announce Type: cross Abstract: Getting accurate, grounded answers out of large enterprise document repositories is a difficult problem. Dense vector retrieval alone frequently performs poorly on queries that mix technical terminology, vendor-specific acronyms, or require reasoning...
62. RecEvolve: A Knowledge-Driven Autonomous Agent System for Recommender Systems ​
Author: Weidi Pan, He Ma, Shuhao Ye, Palaksh Rungta, David McPeek, Junyi Jiao, Arnab Bhadury, Mingyan Gao, Onkar Dalal
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2609.01622v1 Announce Type: cross Abstract: The rise of agentic AI has catalyzed a shift toward self-iterating systems, opening new frontiers for the autonomous optimization of production recommender models. This paper presents the empirical validation of a knowledge-driven autonomous agent sy...
63. The Utility of LLMs in Recommender Systems Explanation Evaluation ​
Author: Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2609.01627v1 Announce Type: cross Abstract: Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing exp...
64. A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platforms ​
Author: Jingwen Luo, Pinrui Zhu, Yiyan Wang, Zilin Xiao, Jingqi Li, Xuebei Kong, Yan Zhan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SI, cs.AI
arXiv:2609.01649v1 Announce Type: cross Abstract: With the rapid growth of live-streaming e-commerce and digital marketing, abnormal marketing behaviors have become increasingly concealed and coordinated across heterogeneous modalities, challenging platform governance and early risk identification. ...
65. From Feature Interaction to Feature Transport - A Unified Block for Scalable Recommendation Models ​
Author: Zichen Luo, Jiachen Guo, Keming Gu, Jie Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2609.01655v1 Announce Type: cross Abstract: Unified recommendation models aim to jointly model non-sequential multi-field features and sequential user behaviors, but existing interaction-centric designs mainly focus on mixing heterogeneous tokens within each layer. We argue that scalable unifi...
66. NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference ​
Author: Aur'elien Lac, Tony Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CV
arXiv:2609.01657v1 Announce Type: cross Abstract: Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models ...
67. PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation ​
Author: MinKeon Kim, Namjun Lee, Jaekwang Kim
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.01658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation, where early retrieval failures confound subsequent steps. Standard outcome-ba...
68. How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making ​
Author: Shubhra Mittal
Published: 9/3/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI
arXiv:2609.01660v1 Announce Type: cross Abstract: Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to...
69. Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration ​
Author: Zekai Jin, Hanrong Zhang, Yihong Tang, Fei Hu, Zhen Dong, Yi Shao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2609.01662v1 Announce Type: cross Abstract: For embodied systems, predictive agreement alone does not determine whether evidence warrants action; evidential origin matters. Repeated inference over one observation can multiply agreement without adding evidence, while source-local values do not ...
70. Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports ​
Author: Safayat Bin Hakim, Houbing Herbert Song
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2609.01671v1 Announce Type: cross Abstract: Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stat...
71. CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction ​
Author: Kewei Li, Rongying Zhang, Peiyu Yang, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Fengfeng Zhou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.BM
arXiv:2609.01673v1 Announce Type: cross Abstract: Activity-cliff ranking remains difficult because local structural changes can cause large activity differences, while high-quality data that resolve the underlying mechanisms remain limited. To use available activity labels more effectively, we combi...
72. Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study ​
Author: Arpan Kumar Mahapatra
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2609.01693v1 Announce Type: cross Abstract: Safety properties assessed separately for Model Context Protocol (MCP) tool use and Agent2Agent (A2A) delegation need not describe behavior when one agent uses both. We measure one such behavior in a single controlled MCP-to-A2A configuration: a test...
73. RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis ​
Author: Amirhosein Azarpour
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01729v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace the fixed scalar weights of a standard network with learnable univariate functions on each edge, but existing variants still fix the \emph{basis} that those functions are built from: B-splines, Chebyshev pol...
74. HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation ​
Author: Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos, Giuseppe Ateniese, Emanuele Rodol`a
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2609.01730v1 Announce Type: cross Abstract: Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplicat...
75. Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives ​
Author: Haibo Jin, Suijin Wang, Xucheng Yu, Haojing Luo, Haohan Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG, cs.MA
arXiv:2609.01736v1 Announce Type: cross Abstract: Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by...
76. Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics ​
Author: Sejuti Basu, Ashima Sood, Vijay Kumar, Sahil Sharma
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.01749v1 Announce Type: cross Abstract: Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult,...
77. Dictionary-Guided Mutation Operators for Automated HDL Repair ​
Author: Maisha Mastora, Dean Sullivan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, cs.AR, cs.NE
arXiv:2609.01775v1 Announce Type: cross Abstract: Automated repair of Hardware Description Language (HDL) designs remains challenging due to the large search space of candidate repairs and the strict syntactic and semantic constraints imposed by HDL grammars. Generic mutation strategies overwhelming...
78. Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks ​
Author: Hatim Chergui, Carolina Fern'{a}ndez-Mart'{i}nez, Mehdi Bennis, Merouane Debbah
Published: 9/3/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.MA
arXiv:2609.01779v1 Announce Type: cross Abstract: Future 6G networks will rely on Large Language Model (LLM) agents to manage the Radio Access Network (RAN). However, current architectures assume inter-agent messages convey objective facts. A message is instead a \emph{trace} of the sender's reasoni...
79. VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages ​
Author: Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.01788v1 Announce Type: cross Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource langua...
80. hLLM: Single Pass Decoding for Generative Reranking ​
Author: Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2609.01807v1 Announce Type: cross Abstract: Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker...
81. Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory ​
Author: Gene Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2609.01818v1 Announce Type: cross Abstract: The browser has become a first-class database host: applications increasingly want to store, query, and reason over structured data entirely on the client - for privacy, offline operation, local-first collaboration, and, most recently, as durable mem...
82. Interpretable Symptom Vectors for Depression in a Large Language Model ​
Author: Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, q-bio.NC
arXiv:2609.01832v1 Announce Type: cross Abstract: Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient s...
83. Agent Memory Is a Surface for Endogenous Authorization Laundering ​
Author: Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2609.01836v1 Announce Type: cross Abstract: Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrictions, and revocations. When memory misrepresents this evolving authorization state, the agent's own records can grant authority that ...
84. Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge ​
Author: Chen Chen, Mohsen Nayebi Kerdabadi, Dongjie Wang, Mei Liu, Zijun Yao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01839v1 Announce Type: cross Abstract: Longitudinal prediction from electronic health records (EHRs) is limited by the sparsity and irregularity in patient trajectories, and knowledge augmentation with external knowledge graphs (KGs) offers a promising way to alleviate these issues. Howev...
85. Thinking effort aligns between humans and reasoning models in abductive reasoning ​
Author: Henry Arthur
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.01867v1 Announce Type: cross Abstract: A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learnin...
86. OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation ​
Author: Yunqin Zhu, Feng Qiu, Yao Xie
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01896v1 Announce Type: cross Abstract: Power-outage planning requires scenarios before an event occurs. These scenarios must represent uncertainty in magnitude, timing, and duration while preserving temporal dependence. However, severe events are rare, and data from any single region cont...
87. Accurate in space, unreliable in time: how LLMs represent national cultural change ​
Author: Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2609.01902v1 Announce Type: cross Abstract: Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model repres...
88. Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens ​
Author: Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.01936v1 Announce Type: cross Abstract: A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix)...
89. On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers ​
Author: Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01947v1 Announce Type: cross Abstract: Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's obse...
90. Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network ​
Author: Lucas Qingyang Fang, Tiyao Liu, Jinhao Jing, Zeji Li, Kaijie Chen, Harikrishna Kuttivelil, Katia Obraczka
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01952v1 Announce Type: cross Abstract: Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) excha...
91. Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight ​
Author: Xinyu Fu, Narayan Ramasubbu, Dennis Galletta
Published: 9/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2609.01976v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alt...
92. InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation ​
Author: Vanshika Vats, Ashwani Rathee, James Davis
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02002v1 Announce Type: cross Abstract: Guideline-consistent semantic segmentation requires more than category recognition, as real-world labeling policies demand fine-grained, task-specific decisions. Recent multi-agent refinement systems improve compliance with such textual guidelines by...
93. InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation ​
Author: Ziquan Liu, Zhewei Zhu, Xuyang Shi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02004v1 Announce Type: cross Abstract: Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based s...
94. Seed-Anchored Budget-Bounded Graph Rendering for Question Answering on Industry-Standard Power-Grid Information and Exchange Models ​
Author: Jayakumar Manoharan, Yamini Sehgal
Published: 9/3/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.IR, cs.SY
arXiv:2609.02011v1 Announce Type: cross Abstract: Large language model question answering over power-grid models must respect a fixed context budget. We introduce seed-anchored graph rendering, a deterministic method that prioritizes query-local graph evidence without adding method-specific tuned or...
95. Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation ​
Author: Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2609.02046v1 Announce Type: cross Abstract: Monolithic world models predict the entire next state at every step, spending capacity re-predicting the static majority of a scene and injecting error into it. We ask whether explicitly modeling change (a per-object change gate plus a residual delta...
96. Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models ​
Author: Tianqi Xiao, Shiyao Cui, Minghao Zhang, Junxiao Yang, Renmiao Chen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CL, cs.CR
arXiv:2609.02082v2 Announce Type: cross Abstract: Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this cross-modal safety drift and ou...
97. Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts ​
Author: Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2609.02101v1 Announce Type: cross Abstract: Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this c...
98. Git4Data: Database-Native Version Control for AI Agents ​
Author: Hongshen Gou, Zuyu Zhang, Yuze Sun, Peng Xu, Feng Tian, Long Wang, Jianguo Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2609.02106v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly explore many candidate states of relational data in parallel, each of which should remain isolated, reproducible, and auditable, preferably through the same SQL interface used for ordinary data work. Exi...
99. Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models ​
Author: Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02108v1 Announce Type: cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit infilling tasks, which require generating a middle span conditioned on...
100. MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs ​
Author: Youssef Ennouri, Soonhoi Ha
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2609.02109v1 Announce Type: cross Abstract: Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially wi...
101. Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap ​
Author: Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh, Jahidul Arafat, Sunil Kumar Gaire
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2609.02111v1 Announce Type: cross Abstract: Dermatology artificial intelligence (AI) models are predominantly trained on light-skinned, cancer-focused image collections, yet they are increasingly proposed for deployment in resource-constrained settings where patients differ from training popul...
102. text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation ​
Author: Ritesh Kumar
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2609.02115v1 Announce Type: cross Abstract: Natural language interfaces to databases have traditionally suffered from three structural limitations: exclusive targeting of relational SQL, unconditional dependence on large language model (LLM) inference at query time, and absence of any runtime ...
103. C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees ​
Author: S M Rafiuddin, Atriya Sen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02131v1 Announce Type: cross Abstract: Sentiment in social-media threads does not only vary across posts; it shifts as users react to claims, corrections, evidence, and hostility within a branching reply tree. We study why sentiment changes in rumor-centric conversation trees by treating ...
104. A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search ​
Author: Sajad Faghfoor Maghrebi, Navid Eslami, Niv Dayan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.IR, cs.LG
arXiv:2609.02143v1 Announce Type: cross Abstract: Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scal...
105. Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor ​
Author: Vaneet Aggarwal, Yiyang Lu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CC, stat.ML
arXiv:2609.02145v1 Announce Type: cross Abstract: We study online maximization of nonnegative, non-monotone DR-submodular functions over compact convex down-closed subsets of the $d$-dimensional unit cube. The best known constructive offline approximation factor is $0.401$ under the corresponding me...
106. OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations ​
Author: Yixiong Xiao, Lang An, Hucheng Yang, Pinxue Ma, Yongquan Chen, Jingjia Cao, Yusai Zhao, Ting Wang, Ting Liu, Siqi Bao, Jingbo Zhou, Hua Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2609.02149v2 Announce Type: cross Abstract: Large language models (LLMs) are increasingly evolving from conversational assistants into agents capable of operating external digital environments. Graphical user interface (GUI) agents play an important role in this transition, as many real-world ...
107. Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation ​
Author: Jialin Liu, Zhaorui Zhang, Ray C. C. Cheung
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2609.02152v1 Announce Type: cross Abstract: Multimodal Recommender Systems (MRSs) typically rely on a flawed "modality harmony" assumption, presuming that multimodal features are inherently beneficial and strictly aligned with users' collaborative interaction patterns. However, modality-topolo...
108. OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction ​
Author: Jaehoon Jeong, Jay-Yoon Lee
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02158v1 Announce Type: cross Abstract: Legal Judgment Prediction (LJP) models are typically trained on documents that describe facts from a prosecutorial perspective. Existing datasets further exhibit severe label imbalance toward guilty outcomes. Consequently, these models suffer from "G...
109. GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories ​
Author: Arpita Joshi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02160v1 Announce Type: cross Abstract: Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes...
110. Schr\"odinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation ​
Author: Shizhe Zhang, Mingyang Zhao, Lei Ma
Published: 9/3/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2609.02196v1 Announce Type: cross Abstract: Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean data, repeated ambient projection, and coordinate inconsistency in Euclidean representations. Schrodinger bridges provide a probabilistic gene...
111. SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework ​
Author: Fang He, Wang-chien Lee
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02203v1 Announce Type: cross Abstract: Time series representation learning (TSRL) has attracted growing research interests in recent years. Two recent explorations in TSRL are: i) exploiting a transformer-based framework to learn time series; ii) instead of using only the targeted dataset...
112. Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging ​
Author: Khawaja Murad ul Hassan, Mehran Ebrahimi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02224v1 Announce Type: cross Abstract: Post-hoc saliency maps such as Grad-CAM are increasingly used to audit why a deployed vision model made a decision, yet the heatmap drifts when the input is rotated, even when the prediction is unchanged. In domains with no canonical orientation, suc...
113. InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models ​
Author: Chengyin Hu, Dingyi Lu, Jiaju Han, Xiang Chen, Weiwen Shi, Jiahuan Long, Yiwei Wei, Jiujiang Guo
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02233v1 Announce Type: cross Abstract: Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch metho...
114. SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation ​
Author: Qin Lu, Zheyang Jing, Yujie Yang, Jianwang Li, Chen Yi, Shaofeng Jiang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02247v1 Announce Type: cross Abstract: Semi-supervised learning has shown great potential for reducing annotation costs in medical image segmentation. However, most existing methods mainly exploit unlabeled data through prediction-level consistency, while the reliability of internal featu...
115. DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space ​
Author: Steffen Hagedorn, Aron Distelzweig, Alexandru P. Condurache
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2609.02252v1 Announce Type: cross Abstract: In trajectory planning for autonomous driving, hybrid planning architectures are often realized as a collection of disparate modules, each with its own objectives. This lack of a unifying principle can lead to inconsistencies between the initial and ...
116. Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics ​
Author: Yijie Lin, Ching-Chun Chang, Isao Echizen, Hui Li, Chin-Chen Chang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV, cs.MM
arXiv:2609.02268v1 Announce Type: cross Abstract: With the rapid proliferation of generative models on Machine Learning as a Service (MLaaS) platforms, reliably tracing the provenance of synthetic media without modifying generator architectures or parameters remains a major challenge. In this work, ...
117. CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation ​
Author: Shucheng Zhang, Yuang Zhang, Bingzhang Wang, Muhammad Monjurul Karim, Kehua Chen, Yinhai Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2609.02270v1 Announce Type: cross Abstract: Generating safety-critical scenarios is essential for evaluating autonomous driving systems. However, existing generators primarily focus on inducing collisions and offer limited control over where contact occurs on the target vehicle. In this paper,...
118. PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation ​
Author: Yunhao Liu, Hong Phuc Pham, Jaehong Yoon
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SE
arXiv:2609.02272v1 Announce Type: cross Abstract: Faithfully translating research papers into repository-level implementations remains challenging because papers often describe methods at a high level, leave implementation assumptions implicit, and require generated repositories to preserve method l...
119. Do Large Language Models Capture the Diversity in their Training Data? ​
Author: Youqi Wu, Farzan Farnia
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.02275v1 Announce Type: cross Abstract: Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an info...
120. Auditory Illusion Benchmark for Large Audio Language Models ​
Author: Hayoon Kim, Eunice Hong, Kyogu Lee
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2609.02277v1 Announce Type: cross Abstract: Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replica...
121. RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification ​
Author: Jierui Li, Zhiyuan Qi, Hao Zhu, Yufan Liu, Jixian Liu, Shaojie Jiang, Jianda Wang, Yaqi Liu, Xiaotong Li, Wei Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02282v1 Announce Type: cross Abstract: Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (...
122. VoRTeC: Taming Foundation Flow for One-step Real time Video Compression ​
Author: Yichong Xia, Qinhong Wu, Bin Chen, Jinpeng Wang, Zeyuan Chen, Haoqian Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02291v2 Announce Type: cross Abstract: Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor te...
123. SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment ​
Author: Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2609.02293v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \text...
124. DiffIE: Diffusion-based Open Information Extraction ​
Author: Konstantin Fedorov, Valentin Malykh
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02315v1 Announce Type: cross Abstract: A single sentence often expresses multiple valid relational triplets, which makes Open Information Extraction (OpenIE) fundamentally a multi-output task. Existing neural systems handle this by autoregressive generation, which is flexible but slow and...
125. What Is Worth Representing? Representational Empowerment for Continual Model Construction ​
Author: Fei Dai, Hanqi Zhou, Alison Gopnik, Charley Wu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02322v1 Announce Type: cross Abstract: The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this problem as continual model construction: an agent maintains an environment-specific...
126. ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans ​
Author: Amirhosein Azarpour
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02333v1 Announce Type: cross Abstract: Brain cancer remains one of the most significant challenges in modern medicine, where the accuracy of early stage diagnosis is a decisive factor in patient survival and treatment efficacy. Although Magnetic Resonance Imaging (MRI) is the established ...
127. AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers ​
Author: Alexey Potapov
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02339v1 Announce Type: cross Abstract: World modeling requires a predictive model to maintain and update an internal state adequate for reasoning about the consequences of actions. We introduce the AGI Maze Prediction Datasets and Benchmark, a lightweight controlled testbed for studying t...
128. Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information ​
Author: Zhen Zhou, Jiachen Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen
Published: 9/3/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI
arXiv:2609.02344v1 Announce Type: cross Abstract: Existing cell embedding methods predominantly rely on transcriptomic or proteomic measurements and represent each cell as a holistic entity, thereby overlooking the subcellular localization of individual molecules. Moreover, they rarely incorporate p...
129. Fair Stable Matching: A Nash Social Welfare Approach ​
Author: Parth Desai, Rasheed M, Ganesh Ghalme, Sujit Gujar
Published: 9/3/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2609.02354v1 Announce Type: cross Abstract: While traditional stable matching algorithms, such as the Gale-Shapley algorithm, prioritize stability, they may fall short of achieving equitable outcomes among participants. We study the role of \emph{Nash social welfare} (NSW) as a fairness object...
130. Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions ​
Author: Maitreyee Tewari, Michele Persiani
Published: 9/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2609.02364v2 Announce Type: cross Abstract: Existing Human-Robot Interaction (HRI) literature has focused on identifying and structuring errors, failures, conflicts, and knowledge issues (called in this work as contradictions) in domain-specific dialogue-based interactions. However, there is s...
131. NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning ​
Author: Meixuan Chen, Hehan Li, Ruizhi Zhao, Xin Lu, peizhi xu, Liwei Qian, LI Meifang, shuanglong li, Hanmeng Liu, Xin Pei, Yanbiao Ma
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02366v1 Announce Type: cross Abstract: Named Entity Recognition (NER) has achieved substantial progress since the advent of large language models (LLMs). Nevertheless, the recognition of long-tail and domain-specific entities remains challenging due to the deficiency in parametric knowled...
132. Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance ​
Author: Sai Niranjan Ramachandran, Suvrit Sra
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cond-mat.stat-mech, cs.AI
arXiv:2609.02373v1 Announce Type: cross Abstract: We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by mo...
133. MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts ​
Author: Matteo Greco, Anudeex Shetty, Andrea Tagarelli, Jey Han Lau
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.DL, cs.IR
arXiv:2609.02379v1 Announce Type: cross Abstract: While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relativel...
134. PolERo: Studying Political Evasion in Romanian ​
Author: Gabriel Stefan, Sergiu Nisioi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02391v1 Announce Type: cross Abstract: Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion s...
135. Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts ​
Author: Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02404v1 Announce Type: cross Abstract: Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, sug...
136. Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking ​
Author: Siyu Chen, Haoran Wang, Xiaojian Li, Yao Huang, Yinpeng Dong, Wei Xu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02414v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a facto...
137. Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment ​
Author: Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02417v1 Announce Type: cross Abstract: Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the ...
138. Towards One-for-All Robustness Across a Continuum of Threat Levels ​
Author: Zhichao Hou, Xiaorui Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02440v1 Announce Type: cross Abstract: Adversarially robust models often overfit to a specific attack budget, necessitating multiple specialized models for diverse and dynamic adversarial environments, a strategy that becomes fundamentally intractable as the threat space grows. This raise...
139. Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression ​
Author: Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2609.02451v1 Announce Type: cross Abstract: In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasibl...
140. Addressing Trust in AI Systems through Education: A Didactic Perspective ​
Author: Pierre Haritz, Hendrik Krone, Thomas Liebig
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2609.02453v1 Announce Type: cross Abstract: Machine learning (ML) education faces two persistent and connected obstacles: many educational tools present ML as an opaque black box, which leaves learners with a superficial understanding, and this same opacity prevents users from forming the cali...
141. DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models ​
Author: Yotam Eshel, Guy Hadad, Guy Feigenblat, Yuri M. Brovman, Matt Gearhart, Bracha Shapira
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02468v1 Announce Type: cross Abstract: We explore predicting eCommerce user preferences for product aspects such as brand, size, and color - a task we define as Aspect Affinity. Solving this task improves customer understanding and enables fine-grained personalization in recommendation, s...
142. ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering ​
Author: Adrien Mialland, Marc Plantevit, Julien Gallois, C'eline Robardet
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.CV
arXiv:2609.02486v2 Announce Type: cross Abstract: Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Langua...
143. RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection ​
Author: Taufikur Rahman Fuad, Md Abrar Jahin, Amir Hussain
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02497v1 Announce Type: cross Abstract: Zero-shot graph anomaly detection seeks to deploy a detector trained on source graphs to unseen, unlabeled targets, yet domain shift can make source-derived notions of normality unreliable. We introduce RINSE (Robust Iterative Normality Self-Estimati...
144. Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models ​
Author: Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02502v1 Announce Type: cross Abstract: Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, re...
145. Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion ​
Author: Md Abrar Jahin, Taufikur Rahman Fuad, Jay Pujara, Craig A. Knoblock
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02519v1 Announce Type: cross Abstract: Uncertain knowledge graphs (UKGs) extend knowledge graphs by assigning each triple a continuous confidence score. Since most possible triples lack observed confidences, recent methods rely on semi-supervised learning to generate pseudo-labels. These ...
146. Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework ​
Author: Yan Zhong, Gefei Chen, Qiufang Ma, Zhen Wang, Zhiwei Fan, Lei Shi, Tingting Jiang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02529v1 Announce Type: cross Abstract: Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--...
147. Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs ​
Author: Xixiang He, Xingming Li, Baiqi Wu, Qiyao Sun, Xuanyu Ji, Ao Cheng, Qingyong Hu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02548v1 Announce Type: cross Abstract: Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose ...
148. ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction ​
Author: Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.02549v1 Announce Type: cross Abstract: Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular p...
149. Competitive Market Behavior of LLMs ​
Author: Pawel Struski, Jakub Swistak, Inez Okulska, Przemyslaw Biecek
Published: 9/3/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, econ.GN, q-fin.EC
arXiv:2609.02580v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes wh...
150. Automated Vulnerability Injection in Smart Contracts Using Large Language Models ​
Author: Luca Migliaccio, Roberto Natella, Naghmeh Ivaki, Nuno Laranjeiro, Marco Vieira
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR
arXiv:2609.02624v1 Announce Type: cross Abstract: Assessing vulnerability detection tools for smart contracts requires datasets with known ground truth, yet such datasets are scarce and difficult to build by hand. We propose an approach that uses Large Language Models (LLMs) to automatically inject ...
151. TaRA: Training-Aware Low-Rank Adaptation Initialization ​
Author: Taehyeon Kim, Eunhyeok Park
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.02639v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches a...
152. From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs ​
Author: Urja Pawar, Rajitha Ramanayake, Owen O'Neill, Nabeel Kemal, Abhishek Mandal, Houssem Chatbri, Christopher Martin
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02679v1 Announce Type: cross Abstract: When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no trusted context or reference document is available, we study two signals ...
153. DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models ​
Author: Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi, Gaurav Pandey, Sonam Gupta, Vineet Kumar, Jaydeep Sen, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.02685v1 Announce Type: cross Abstract: RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading...
154. RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models ​
Author: Canjie Liu, Jiawen Kang, Jinbo Wen, Zishao Zhong
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2609.02731v1 Announce Type: cross Abstract: Large vision-language models have achieved remarkable success in vision-language tasks. However, they remain prone to Visual Hallucinations (VHs), undermining their reliability in real-world applications. Existing solutions typically require curated ...
155. Language Models Can Control Their Own Attention ​
Author: Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.02737v1 Announce Type: cross Abstract: Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan...
156. HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design ​
Author: Ge Sun, Gervasio Zaldivar, Yuan Tian, Gustavo Perez Lemus, Juhae Park, Dasha Safarian, Ming Han, Juan J. de Pablo
Published: 9/3/2026, 4:00:00 AM
Categories: physics.chem-ph, cond-mat.mtrl-sci, cs.AI, cs.LG, physics.comp-ph
arXiv:2609.02746v1 Announce Type: cross Abstract: Polymeric materials are central to modern technologies, with applications ranging from energy to health and transportation. Although AI has made significant advances in materials discovery, the hierarchical structure of polymers across multiple lengt...
157. Untangling the Mechanisms of Misleading Context in Medical Question Answering ​
Author: Robin Linzmayer, No'emie Elhadad
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.02754v1 Announce Type: cross Abstract: Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupt...
158. From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution ​
Author: Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2609.02771v1 Announce Type: cross Abstract: Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under...
159. Dutch Books for Language Models ​
Author: Isaiah Andrews, Suproteem Sarkar
Published: 9/3/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.CL, cs.LG, q-fin.EC
arXiv:2609.02797v1 Announce Type: cross Abstract: People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that ...
160. frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study ​
Author: Onur U\u{g}urlu (.Izmir Bak{\i}r\c{c}ay University)
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DM, cs.AI
arXiv:2609.02804v1 Announce Type: cross Abstract: For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record had stood at 99 of 100 variables. We give a directly checkable 100-vertex independent set for its 4,000-vertex graph. Together with a v...
161. Post-Training Language Models for Gold-Medal Performance in Coding Competitions ​
Author: Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, Boris Ginsburg
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.MA, cs.SE
arXiv:2609.02849v1 Announce Type: cross Abstract: Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale p...
162. Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework ​
Author: Cagri Temel
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2609.02861v1 Announce Type: cross Abstract: Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions. This paper presents TRACE (Transparent Reasoning Architecture for Cre...
163. When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning ​
Author: Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2505.15276v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to "overthink" leads to inefficiencies. Although "save-thinking" prompts are intended to mitigate this issue, we find that LRMs still frequently ent...
164. Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy ​
Author: Saleh Afzoon, Ali Shahsavandi, Phuong Thao Huynh, Melika Zare, Zahra Jahanandish, Amin Beheshti, Usman Naseem
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2505.21907v3 Announce Type: replace Abstract: AI copilots represent a new generation of AI-powered systems designed to assist users, particularly knowledge workers and developers, in complex, context-rich tasks. As these systems become more embedded in daily workflows, personalization has emer...
165. AI Mathematician: Towards Fully Automated Frontier Mathematical Research ​
Author: Yuanhang Liu, Yanxing Huang, Yanqiao Wang, Peng Li, Yang Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2505.22451v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have made significant progress in mathematical capabilities in recent times. However, these successes have been primarily confined to competition-level problems. In this work, we propose AI Mathematician (AIM) framewor...
166. Achieving Olympiad-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning ​
Author: Haiteng Zhao, Junhao Shen, Yiming Zhang, Songyang Gao, Kuikun Liu, Tianyou Ma, Fan Zheng, Dahua Lin, Wenwei Zhang, Kai Chen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.10534v4 Announce Type: replace Abstract: Large language model (LLM) agents exhibit strong mathematical problem-solving abilities and can even solve International Mathematical Olympiad (IMO) level problems with the assistance of formal proof systems. However, due to weak heuristics for aux...
167. Stepwise Think-Critique: Interleaved Reasoning and Self-Critique in a Single LLM ​
Author: Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.15662v4 Announce Type: replace Abstract: Human beings solve complex problems through critical thinking, where reasoning and evaluation are intertwined to converge toward correct solutions. However, most existing large language models (LLMs) treat the reasoning and verification as separate...
168. What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? ​
Author: Basile Terver, Tsung-Yen Yang, Jean Ponce, Adrien Bardes, Yann LeCun
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.RO, stat.ML
arXiv:2512.24497v4 Announce Type: replace Abstract: A long-standing challenge in AI is to develop agents capable of solving a wide range of physical tasks and generalizing to new, unseen tasks and environments. A popular recent approach involves training a world model from state-action trajectories ...
169. Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents ​
Author: Miao Su, Yucan Guo, Zhongni Hou, Long Bai, Zixuan Li, Yufei Zhang, Guojun Yin, Wei Lin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.07468v2 Announce Type: replace Abstract: Memory enables Large Language Model (LLM) agents to perceive, store, and use information from past dialogues, which is essential for personalization. However, existing methods fail to properly model the temporal dimension of memory in two aspects: ...
170. Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories ​
Author: Ya Gao, Kalle Kujanp"a"a, Pekka Marttinen, Harri Valpola, Alexander Ilin
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.02028v3 Announce Type: replace Abstract: Enabling artificial intelligence systems, particularly large language models, to update knowledge and flexibly apply it during reasoning remains a central challenge. Existing knowledge editing approaches emphasize atomic facts, improving factual re...
171. TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning ​
Author: Christian Greisinger, Steffen Eger
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2603.03072v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist scientists across diverse workflows. A key challenge is generating high-quality figures from textual descriptions, often represented as TikZ programs that can be rendered as scientific im...
172. FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization ​
Author: Haijian Lu, Wei Wang, Jing Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.19828v4 Announce Type: replace Abstract: Autoformalization aims to produce formal statements that compile and faithfully preserve the intended meaning of informal mathematics. Yet standard single-output evaluation collapses this many-to-many structure into a single prediction. For downstr...
173. BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA ​
Author: Taeyun Roh, Suhyeong Park, Dongyoung Lee, Eunyeong Jo, Wonjune Jang, Junha Jung, Jaewoo Kang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.28026v3 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs). However, because the MCQA format incorporates the candidate choices into the input context...
174. From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems ​
Author: Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K"oster, Sven Hallerbach
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2604.02198v2 Announce Type: replace Abstract: While Artificial Intelligence (AI) offers transformative potential for operational performance, its deployment in safety-critical domains such as aviation requires strict adherence to rigorous certification standards. Current EASA guidelines mandat...
175. UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents ​
Author: Yijuan Liang, Xinghao Chen, Yifan Ge, Ziyi Wu, Hao Wu, Changyu Zeng, Wei Xing, Xiaoyu Shen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.11557v3 Announce Type: replace Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the str...
176. Unifying biomedical knowledge in a modern multimodal graph ​
Author: Lucas Vittor, Ayush Noori, I~naki Arango, Joaqu'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.27269v3 Announce Type: replace Abstract: Biomedical knowledge graphs (KGs) are widely used in the life sciences, yet many are derived from unstructured documents and therefore lack schema-level constraints, whereas graphs assembled from structured resources are difficult to harmonize into...
177. Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework ​
Author: Ali \c{S}enol, Garima Agrawal, Huan Liu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.24661v4 Announce Type: replace Abstract: Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficie...
178. OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation ​
Author: Haochen Yang, Ke Zhao, Mengyuan Ma, Xingyu Lu, Xiangfeng Wang, Hong Qian
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2605.29829v2 Announce Type: replace Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient paradigm for automated optimization. However, existing methods still exhibit limited generalization...
179. From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds ​
Author: Louis Nisiotis, Aimilios Hadjiliasi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2606.03557v2 Announce Type: replace Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge. Users interact through in-world interfaces in multimodal ways, yet their requests demand fundamentally different AI backend models and computatio...
180. Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules ​
Author: Wei Xu, Ke Yang, Gang Luo, Keli Zheng, Lingyan Hu, Jing Wang, Kefeng Li
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG
arXiv:2606.16337v4 Announce Type: replace Abstract: Predictive modeling for clinical decision support requires both strong predictive performance and transparent, auditable, and human-reviewable decision logic. Although deep learning and tree-based ensemble methods can achieve high accuracy, their b...
181. EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent ​
Author: Zeyao Du, Tong Li, Haibo Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.17698v3 Announce Type: replace Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose...
182. Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability? ​
Author: Ayan Antik Khan, Harsh Kohli, Yuekun Yao, Huan Sun, Ziyu Yao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.24026v2 Announce Type: replace Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) ag...
183. Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models ​
Author: MY Pitsane, Hope Mogale
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2607.04562v3 Announce Type: replace Abstract: Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism f...
184. TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting ​
Author: Xiachong Lin, Du Yin, Arian Prabowo, Hao Xue, Wen Hu, Imran Razzak, Matthew Amos, Sam Behrens, Flora D. Salim
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.06349v2 Announce Type: replace Abstract: Building sensors are embedded in physical topology, spatial hierarchy, and operational context, yet existing forecasters often treat them as isolated time series or rely on fixed covariate sets. We present TopoBrick, a training-free framework for z...
185. CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models ​
Author: Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.10110v2 Announce Type: replace Abstract: Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, l...
186. Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions ​
Author: Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, Hongyu Luo, Wun Yu Chan, Wei Fan, Haoran Li, Yangqiu Song
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.12406v2 Announce Type: replace Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behav...
187. SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction ​
Author: Xue Yu, Bo Yuan, Kailin Zhao, Pengshuai Yang, Hong Hu, Junlan Feng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15550v3 Announce Type: replace Abstract: Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks because a single erroneous action can lead to irreversible consequences. Existing safety me...
188. Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation ​
Author: Xiaohan Ye, Xu Chen, Zihan Gong, Jian Ding, Lianyu Du, Baicheng Chen, Yunmeng Shu, Jingqian Zhao, Zhixiang Zhao, Shuaiqi Jia, Chong Ma, Shuwen Xiao, Xiangheng Kong, Yuan Gao, Jun Song, Jinsong Lan, Xiaoyong Zhu, Bo Zheng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.17499v2 Announce Type: replace Abstract: The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to complex multimodal interactions that seamlessly combine product images, natural language descriptions, and m...
189. CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning ​
Author: El Hassane Ettifouri, Ayoub Belfatmi, Mahaman Sanoussi Yahaya Alassan, Walid Dahhane
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.20129v2 Announce Type: replace Abstract: Quantized small reasoning models can enter repetitive or otherwise unproductive trajectories, yet standard decoding does not adapt to the trajectory as it unfolds. We study MGT-B, a fixed, weight-preserving controller that converts overlapping wind...
190. Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models ​
Author: Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21617v2 Announce Type: replace Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend to rewrite it into...
191. Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance ​
Author: Andrei Starodubov, Yaqub Aris Prabowo, Andreas Hadjipieris, Roberto Galeazzi, Ioannis Kyriakides
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, cs.LG, cs.RO, cs.SY, eess.SP, eess.SY, math.IT
arXiv:2607.22667v2 Announce Type: replace Abstract: This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: in...
192. Adaptive Graph-of-Islands Evolution for Automatic Feature Engineering with LLMs ​
Author: Sha Li, Naren Ramakrishnan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23286v2 Announce Type: replace Abstract: Automatic feature engineering (AutoFE) for tabular data requires discovering informative transformations from a large program space. Existing approaches suffer from three limitations: classical methods rely on fixed operator libraries with limited ...
193. LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation ​
Author: Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.24780v2 Announce Type: replace Abstract: Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those observations into an evaluation process. To study this question, we i...
194. Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings ​
Author: Joseph Walusimbi, Ann Move Oguti, Abubakhari Sserwadda, Precious Boss Kasasira, Charles Brian Okoboi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-bio.OT
arXiv:2607.24814v2 Announce Type: replace Abstract: Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where physician-to-patient ratios can fall below 1:25,000 in rural settings. Existing AI-assisted diagnostic tools predominantly require reliable internet c...
195. PIE-APT: Abductive Planning over Temporal Dynamic Knowledge Graphs via Incremental Reasoning ​
Author: Amir Hossein Sharafi, Alireza Shahbazi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2607.27287v2 Announce Type: replace Abstract: Planning over Temporal Dynamic Knowledge Graphs (TDKGs) presents theoretical challenges in open-world environments with incomplete information. Existing action formalisms often face decidability issues and the Ramification Problem, while structural...
196. Nova: An End-to-End MLIR Compiler for Deep Learning ​
Author: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, cs.LG, cs.PL
arXiv:2608.00029v3 Announce Type: replace Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions, their execution mode...
197. FemWear: A Parameter-Efficient Wearable Foundation Model for Women's Health ​
Author: Yifan Wang, Chenzhong Li
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08244v2 Announce Type: replace Abstract: General-purpose wearable foundation models are pretrained on broad sensor streams and populations, but their representations are not organized around women's health. FemWear is a women's wearable foundation model, obtained by parameter-efficiently ...
198. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents ​
Author: Harshitha Kolukuluru, Reshma Ashok, Kirat Arora, Evan William Ciccarelli, Nischal Ashok Kumar, Lunyiu Nie, Franck Dernoncourt, Samyadeep Basu, Ryan A. Rossi, Nedim Lipka
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.MA
arXiv:2608.08389v2 Announce Type: replace Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher late...
199. FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training ​
Author: Josef Chen (Independent Researcher), Erim Hayretci (Imperial College London)
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.SE
arXiv:2608.20574v3 Announce Type: replace Abstract: We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model. We test 27 frontier large language model endpoints on 534 substitution, pairing and constraining tasks for tasks that re...
200. SKILL.state: Scalable Long-Horizon Agent Skills ​
Author: Sanket Badhe, Priyanka Tiwari, Jonghyun Chung
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.26263v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an...
201. Rating the Raters: Rasch Measurement Theory for LLM Evaluation ​
Author: Pratik S. Sachdeva, Nathan Boudol
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27463v2 Announce Type: replace Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probe...
202. When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems ​
Author: Yangxiao Jiang, Jiarun Fan, Mingcong Xu, Yanxi Guo, Jiwen Feng, Shanqing Xu, Mengchen Qian, Wei Chen, Xiaojin Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27984v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) have recently moved from static workflows toward dynamically generated collaboration topologies. However, existing topology generation methods rely primarily on the parametric knowledge of large language models, with exter...
203. Automated Researchers Can Mitigate Well-characterized Alignment Failures ​
Author: Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.28945v3 Announce Type: replace Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study ...
204. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling ​
Author: Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.29291v3 Announce Type: replace Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structur...
205. Can escalation channels redirect reward hacking toward defect disclosure? ​
Author: Francesca Gomez
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY
arXiv:2608.29460v2 Announce Type: replace Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-age...
206. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets ​
Author: Dongsheng Chen, Xiangyu Zhao, Xin Yao, Xuetao Wei
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2609.00015v2 Announce Type: replace Abstract: AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, tools, and execution backends operate over shared environments. In such settings, safety becomes a syste...
207. VoiceLongMemEval: Do Assistants Remember How You Sounded? ​
Author: Ramit Pahwa, Parivesh Priye, Apoorva Beedu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2609.00570v2 Announce Type: replace Abstract: With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue hist...
208. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs ​
Author: Seungwoo Jung, Dohyeok Kwon, Seungmin Cha, Junseok Lee, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2609.00575v2 Announce Type: replace Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representati...
209. Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate ​
Author: Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MM
arXiv:2609.01168v2 Announce Type: replace Abstract: Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and cross-moda...
210. FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue ​
Author: Hangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim, Jaeik Park, Hyunkyu Kim, Jungmin Son, Pilsung Kang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2609.01198v2 Announce Type: replace Abstract: Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or targeted ...
211. Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations ​
Author: Qingde Li, Qingqi Hong, Zihan Li, Jie Tian
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.GR
arXiv:2609.01408v2 Announce Type: replace Abstract: A fundamental challenge in artificial intelligence is the transformation of observations into explicit symbolic representations suitable for abstraction, interpretation, and reasoning. While modern AI systems achieve remarkable perceptual capabilit...
212. Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers ​
Author: Giovanni Bonetta, Matteo Merler, Davide Zago, Rossella Cancelliere, Bernardo Magnini
Published: 9/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2609.01567v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat syste...
213. Why we need an AI-resilient society- Profiling Large Language Models ​
Author: Thomas Bartz-Beielstein, Eva Bartz
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:1912.08786v5 Announce Type: replace-cross Abstract: Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn...
214. Deep denoising autoencoder-based non-invasive blood flow detection for arteriovenous fistula ​
Author: Li-Chin Chen, Yi-Heng Lin, Li-Ning Peng, Feng-Ming Wang, Yu-Hsin Chen, Po-Hsun Huang, Shang-Feng Yang, Yu Tsao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP
arXiv:2306.06865v2 Announce Type: replace-cross Abstract: Clinical guidelines underscore the importance of regularly monitoring and surveilling arteriovenous fistula (AVF) access in hemodialysis patients to promptly detect any dysfunction. Although phono-angiography/sound analysis overcomes the limi...
215. A Survey of Transformer-based Language Models with Focus on Efficiency ​
Author: Wazib Ansar, Saptarsi Goswami, Amlan Chakrabarti
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2406.16893v3 Announce Type: replace-cross Abstract: The emergence of Transformer-based Large Language Models (LLMs) has substantially augmented the capabilities of Natural Language Processing (NLP), thereby intensifying the demand for computational resources. Therefore, enhancing efficiency ba...
216. Doubly Stochastic Adaptive Neighbors Clustering via the Marcus Mapping ​
Author: Jinghui Yuan, Chusheng Zeng, Fangyuan Xie, Zhe Cao, Mulin Chen, Rong Wang, Feiping Nie, Yuan Yuan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2408.02932v3 Announce Type: replace-cross Abstract: Clustering is a fundamental task in machine learning and data science, and similarity graph-based clustering is an important approach within this domain. Doubly stochastic symmetric similarity graphs provide numerous benefits for clustering p...
217. Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations ​
Author: Garima Agrawal, Sashank Gummuluri, Cosimo Spera
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2410.10136v2 Announce Type: replace-cross Abstract: In customer contact centers, human agents often struggle with long average handling times (AHT) due to the need to manually interpret queries and retrieve relevant knowledge base (KB) articles. While retrieval augmented generation (RAG) syste...
218. Action abstractions for amortized sampling ​
Author: Oussama Boussif, L'ena N'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2410.15184v2 Announce Type: replace-cross Abstract: As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and ...
219. Nonasymptotic CLT and Error Bounds for Linear Two-Time-Scale Stochastic Approximation ​
Author: Seo Taek Kong, Sihan Zeng, Thinh T. Doan, R. Srikant
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.09884v4 Announce Type: replace-cross Abstract: We consider linear two-time-scale stochastic approximation algorithms driven by martingale noise. Recent applications in machine learning motivate the need to understand finite-time error rates, but conventional stochastic approximation analy...
220. Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English ​
Author: Jeremy Barnes, Naiara Perez, Alba Bonet-Jover, Bego~na Altuna
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they captur...
221. Multimodal Language Models as Text-to-Image Model Evaluators ​
Author: Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2505.00759v3 Announce Type: replace-cross Abstract: The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Mu...
222. No Data Wasted: A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels ​
Author: Yiyang Shen, Weiran Wang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2508.11180v2 Announce Type: replace-cross Abstract: Multi-view learning is widely applied to real-life datasets, but it often suffers from both missing views and missing labels. Prior probabilistic approaches addressed the missing view problem by using a product-of-experts scheme to aggregate ...
223. General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population ​
Author: Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2509.07330v3 Announce Type: replace-cross Abstract: Foundation models for healthcare require balancing robust generalization across heterogeneous clinical populations and disease settings with the architectural simplicity needed for deployment. We present a pre-trained model focused on demogra...
224. OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling ​
Author: Jihu Guo, Tenghui Ma, Wei Gao, Peng Sun, Xun Chen, Jiaxing Li, Zhisheng Ye, Yuyang Jin, Dahua Lin
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2509.23722v3 Announce Type: replace-cross Abstract: Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerbates pipeline bubbles, thereby reducing training efficiency. Prior approaches typically optimize a sing...
225. Fetch.ai: An Architecture for Modern Multi-Agent Systems ​
Author: Michael J. Wooldridge, Attila Bagoly, Jonathan J. Ward, Emanuele La Malfa, Gabriel Paludo Licks
Published: 9/3/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2510.18699v2 Announce Type: replace-cross Abstract: Recent surges in LLM-driven intelligent systems largely overlook decades of foundational multi-agent systems (MAS) research, resulting in frameworks with critical limitations such as centralization and inadequate trust and communication proto...
226. Inference-Time Optimization of Prompt Embeddings in Diffusion Models: A Comparison of sep-CMA-ES and Adam ​
Author: Dom'icio Pereira Neto, Jo~ao Correia, Penousal Machado
Published: 9/3/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2511.03913v3 Announce Type: replace-cross Abstract: Deep diffusion models have revolutionized image generation by producing high-quality outputs. However, achieving specific objectives with these models often requires costly adaptations such as fine-tuning, which can be resource-intensive and ...
227. SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning ​
Author: Tairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye, Haibo Hu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.09681v3 Announce Type: replace-cross Abstract: Visual reinforcement learning has achieved remarkable progress in visual control and robotics, but its vulnerability to adversarial perturbations remains underexplored. Most existing black-box attacks focus on vector-based or discrete-action ...
228. An Energy-Based Mechanism for Compositional Behavior ​
Author: Francesca Rossi, Veronica Centorrino, Francesco Bullo, Giovanni Russo
Published: 9/3/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.SY, eess.SY, nlin.AO
arXiv:2512.04745v4 Announce Type: replace-cross Abstract: Flexible intelligence relies on the ability to reuse previously acquired behaviors and combine them differently as circumstances change. In biological and artificial systems, this ability is often attributed to gating mechanisms that determin...
229. ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation ​
Author: Ziyang Mai, Yu-Wing Tai
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2512.07328v2 Announce Type: replace-cross Abstract: Text-to-video generation has advanced rapidly, yet preserving a character's holistic appearance from a single reference image remains challenging, particularly when the character undergoes large pose, motion, and scene changes. Existing refer...
230. Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation ​
Author: Yuxuan Qiao, Dongqin Liu, Hongchang Yang, Wei Zhou, Songlin Hu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2512.16310v4 Announce Type: replace-cross Abstract: LLM agents can combine individually non-revealing tool returns and disclose a sensitive conclusion, creating Tools Orchestration Privacy Risk (TOP-R). We formalize TOP-R through three conditions: conclusion sensitivity, single-source non-infe...
231. Beyond Transfer Accuracy: Mechanism-Guided Controlled Adaptation for Low-Resource Languages ​
Author: Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2601.08146v4 Announce Type: replace-cross Abstract: Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformers (CD-T) for unstructured settings via label-balanced activati...
232. Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks ​
Author: Candida M. Greco, Lucio La Cava, Andrea Tagarelli
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC, physics.soc-ph
arXiv:2601.22396v3 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain...
233. Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers ​
Author: Jiaji Zhang, Hailiang Zhao, Jiaju Wu, Ruichao Sun, Xinkui Zhao, Shuiguang Deng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2602.05605v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) are costly at high resolution because self-attention scales quadratically with token sequence length. Existing pruning methods do not jointly provide end-to-end learnability, low training overhead, and determinis...
234. FlatLands: Generative Floormap Completion From a Single Egocentric View ​
Author: Subhransu S. Bhattacharjee, Dylan Campbell, Rahul Shome
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO, eess.IV
arXiv:2603.16016v3 Announce Type: replace-cross Abstract: A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surroundings would better serve applications such as indoor navigation. We introduce FlatLands, a dataset and benc...
235. ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs ​
Author: Abhinaba Basu, Pavan Chakraborty
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2603.18579v2 Announce Type: replace-cross Abstract: Evaluating whether explanations faithfully reflect a model's reasoning remains an open problem. Existing benchmarks use single interventions without statistical testing, making it impossible to distinguish genuine faithfulness from chance-lev...
236. FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment ​
Author: Betty Xiong, Jillian Fisher, Benjamin Newman, Meng Hu, Shivangi Gupta, Yejin Choi, Lanyan Fang, Russ B Altman
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.19539v2 Announce Type: replace-cross Abstract: We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain ri...
237. Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model ​
Author: Jiahao Wu, Ning Lu, Shengcai Liu, Kun Wang, Yanting Yang, Baijiong Lin, Chen Jason Zhang, Li Qing, Ke Tang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.25184v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance performance, the computational overhead is a critical issue. In algo...
238. Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent ​
Author: Josef Hardi, Martin J. O'Connor, Marcos Martinez-Romero, Jean G. Rosario, Stephen A. Fisher, Mark A. Musen
Published: 9/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2604.08552v3 Announce Type: replace-cross Abstract: Descriptive scientific metadata in public repositories are often incomplete and inconsistent with community standards and ontologies, limiting data FAIRness. Large language models (LLMs) offer a promising approach to automatically standardizi...
239. On the Expressive Power and Limitations of Multi-Layer SSMs ​
Author: Nikola Zubi'c, Qian Li, Yuyi Wang, Davide Scaramuzza
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CC
arXiv:2604.14501v2 Announce Type: replace-cross Abstract: We study how depth, finite precision, state dimension, and chain-of-thought (CoT) affect the expressive power of multi-layer state-space models (SSMs). For the explicit-table $K$-function-composition problem, a canonical benchmark for sequent...
240. CaST-POI: Candidate-Conditioned Spatiotemporal Modeling for Next POI Recommendation ​
Author: Zhenyu Yu, Chunlei Meng, Yangchen Zeng, Mohd Yamani Idna Idris, Jihong Guan, Shuigeng Zhou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2604.20845v2 Announce Type: replace-cross Abstract: Next Point-of-Interest (POI) recommendation ranks a user's likely next location based on check-in history. Most recent rankers compress the trajectory into a single user vector and score every candidate through the same representation, ignori...
241. Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data ​
Author: Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2604.26841v2 Announce Type: replace-cross Abstract: When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as ...
242. Can Coding Agents Reproduce Findings in Computational Materials Science? ​
Author: Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cond-mat.mtrl-sci, cs.AI, cs.SE
arXiv:2605.00803v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific ...
243. The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting ​
Author: Lauri Lov'en, Sasu Tarkoma
Published: 9/3/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA, econ.TH, math.OC
arXiv:2605.07671v2 Announce Type: replace-cross Abstract: An agent's probability report is paid for twice: by a strictly proper scoring rule, and by an approval rule for the decision it triggers. In this classical decision-coupled setting, non-affine approval is known to defeat truthful reporting. W...
244. Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning ​
Author: Jan Net'ik, Patr'icia Martinkov'a
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.16991v2 Announce Type: replace-cross Abstract: Item difficulty must often be estimated before test administration, when no responses are yet available for calibration. While most response-free difficulty modelling approaches derive item-text features by hand for a separate statistical mod...
245. Bernini: Latent Semantic Planning for Video Diffusion ​
Author: Bernini Team, Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2605.22344v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models synthesize images and vi...
246. CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations ​
Author: Mike Zhang, Ali Basirat, Desmond Elliott
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.26293v2 Announce Type: replace-cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evalu...
247. FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies ​
Author: Xintong Hu, Xuhong Huang, Jinyu Zhang, Yutong Yao, Yuchong Sun, Qiuyue Wang, Mingsheng Li, Sicheng Xie, Yitao Liu, Junhao Chen, Yixuan Chen, Yingming Zheng, Shuai Bai, Tao Yu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2605.27284v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with coars...
248. BioELX: Context-Aware Cross-lingual Biomedical Entity Linking without Task-Specific Supervision ​
Author: Yi Wang, Corina Dima, Liangyu Zhong, Steffen Staab
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.27380v2 Announce Type: replace-cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base, supporting clinical and biomedical NLP applications. We identify two issues affecting current systems. First, th...
249. Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders ​
Author: Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.30022v2 Announce Type: replace-cross Abstract: Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks such ...
250. TUX: Measuring Human--AI Tacit Understanding ​
Author: Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha
Published: 9/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY
arXiv:2605.30930v2 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accuracy, or reward optimization. Yet many collaborative settings depend on tacit understanding:...
251. Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 ​
Author: Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.02255v3 Announce Type: replace-cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled. We provide the first ...
252. Enabling KV Caching of Shared Prefix for Diffusion Language Models ​
Author: Younghun Go, Jaehoon Han, Changyong Shin, Chuck Yoo, Gyeongsik Yang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.07571v4 Announce Type: replace-cross Abstract: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). In DLMs, bidirectional attention means that updatin...
253. DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment ​
Author: Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.07678v4 Announce Type: replace-cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data selection methods typically score each preference pair independently, collapsing directional p...
254. WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing ​
Author: Young D. Kwon, Miles Williams, Rui Li, Alexandros Kouris, Stylianos I. Venieris
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.07710v2 Announce Type: replace-cross Abstract: The autoregressive nature of large language models (LLMs) remains a significant bottleneck for inference, particularly in complex agentic workloads. While speculative decoding (SD) accelerates inference, current approaches rely on static draf...
255. Emotional regulation improves deep learning-based image classification ​
Author: Riccardo Emanuele Landi, Jo~ao M. F. Rodrigues, Marta Chinnici
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.13081v2 Announce Type: replace-cross Abstract: Emotion significantly influences cognition, enhancing memory and learning under certain conditions. Drawing on this principle, emotion-augmented deep learning investigates how affective states can improve neural network architectures and lear...
256. Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens ​
Author: Yizhen Yao, Qinglin Zhu, Runcong Zhao, Xiangxiang Dai, Yanzheng Xiang, Yulan He, Lin Gui
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.16847v4 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens...
257. Do Large Language Models Always Tell The Same Stories? ​
Author: Thennal DK, Hans Ole Hatzel
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.17350v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet whether these models are capable of generating diverse or creative artifacts remains a contested question. In this work, we investigate the...
258. Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication ​
Author: Peter Zeng, Amie J. Paige, Weiling Li, Susan E. Brennan, Owen Rambow, Cameron R. Jones
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.17372v3 Announce Type: replace-cross Abstract: Two recent studies \citep{jones2026llms, zeng2026lvlms} reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions. We control for task d...
259. Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning ​
Author: Yongbin Huang, Xihao Xie, Jia Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR
arXiv:2606.21052v2 Announce Type: replace-cross Abstract: Speech Emotion Recognition (SER) systems increasingly leverage self-supervised acoustic representations, yet their vulnerability to training-time attacks remains largely underexplored. This paper presents the first systematic study of poisoni...
260. AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents ​
Author: Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.21144v2 Announce Type: replace-cross Abstract: Long-term memory systems allow LLM agents to preserve information beyond a single context window, but most systems focus on storing and retrieving facts after extraction, leaving the write decision under-specified. What deserves memory can de...
261. SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics ​
Author: Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov, Martin Vechev
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.LG
arXiv:2606.29894v2 Announce Type: replace-cross Abstract: As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains diffic...
262. Training nGPT ​
Author: Ilya Loshchilov, Boris Ginsburg
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01284v2 Announce Type: replace-cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and eval...
263. Scaling an Autoregressive Transformer for Single-Cell Generation ​
Author: Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.GN
arXiv:2608.02961v2 Announce Type: replace-cross Abstract: We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both the biolo...
264. Direct Construction of Disambiguated Knowledge Bases from Large Language Models ​
Author: Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.03729v3 Announce Type: replace-cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess...
265. GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base ​
Author: Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.06992v2 Announce Type: replace-cross Abstract: We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and...
266. Open-World Semantic Segmentation with Sensitivity Modeling ​
Author: Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.08308v2 Announce Type: replace-cross Abstract: Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content. Conventional semantic segmentation models operate under a "closed-world" assumption, often produ...
267. Three Necessary Principles for Self-Supervised Visual Representation Learning ​
Author: Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania Stathaki
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.08309v2 Announce Type: replace-cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational...
268. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations ​
Author: Lior Baruch, Moshe Butman, Kfir Bar, Doron Friedman
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12062v2 Announce Type: replace-cross Abstract: Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree O...
269. LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment ​
Author: Haonan He, Xinyue Fan
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.19800v2 Announce Type: replace-cross Abstract: Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhead. However, a persistent performance gap remains between LoRA and full fine-tuning. Recent studies hav...
270. Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models ​
Author: Thantham Jittham
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2608.21377v2 Announce Type: replace-cross Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjectin...
271. Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules ​
Author: Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero, Carsten Eickhoff
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.22642v4 Announce Type: replace-cross Abstract: Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments. To address these challenges, we presen...
272. SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts ​
Author: Shyam Agarwal, Anmol Singhal, Travis Breaux, Bogdan Vasilescu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.25202v3 Announce Type: replace-cross Abstract: Spec-Driven Development (SDD) is a fast-emerging practice in which a structured natural-language specification, written by a developer, or (more often) drafted by an AI tool and then curated by the developer, drives an AI coding agent's imple...
273. MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration ​
Author: Miseon Yu, Jaehoon Choi, Younghan Lee, Yunheung Paek
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA
arXiv:2608.25457v3 Announce Type: replace-cross Abstract: Despite their strong ability to generate code, large language models often fail to produce secure code, as their outputs frequently contain security vulnerabilities. Secure code generation is inherently challenging because it requires solving...
274. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents ​
Author: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27141v4 Announce Type: replace-cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iteratio...
275. Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit ​
Author: Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.27309v2 Announce Type: replace-cross Abstract: Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded ra...
276. The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era ​
Author: Kiyan Rezaee
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28980v2 Announce Type: replace-cross Abstract: Can the specialized architectures that machine learning has traditionally built for structured data be replaced by language-based models? This question is examined through a review of 159 papers (2016--2026) across nine modalities, with predi...
277. SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization ​
Author: Hojae Han, Jongyoon Kim, Sanghyeok Park, Dongwook Cheon, Yeachan Park, Myung Jae Jeon, Sunjong Choe, Soonho Kong, Wonseok Hur, Seung-won Hwang, Donghoon Hyeon
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.29270v2 Announce Type: replace-cross Abstract: Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-correct but misaligned statements or reject correct statements wri...
278. A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction ​
Author: Nanxing Nick Deng, Qing Cheng, Niclas Zeller, Daniel Cremers
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.29705v2 Announce Type: replace-cross Abstract: Feed-forward 3D reconstruction models emit a per-pixel confidence that downstream systems read as a reliability signal. It is trained as a loss weight, not as an uncertainty magnitude, and whether it can be used as an error prediction has not...
279. PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies ​
Author: Botong Zhao, Fang Yu, Tim Yu, Senhua Zhu, Xinyuan Chen, Yue Lu
Published: 9/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.30378v2 Announce Type: replace-cross Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: their representations are not explicitly required to describe how the scene evolves over multip...
280. Lot Machine: Multimodal Lot Extraction from Auction Catalogs ​
Author: Mathias Zinnen, Alisha Mund, Sabine Lang, Lukas H"uttner, Thomas Gorges, Vincent Christlein
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.DL
arXiv:2608.30510v2 Announce Type: replace-cross Abstract: For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remai...
281. CogEvol: Towards Efficient and Reliable Learning Environment Generation ​
Author: Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.30968v2 Announce Type: replace-cross Abstract: We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Acro...
282. RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery ​
Author: Yifan Yang, Lei Zou
Published: 9/3/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY
arXiv:2609.00046v2 Announce Type: replace-cross Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations is essential for emergency response and recovery. However, existing AI-based approaches often require extensive manual annotation, lack cr...
283. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization ​
Author: Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.00224v2 Announce Type: replace-cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2...
284. Exploring Collaboration between a language and a non-language agent ​
Author: Harini S I, Somesh Singh, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy
Published: 9/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2609.00474v2 Announce Type: replace-cross Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not l...
285. EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection ​
Author: Guanzhong Sun, Junyi Ma, Yuxuan Wu, Yanzi Miao
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.00566v2 Announce Type: replace-cross Abstract: We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and ...
286. Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search ​
Author: Jincheng Zhang, Chen Huang, Wenqiang Lei, See-Kiong Ng, Yang Deng
Published: 9/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2609.00618v2 Announce Type: replace-cross Abstract: We investigate the role of conversational context modeling in user preference tracking for Conversational Recommendation Systems (CRSs). In this regard, we propose DREAMS, a novel tree-structured context modeling framework that explicitly cap...
287. Bandits in Prod: Hyperparameter Optimization at Inference Time ​
Author: Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2609.01335v2 Announce Type: replace-cross Abstract: Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting...
288. Rethinking Learnability in Offline Data-driven Optimization ​
Author: Chao Qian, Chen-Guang Wang, Rong-Xi Tan, Ke Xue
Published: 9/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2609.01493v2 Announce Type: replace-cross Abstract: Black-Box Optimization (BBO) has broad applications, while traditional algorithms such as evolutionary algorithms and Bayesian optimization face efficiency challenges as real-world BBO problems grow increasingly complex. Data-driven optimizat...