Skip to content

arXiv cs.AI - 2026-07-13 ​

177 items collected.


1. Interval Certifications for Multilayered Perceptrons via Lattice Traversal ​

Author: Merkouris Papamichail, Konstantinos Varsos, Giorgos Flouris, Jo~ao Marques-Silva
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.08773v1 Announce Type: new Abstract: In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness. In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal problem. Each element o...

📖 Read original article


2. CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions ​

Author: Vanessa Figueiredo, Wilter Franceschi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2607.08774v1 Announce Type: new Abstract: Reliability in large language model (LLM) systems is typically framed as a function of model capability. We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} -- the computational layer governi...

📖 Read original article


3. GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning ​

Author: Maureese Williams, Dymitr Nowicki
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.08894v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational costs and stochasti...

📖 Read original article


4. Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading ​

Author: Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.08964v1 Announce Type: new Abstract: AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks ...

📖 Read original article


5. A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game ​

Author: Joseph K. Miller
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, math-ph, math.AP, math.MP

arXiv:2607.08986v1 Announce Type: new Abstract: We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a formalization game. The objective is to turn a LaTeX document into Lean. The game is won when the development compil...

📖 Read original article


6. ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning ​

Author: Kunbo Zhang, Lei Fu, Zeyu Wang, Zijing Liu, Kejian Tong
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09059v1 Announce Type: new Abstract: We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective refin...

📖 Read original article


7. Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls ​

Author: Saroj Gopali, Bipin Chhetri, Deepika Giri, Sima Siami-Namini, Akbar Siami Namin
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09076v1 Announce Type: new Abstract: Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditional rule-based monitoring in industrial IoT environments. While Large Language Models (LLMs) have strong semantic r...

📖 Read original article


Author: Tan-Minh Nguyen, Hoang-Trung Nguyen, Huu-Dong Nguyen, Dinh-Truong Do, Thi-Hai-Yen Vuong, Le-Minh Nguyen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09099v1 Announce Type: new Abstract: While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-...

📖 Read original article


9. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation ​

Author: Runhan Shi, Quan Zhou, Yuqian Xu, Shuai Yang, Xin Wu, Zitong Zhou, Hui Liu, Bin Cha, Zheming Wang, Liya Li, Wei Wei, Haoyuan Hu, Jun Xu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2607.09142v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical ...

📖 Read original article


10. KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling ​

Author: Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09153v1 Announce Type: new Abstract: Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the enti...

📖 Read original article


11. Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift ​

Author: Dan C. Hsu, Luke Lu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.09175v1 Announce Type: new Abstract: Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational e...

📖 Read original article


12. Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents ​

Author: Izumi Takahara, Teruyasu Mizoguchi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci

arXiv:2607.09195v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific ...

📖 Read original article


13. OpenProver: Agentic and Interactive Theorem Proving with Lean 4 ​

Author: Mat\v{e}j Kripner, Milan Straka
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.MS

arXiv:2607.09217v1 Announce Type: new Abstract: In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic sy...

📖 Read original article


14. LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making ​

Author: Yanzhen Chen, Zihan Xu, Xiaocheng Zhang, Zhiting Fan, Weiqi Zhai, Hongxia Xu, Zuozhu Liu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medi...

📖 Read original article


15. Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks ​

Author: Nuocheng Yang, Sihua Wang, Zihan Chen, Tony Q. S. Quek, Changchuan Yin
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.09330v1 Announce Type: new Abstract: Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intelligence such as smart factories, warehouses, and service robotics. To enable collaboration among such an agent team, effici...

📖 Read original article


16. Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review ​

Author: Jingbo Chen, He Wang, Wei Yuan, Yuqiao Lai, Zhenyan Lu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09403v1 Announce Type: new Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application to worldbuilding fac...

📖 Read original article


17. How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding ​

Author: Debargha Ghosh, Silja Renooij, Anna Kononova
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09449v1 Announce Type: new Abstract: Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) through posterior inference. However, its behaviour under latent confounding remains poorly understood, as existing work typi...

📖 Read original article


18. ProofCouncil: An LLM Agent for Solving Open Mathematical Problems ​

Author: Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely B'erczi, Uri Kreitner, Liam Price, David Holmes
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09474v1 Announce Type: new Abstract: Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce Pro...

📖 Read original article


19. Ceci n'est pas une pipe: AI systems as semantic abstractions ​

Author: Jade Alglave, Patrick Cousot
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.PL

arXiv:2607.09489v1 Announce Type: new Abstract: An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation. We propose a semantic framework to describe AI systems, to be able to examine the correctness of such representations. To do so, we di...

📖 Read original article


20. Multimodal Reward Hacking in Reinforcement Learning ​

Author: Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09492v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-only or weakly grounded...

📖 Read original article


21. Shared Selective Persistent Memory for Agentic LLM Systems ​

Author: Sanjana Pedada, Aditya Dhavala, Neelraj Patil
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE

arXiv:2607.09493v1 Announce Type: new Abstract: Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessi...

📖 Read original article


22. SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction ​

Author: Chongyu Qu, Can Cui, Zhengyi Lu, Junchao Zhu, Tianyuan Yao, Junlin Guo, Juming Xiong, Yanfan Zhu, Yuechen Yang, Bennett A. Landman, Yuankai Huo
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09521v1 Announce Type: new Abstract: Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from demographics collected at intake ...

📖 Read original article


23. Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI ​

Author: Yuan Cao, Haiqian Yang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.09560v1 Announce Type: new Abstract: Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within...

📖 Read original article


24. Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining ​

Author: Jan Gronewald, Andreas Emrich, Nijat Mehdiyev
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09578v1 Announce Type: new Abstract: Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in which AI support must serve qualified auditors who remain accountable for the decisions taken. The relevant unit of value is not predictio...

📖 Read original article


25. TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems ​

Author: Hannah M. Liu, Rhea Saxena, Shiv Asthana
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09586v1 Announce Type: new Abstract: The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classification Framework...

📖 Read original article


26. Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation ​

Author: Kaiji Zhou, Ales Leonardis, Yue Feng
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.09600v1 Announce Type: new Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the funct...

📖 Read original article


27. ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI ​

Author: Mohadeseh Mollapour, Koorosh Aslansefat, Zeinab Dehghani, Bhupesh Kumar Mishra, Tejal Shah, Zhibao Mian
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.09649v1 Announce Type: new Abstract: Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework fo...

📖 Read original article


28. Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model ​

Author: Song-Ju Kim
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, q-bio.NC

arXiv:2601.10034v2 Announce Type: cross Abstract: Decision making often exhibits context dependence that challenges classical probability theory. This paper develops a quantum-like extension of the Tug-of-War (QTOW) decision-making model to clarify when such context dependence can be represented by ...

📖 Read original article


29. REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming ​

Author: Nicolas Koller, Andreas u. Schmidt
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR, cs.PL

arXiv:2607.07738v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security workflows. Claims about their capability, however, outpace our ability to...

📖 Read original article


30. A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions ​

Author: Qingzhuo Wang, Ruiyang Qin, Zhenxin Qin, Wen Shen, Zhihua Wei
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.GT

arXiv:2607.08776v1 Announce Type: cross Abstract: Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unified approach to explore the common mechanism of various KD methods using...

📖 Read original article


31. iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis ​

Author: Farica Zhuang, Seong Woo Han, Zixuan Wen, Shu Yang, Yize Zhao, Li Shen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08778v1 Announce Type: cross Abstract: Alzheimer's Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide. Predicting AD conversion during the prodromal stage remains critical for disease understanding and patient care. As such, survival...

📖 Read original article


32. Signed Symmetric Quantization for Few-Bit Integers ​

Author: Ian Colbert, Eashan Dash, Pablo Monteagudo-Lago, Juan Amboage, Srinidhi N, Giuseppe Franco, Nicholas J. Fraser, Arun Ramachandran
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08779v1 Announce Type: cross Abstract: The signed integer alphabet contains one more negative representable value than positive. Yet, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negati...

📖 Read original article


33. Sticky Routing: Training MoE Models for Memory-Efficient Inference ​

Author: Ali Kayyam
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.08780v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping between slow storage and fast memory on edge devices. Existing remedi...

📖 Read original article


34. Reward Transport: Property Control in Flow Matching via Noise-Space Alignment ​

Author: Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.QM

arXiv:2607.08781v1 Announce Type: cross Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a targ...

📖 Read original article


35. Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement ​

Author: Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li, Fahao Chen, Haodong Wang, Jian Lin, Song Guo
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08782v1 Announce Type: cross Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing work...

📖 Read original article


36. LieBN: Batch Normalization over Lie Groups ​

Author: Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, Nicu Sebe
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08783v1 Announce Type: cross Abstract: Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, collectively re...

📖 Read original article


37. HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning ​

Author: Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC

arXiv:2607.08784v1 Announce Type: cross Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client ...

📖 Read original article


38. DaDaDa: A Dataset for Data Pricing in Data Marketplaces ​

Author: Qiheng Sun, Hongwei Zhang, Junxu Liu, Xiaokai Mao, Jinfei Liu, Kui Ren, Haibo Hu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08785v1 Announce Type: cross Abstract: High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplaces, e.g., AWS Marketplace, Databricks, and Datarade. However, determini...

📖 Read original article


39. Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices ​

Author: Tao Lu, Haoyu Wang, Zonghui Wang, Keshen Xiang, Jiaheng Zhang, Wenzhi Chen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2607.08786v1 Announce Type: cross Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits ...

📖 Read original article


40. LLM-Driven Evolutionary Generation of Multi-Objective Bayesian Optimization Algorithms ​

Author: Georgios Laskaris, Reuben Brasher, Niki van Stein, Elena Raponi, Thomas B"ack, Florian Neukart
Published: 7/13/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.08791v1 Announce Type: cross Abstract: Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdependent design choices whose optimal configuration is problem-dependent and typically demands deep expertise. We extend the LLaMEA framework to...

📖 Read original article


41. EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins ​

Author: Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
Published: 7/13/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, cs.SY, eess.SY, math.OC

arXiv:2607.08793v1 Announce Type: cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inferen...

📖 Read original article


42. Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection ​

Author: Padam Jung Thapa, Abdullah Bin Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque
Published: 7/13/2026, 4:00:00 AM
Categories: cs.GR, cs.AI

arXiv:2607.08794v1 Announce Type: cross Abstract: Sand boils on earthen levees are safety-critical defects, but pixel-level detection is limited by scarce annotations. We present a diffusion-based synthesis pipeline for low-resource sand-boil imagery. Using Stable Diffusion XL fine-tuned with DreamB...

📖 Read original article


43. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology ​

Author: Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
Published: 7/13/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.LG

arXiv:2607.08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories...

📖 Read original article


44. Prompt-Driven Exploration ​

Author: Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08837v1 Announce Type: cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a we...

📖 Read original article


45. A Novel Parallel QCNN Architecture with Efficient Classical Simulability ​

Author: Lawrence Nguyen, Hiu Yung Wong
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2607.08928v1 Announce Type: cross Abstract: This work presents a study of an implementation of a novel Quantum Convolutional Neural Network (QCNN) for binary classification of images from the Modified National Institute of Standards and Technology (MNIST) dataset. Using a novel architecture in...

📖 Read original article


46. Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution ​

Author: Ning Liu, Kalle Kujanp"a"a, Zhaoxuan Zhu, P Aditya Sreekar, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez, Matthew Bales, Tianyu Yang, Shahnawaz Alam, Rose Yu, Baoyuan Liu, Kristina Klinkner, Shervin Malmasi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08960v1 Announce Type: cross Abstract: Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance...

📖 Read original article


47. NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision ​

Author: Berkay Anahtarci
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.ST, stat.TH

arXiv:2607.08961v1 Announce Type: cross Abstract: Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels redu...

📖 Read original article


48. MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs ​

Author: Hantao Zhang, Jinru Sui, Ed Li, Dirk Bergemann, Zhuoran Yang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.08970v1 Announce Type: cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView...

📖 Read original article


49. CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding ​

Author: Yuri Ishitoya, Jeremy Siburian, Masashi Hamaya, Kuniaki Saito, Cristian C. Beltran-Hernandez, Mai Nishimura
Published: 7/13/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.08974v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the backbone so extensively that it becomes difficult to isolate what the...

📖 Read original article


50. The Patchwork Problem in LLM-Generated Code ​

Author: Viraaji Mothukuri, Reza M. Parizi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.08981v1 Announce Type: cross Abstract: LLM-generated code often compiles, passes tests, and appears correct, yet breaks once deployed. The root cause is frequently structural rather than logical. A generated endpoint references configuration keys never declared in the project, an import t...

📖 Read original article


51. SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation ​

Author: Sijia Gu, Noor Nashid, Ali Mesbah
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.08983v1 Announce Type: cross Abstract: While autonomous coding agents have significantly advanced automated test generation, they remain fundamentally limited by lazy generation, a phenomenon where agents prematurely terminate tasks and systematically avoid complex programmatic logic, res...

📖 Read original article


52. AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision ​

Author: Brent Kong, Tejas Ram, Tony Yue Yu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT, math.CO

arXiv:2607.08984v1 Announce Type: cross Abstract: AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect...

📖 Read original article


53. Model Agnostic Graph Prompt Learning for Crystal Property Prediction ​

Author: Shrimon Mukherjee, Kishalay Das, Partha Basuchowdhuri, Pawan Goyal, Niloy Ganguly
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.08996v1 Announce Type: cross Abstract: Graph Neural Networks have emerged as a powerful tool for the fast and accurate prediction of various crystal properties. These models often encode domain-specific knowledge into their graph encoding modules, which increases their parameter size and ...

📖 Read original article


54. Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing ​

Author: Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi, Vijaykrishnan Narayanan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09015v1 Announce Type: cross Abstract: We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely ...

📖 Read original article


55. Phone Segmentation and Recognition through Phonological Activation Mapping ​

Author: Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
Published: 7/13/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD

arXiv:2607.09020v1 Announce Type: cross Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only ne...

📖 Read original article


56. Video Generation Models are General-Purpose Vision Learners ​

Author: Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that lar...

📖 Read original article


57. Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems ​

Author: Chao Wang, Lingling Li, Fang Liu, Licheng Jiao
Published: 7/13/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CE

arXiv:2607.09025v1 Announce Type: cross Abstract: Artificial intelligence (AI) is shifting scientific discovery from task-specific workflows towards autonomous systems that organize exploration with experimental and human feedback in open-ended candidate spaces. Evolutionary computation (EC) provide...

📖 Read original article


58. Quantum Logic as the Logic of Contexts ​

Author: Haruki Emori, Atsushi Iriki, Andrei Khrennikov, Kazunori Kondo
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, math.LO, q-bio.NC

arXiv:2607.09032v1 Announce Type: cross Abstract: Quantum logic is usually presented as a non-classical departure from ordinary reasoning forced on us by quantum mechanics, with classical logic kept as the secure starting point. We argue for the opposite order of explanation in a finite and fully co...

📖 Read original article


59. On Locality and Length Generalization in Visual Reasoning ​

Author: Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.09061v1 Announce Type: cross Abstract: A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision m...

📖 Read original article


60. Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills ​

Author: Jialun Cao, Xinru Yan, Songqiang Chen, Yaojie Lu, Zhongxin Liu, Shing-Chi Cheung
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.09065v1 Announce Type: cross Abstract: Software engineering (abbrev. SE) has continuously evolved through increasingly powerful forms of reuse, from source code and libraries to components and services. Recent advances in AI agents have introduced a potentially new reusable artifact: skil...

📖 Read original article


61. OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents ​

Author: Yang Chen, Yunwen Li, Yufan Shen, Minghao Liu, Tianyu Zheng, Bin Fu, Qunshu Lin, Zhi Yu, Botian Shi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09068v1 Announce Type: cross Abstract: Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance w...

📖 Read original article


Author: Devanshu Verma, Vasudha Bhatnagar, Vikas Kumar, Balaji Ganesan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09094v1 Announce Type: cross Abstract: Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research. Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and compute sim...

📖 Read original article


63. A Coreset Selection Framework with Ensemble Aggregation for Image Classification ​

Author: Pedro Rocha Dantas, Lucas Pascotti Valem
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09100v1 Announce Type: cross Abstract: The rapid growth of image data has produced large-scale datasets, raising concerns about the time and memory costs of model training. Selecting representative training subsets, however, remains challenging: individual sample contributions are unclear...

📖 Read original article


64. Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging ​

Author: Yawen Li, Yan Li, Zhe Xue, Yingxia Shao, Meiyu Liang, Guanhua Ye
Published: 7/13/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.MM

arXiv:2607.09102v1 Announce Type: cross Abstract: Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and man...

📖 Read original article


65. Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification ​

Author: Camila Piscioneri Magalh~aes, Lucas Pascotti Valem
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09104v1 Announce Type: cross Abstract: While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and un...

📖 Read original article


66. Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms ​

Author: Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2607.09114v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these ...

📖 Read original article


67. Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs ​

Author: Bartosz Zi'o{\l}ko, Kacper Dobrzeniewski
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, q-fin.PM, q-fin.TR

arXiv:2607.09121v1 Announce Type: cross Abstract: In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation like GDP and inflation...

📖 Read original article


68. IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation ​

Author: Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09133v1 Announce Type: cross Abstract: While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (C...

📖 Read original article


69. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models ​

Author: Sang-Hoon Lee, Ha-Yeong Choi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS, eess.SP

arXiv:2607.09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To ad...

📖 Read original article


70. A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models ​

Author: Qingchu Jin, Felistas Mazhude, Jamie B. Rabb, Robert S. Kramer, Douglas B. Sawyer, Raimond L. Winslow
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09165v1 Announce Type: cross Abstract: Achieving early and timely diagnosis and treatment for disease is a major challenge. Recent applications of machine learning (ML) algorithms trained on patient data have shown promise in many different settings for predicting the patient health state...

📖 Read original article


71. Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations ​

Author: Nada Zine, Tristan Coignion, Vincenzo Stoico, Cl'ement Quinton, Romain Rouvoy, Patricia Lago
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.PF

arXiv:2607.09172v1 Announce Type: cross Abstract: Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior work has fo...

📖 Read original article


Author: Wenjun Zhang, Zhiyong Chen, Tong Wu, Guo Lu, Li Song, Feng Yang, Meixia Tao
Published: 7/13/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT

arXiv:2607.09183v1 Announce Type: cross Abstract: The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as images and videos, reshaping communication paradigms. This article introduces generative communications (GenCom), a ...

📖 Read original article


73. Interference and Retention in Continual Learning ​

Author: Julius St"ork
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2607.09202v1 Announce Type: cross Abstract: Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forget...

📖 Read original article


74. Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation ​

Author: Rishabh Madan, Angchen Xie, Samantha Saak, Andres Blanco, Dohyeok Lee, Sarah Grace Brown, Yunting Yan, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee
Published: 7/13/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09218v1 Announce Type: cross Abstract: Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multiple links as contacts form, slide, and break. This setting breaks common implicit assumptions in many learning-ba...

📖 Read original article


75. Git-Assistant: Planning-Based Support for Updating Git Repositories ​

Author: Alfredo Garrach'on Ruiz, Tom'as de la Rosa, Daniel Borrajo
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.09224v1 Announce Type: cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Recent advances in Large Language Models (LLMs) offer promising capabilities for interpreting developer intent...

📖 Read original article


76. All you need is SAMPAT ​

Author: Jayadeva, Madhur Aswani
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, math.FA

arXiv:2607.09235v1 Announce Type: cross Abstract: The current state of the art in AI/ML rests on deep neural architectures, which, in general, suffer from a lack of interpretability. Interpretability is crucial to gleaning insights while analyzing experimental data, where quantitative predictions ma...

📖 Read original article


77. LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics ​

Author: Gwenn Beets, Anniek Jansen, Saar Hommes, Ruben D. Vromans, Leonie Westerbeek, Supraja Sankaran, Julia C. M. van Weert, Emiel J. Krahmer, Nadine Bol
Published: 7/13/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.09253v1 Announce Type: cross Abstract: AI chatbots are increasingly used for answering health-related questions. This study examines the role of topic type discussed with an AI chatbot and individual characteristics on perceived benefits and risks, intention to use an AI chatbot, and will...

📖 Read original article


78. Blockchain-Linked Auditable Decision Management for Telecom/IoT Fraud-Control Requests ​

Author: Saviz Changizi, Nasibeh Mohammadzadeh, Mohammad Shojafar, Rahim Tafazolli
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.09259v1 Announce Type: cross Abstract: Telecom fraud-control studies often stop at detector-level classification, but deployment use requires request-level policy resolution, lifecycle traceability, and auditability. This paper reframes fraud control as blockchain-linked auditable decisio...

📖 Read original article


79. Geopolitical alignment: Endorsement effects in large language models ​

Author: Maxim Chupilkin
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.09262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment ...

📖 Read original article


80. Risk-Aware General-Utility Markov Decision Processes ​

Author: Pedro P. Santos, F'abio Vital, Alberto Sardinha, Francisco S. Melo
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09298v1 Announce Type: cross Abstract: We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visi...

📖 Read original article


81. Creativity, honesty and designed forgetting emerge in small hyperbolic language models ​

Author: Kwan Soo Shin, In Seok Kang, Yunkyung Min
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.LG

arXiv:2607.09306v1 Announce Type: cross Abstract: Language models are optimised for scale, yet remain functional rather than companionable, and as an assistant personalises into a companion, accumulating memory of one user, it quietly becomes someone, and can silently acquire traits that harm that u...

📖 Read original article


82. Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works ​

Author: Miguel Arana-Catania, Gillian Pink, Glenn Roe
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL, cs.IR, cs.LG

arXiv:2607.09316v1 Announce Type: cross Abstract: Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper...

📖 Read original article


83. Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI ​

Author: Miguel Arana-Catania, Catherine Conisbee, Matthew Kidd
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL, cs.IR, cs.LG

arXiv:2607.09324v1 Announce Type: cross Abstract: Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" project, which used the Their Fine...

📖 Read original article


84. WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning ​

Author: Zixin Chen, Peng Liu, Haobo Li, Rui Sheng, Jianhong Tu, Xiaodong Deng, Fei Huang, Kashun Shum, Dayiheng Liu, Huamin Qu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09328v1 Announce Type: cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that joi...

📖 Read original article


85. Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning ​

Author: Guanquan Wang, Yoshimasa Tsuruoka
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2607.09336v1 Announce Type: cross Abstract: Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typi...

📖 Read original article


86. Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation ​

Author: Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.09349v1 Announce Type: cross Abstract: Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated ch...

📖 Read original article


87. CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation ​

Author: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09362v1 Announce Type: cross Abstract: Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or u...

📖 Read original article


88. Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability ​

Author: Shirley Yu, Ruben Martins
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LO

arXiv:2607.09366v1 Announce Type: cross Abstract: Program verification is crucial for software correctness, but producing fully verified programs remains difficult in practice. This paper studies whether implementation structure affects automated verifiability when multiple generated programs are in...

📖 Read original article


89. When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks ​

Author: Brennan Bell, Inti Gabriel Mendoza Estrada, Andreas Tr"ugler, Paul Erker
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.CR

arXiv:2607.09378v1 Announce Type: cross Abstract: We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move, while Eve selects an attack surface, ei...

📖 Read original article


90. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ​

Author: Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini
Published: 7/13/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.PF

arXiv:2607.09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliab...

📖 Read original article


91. Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks ​

Author: Wout Mommen, Lars Keuninckx, Matthias Hartmann, Werner Van Leekwijck, Piet Wambacq
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09399v1 Announce Type: cross Abstract: We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connectio...

📖 Read original article


92. On-Device Adaptive Battery Power Prediction for Electric Vehicles ​

Author: Avik Bhatnagar, Anton Paule, Tobias Schuermann, Sebastian Reiter, Oliver Bringmann
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR, cs.PF

arXiv:2607.09400v1 Announce Type: cross Abstract: Adaptive power management in Electric Vehicles (EVs) requires accurate power prediction. Although deep learning models have emerged as highly effective for time-series forecasting in this domain, their performance is prone to degradation when exposed...

📖 Read original article


93. Self-Guided Test-Time Training for Long-Context LLMs ​

Author: Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating th...

📖 Read original article


94. SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition ​

Author: Hyein Park, Namho Kim, Junhwa Kim
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09417v1 Announce Type: cross Abstract: Ambivalence and hesitancy are subtle behavioral states that are expressed through a combination of verbal content, facial behavior, visual context, and acoustic cues. Effective recognition therefore requires not only extracting informative unimodal r...

📖 Read original article


95. A Sovereign, Open-Source Foundation Model for German and English ​

Author: The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben H"arle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom R"ohr, Sebastian von Rohrscheidt, J"org Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim K"ohler, Alexander L"oser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max L"ubbering
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.09424v1 Announce Type: cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constan...

📖 Read original article


96. Test-Time Scaling for Small VLMs on Multilingual Visual MCQ ​

Author: Spiros Baxevanakis, Peng-Jian Yang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.09438v1 Announce Type: cross Abstract: Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-c...

📖 Read original article


97. Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification ​

Author: Anil Osman Tur, Tonje Knutsen Sordalen, Kim Tallaksen Halvorsen, Cigdem Beyan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09443v1 Announce Type: cross Abstract: Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although recent vision-language models provide strong pretrained visual representations, adapting them to longitudinal eco...

📖 Read original article


98. Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning ​

Author: Charles Edward Gagnon, Steven H. H. Ding, Philippe Charland, Benjamin C. M. Fung
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.09452v1 Announce Type: cross Abstract: We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based source code retrieval, and large language model reasoning. Our binary-to-source-code retrieval method attempts to...

📖 Read original article


99. Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation ​

Author: Yungeng Liu, Xuanzi Fang, Haijin Zeng, Qi Dai, Yongyong Chen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09481v1 Announce Type: cross Abstract: Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture. Such tight coupling makes it diffi...

📖 Read original article


100. All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models ​

Author: Pan Li
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2607.09502v1 Announce Type: cross Abstract: Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We arg...

📖 Read original article


101. What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility ​

Author: Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09503v1 Announce Type: cross Abstract: A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap. We demonstrate that VGGT implicitly encodes co-v...

📖 Read original article


102. Failure as a Process: An Anatomy of CLI Coding Agent Trajectories ​

Author: Xiangxin Zhao, Han Li, Shuaiting Li, Tianyi Zhao, Earl T. Barr, Federica Sarro, He Ye
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.09510v1 Announce Type: cross Abstract: Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making their reliability a growing concern. Existing empirical studies investigate why coding agents...

📖 Read original article


103. Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference ​

Author: Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09520v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood. Existing efficiency efforts focus predominantly on reducing visual tokens, implicitly treating visual pro...

📖 Read original article


104. ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts ​

Author: Jiawen Li, Tian Guan, Huijuan Shi, Xitong Ling, Mingxi Fu, Anjia Han, Chao He, Yonghong He
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09526v1 Announce Type: cross Abstract: Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones. Here we present ALICE, a unified...

📖 Read original article


105. TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models ​

Author: Tianyou Jiang, Ziyu Zhou
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09562v1 Announce Type: cross Abstract: Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptat...

📖 Read original article


106. Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary Algorithms ​

Author: Danial Ramezani, Mostafa Abouei Ardakan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CE, cs.AI, math.OC, q-fin.PM, q-fin.RM

arXiv:2607.09566v1 Announce Type: cross Abstract: Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternatives available in financial markets. A hot area of research over the past few decades has been portfolio optimization that seeks to d...

📖 Read original article


107. Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach ​

Author: Kiran Pala, Punam Silu, Lixun Yu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.ET

arXiv:2607.09576v1 Announce Type: cross Abstract: We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic. Each expression is a...

📖 Read original article


108. PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers ​

Author: Yujie Pang, Zudong Li
Published: 7/13/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.09590v1 Announce Type: cross Abstract: Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost...

📖 Read original article


109. Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026 ​

Author: Nirjhar Das, Md. Al-Mamun Provath
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.09623v1 Announce Type: cross Abstract: We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed t...

📖 Read original article


110. 4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception ​

Author: Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion n...

📖 Read original article


111. Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory ​

Author: Chengkai Zhu, Ziao Tang, Guocheng Zhen, Yimeng Cao, Yusheng Zhao, Ranyiliu Chen, Xuanqiang Zhao, Lei Zhang, Xin Wang
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2607.09632v1 Announce Type: cross Abstract: Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correction. Formalizing its coding theorems requires connecting finit...

📖 Read original article


112. Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection ​

Author: Cl'audio L'ucio do Val Lopes, Lucca Machado da Silva
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.09641v1 Announce Type: cross Abstract: Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To o...

📖 Read original article


113. VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents ​

Author: Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.09653v1 Announce Type: cross Abstract: Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions of Large Lan...

📖 Read original article


114. Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models ​

Author: Shravan Murlidaran, Miguel P. Eckstein
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09654v1 Announce Type: cross Abstract: Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated hum...

📖 Read original article


115. Scalable Visual Pretraining for Language Intelligence ​

Author: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2607.09657v1 Announce Type: cross Abstract: The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts c...

📖 Read original article


116. PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis ​

Author: Ren Takahashi, Emre Yusuf, Jayabrata Bhaduri
Published: 7/13/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG, eess.SP, math.AT

arXiv:2607.09662v1 Announce Type: cross Abstract: Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achieving a state-of-the-art area under the receiver operating characteristic curve (AUC) of approximately 0.70 on the ...

📖 Read original article


117. IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs ​

Author: Jinwei He, Feng Lu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2409.05559v3 Announce Type: replace Abstract: Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is still limited exploration of abductive reasoning. The multi-perspective and multi-level of causes is one of th...

📖 Read original article


118. Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors ​

Author: Maheep Chaudhary, Fazl Barez
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2505.14300v2 Announce Type: replace Abstract: White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior. However, white-box monitors can be circumvented, and the mechanisms underlying such eva...

📖 Read original article


119. A Descriptive and Normative Theory of Human Beliefs in RLHF ​

Author: Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2506.01692v2 Announce Type: replace Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human beliefs about the capabilities of the agent being trained also play a key...

📖 Read original article


120. QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming ​

Author: Zhenxiao Fu, Lei Jiang, Yilun Xu, Gang Huang, Fan Chen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, quant-ph

arXiv:2508.20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific plan...

📖 Read original article


121. Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity ​

Author: Charles E. Gagnon, Steven H. H. Ding, Philippe Charland, Benjamin C. M. Fung
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.SE

arXiv:2509.23449v2 Announce Type: replace Abstract: Binary code similarity detection is a core task in reverse engineering. It supports malware analysis and vulnerability discovery by identifying semantically similar code in different contexts. Modern methods have progressed from manually engineered...

📖 Read original article


122. Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting ​

Author: Chenhua Shi, Bhavika Jalli, Gregor Macdonald, John Zou, Wanlu Lei, Mridul Jain, Joji Philip
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IT, cs.MA, cs.NI, math.IT

arXiv:2511.00651v2 Announce Type: replace Abstract: Telecom networks are rapidly growing in scale and complexity, making effective management, operation, and optimization increasingly challenging. Although Artificial Intelligence (AI) has been applied to many telecom tasks, existing models are often...

📖 Read original article


123. Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge ​

Author: Shashank Kirtania, Param Biyani, Priyanshu Gupta, Yasharth Bajpai, Roshni Iyer, Sumit Gulwani, Gustavo Soares
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.20297v2 Announce Type: replace Abstract: Large Language Model (LLM)-based agents are increasingly capable of complex, multi-step tasks such as GUI automation, tool use, and data manipulation, yet they cannot learn from experience: each new session rediscovers solutions from scratch. We in...

📖 Read original article


124. Programming over Thinking: Efficient and Robust Multi-Constraint Planning ​

Author: Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.09097v4 Announce Type: replace Abstract: Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints. Existing large language model (LLM) approaches face fundamental limitations in this domain. Pur...

📖 Read original article


125. PACE: A Personalized Adaptive Curriculum Engine for 9-1-1 Call-taker Training ​

Author: Zirong Chen, Hongchao Zhang, Meiyi Ma
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.05361v2 Announce Type: replace Abstract: 9-1-1 call-taking training requires mastery of over a thousand interdependent skills, covering diverse incident types and protocol-specific nuances. A nationwide labor shortage is already straining training capacity, but effective instruction still...

📖 Read original article


126. A Self-Evolving Agentic Framework for Metasurface Inverse Design ​

Author: Yi Huang, Bowen Zheng, Yunxi Dong, Hong Tang, Huan Zhao, S. M. Rakibul Hasan Shawon, Hualiang Zhang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, physics.comp-ph

arXiv:2604.01480v2 Announce Type: replace Abstract: Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization code still requires substantial expertise in computational electromagnetics and solver-specific software engine...

📖 Read original article


127. Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys ​

Author: Zikun Ye, Hema Yoganarasimhan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, stat.AP

arXiv:2604.17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LL...

📖 Read original article


128. Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs ​

Author: Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.17502v4 Announce Type: replace Abstract: Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Same-Length Trajectories (DReST) reward function does this by penalizi...

📖 Read original article


129. HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs ​

Author: Darsh Kachroo, Arjun Prasaath Anbazhagan, Adriana Caraeni, Brennan Lagasse, Kevin Zhu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.20140v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex reasoning tasks. DPO optimizes for the likelihood of generating preferred over dispreferred resp...

📖 Read original article


130. Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning ​

Author: Wei Duan, Junyu Xuan, En Yu, Xiaoyu Yang, Jie Lu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2605.17393v2 Announce Type: replace Abstract: Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners lack a theoretically grounded mechanism to decide which edges should exist and how much information each edge...

📖 Read original article


131. Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers ​

Author: Carmen Quiles-Ram'irez, Leticia L. Rodr'iguez, Nicol'as Martorell, Natalia D'iaz-Rodr'iguez
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.LO, cs.MA

arXiv:2605.28215v2 Announce Type: replace Abstract: In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use the provided context remains opaque. While Chain-of-Thought prompting is widely used, recent work ...

📖 Read original article


132. Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs ​

Author: Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang, Dexu Yu, Ronghao Chen, Yang Zhou, Hongwu Peng, Xuanqi Lan, Dimitris N. Metaxas, Youhua Li
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.00726v2 Announce Type: replace Abstract: Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures a...

📖 Read original article


133. SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition ​

Author: Jayanta Dey, Shikhar Srivastava, Itamar Lerner, Christopher Kanan, Dhireesha Kudithipudi
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.00732v4 Announce Type: replace Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings. In these settings, data arrive sequentially and must be processed in a single pass without simultan...

📖 Read original article


134. Coding-agents can replicate scientific machine learning papers ​

Author: Atharva Hans, Ilias Bilionis
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.02134v2 Announce Type: replace Abstract: Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be prompted to replicate those c...

📖 Read original article


Author: Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin, M{\aa}ns Magnusson, Huiyuan Xie, Hao Tian Yeung, Christine Carter, Jonathan Rutherford, Felix Steffek
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.04261v2 Announce Type: replace Abstract: Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shor...

📖 Read original article


136. Large Behavior Model: A Promptable Digital Twin of the Retail Customer ​

Author: Wachiravit Modecrua, Krittin Pachtrachai, Touchapon Kraisingkorn
Published: 7/13/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.06993v2 Announce Type: replace Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We pres...

📖 Read original article


137. Projection Methods for Operator Learning and Universal Approximation ​

Author: Emanuele Zappala
Published: 7/13/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.LG, cs.NA

arXiv:2406.12264v5 Announce Type: replace-cross Abstract: We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Schauder mapping. Moreover, we introduce and study a method for operator learning in Banach spaces $L^p$ ...

📖 Read original article


138. Multi-Attribute Steering of Language Models via Targeted Intervention ​

Author: Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2502.12446v3 Announce Type: replace-cross Abstract: Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to the...

📖 Read original article


139. Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning ​

Author: Cyril Shih-Huan Hsu, Anestis Dalgkitsis, Paola Grosso, Chrysa Papagianni
Published: 7/13/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.LG, cs.NE

arXiv:2504.18902v3 Announce Type: replace-cross Abstract: In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and ubiquitous connectivity, effective management of Virtualized Network Functions (VNFs) is essential. VNFs are software-based counterparts ...

📖 Read original article


140. M4V: Multimodal Mamba for Efficient Text-to-Video Generation ​

Author: Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao, Yinlong Qian, Ling Chen, Yunchao Wei, Lin Ma
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2506.10915v2 Announce Type: replace-cross Abstract: Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when employ...

📖 Read original article


141. Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching ​

Author: Yu Han, Zhiwei Huang, Yanting Zhang, Fangjun Ding, Shen Cai, Xiaoyu Tang, Yanchao Dong, Rui Fan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2506.22784v2 Announce Type: replace-cross Abstract: Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and structur...

📖 Read original article


142. GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs ​

Author: Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2507.18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights. However, most existin...

📖 Read original article


143. Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs ​

Author: Skatje Myers, Dmitriy Dligach, Timothy A. Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew Churpek, Anoop Mayampurath, Majid Afshar
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.14817v2 Announce Type: replace-cross Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Methods: We defined three EHR-based tasks that a...

📖 Read original article


144. REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression ​

Author: Mengjie Li, Yuan Feng, Xike Xie, William J. Song
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.15806v2 Announce Type: replace-cross Abstract: The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cache eviction methods primarily analyze the inference behavior of attention heads in successful retrieval-...

📖 Read original article


145. Contrastive Weak-to-strong Generalization ​

Author: Houcheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang, Chen Gao, Xiang Wang, Xiangnan He, Yang Deng
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.07884v2 Announce Type: replace-cross Abstract: Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling. However, its r...

📖 Read original article


146. Explaining Human Choice Probabilities with Simple Vector Representations ​

Author: Peter A. V. DiBerardino, Britt Anderson
Published: 7/13/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2511.03643v3 Announce Type: replace-cross Abstract: We formalize human choice behavior in a probabilistic hide-and-seek task. In our geometric construction, vectors represent participant choice frequencies as well as probability matching and maximizing strategies. We measured choice behavior n...

📖 Read original article


147. H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification ​

Author: Yongji Zhang, Siqi Li, Kuiyang Huang, Yue Gao, Yu Jiang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2511.10260v2 Announce Type: replace-cross Abstract: Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to l...

📖 Read original article


148. AutoGraphAD: Unsupervised network anomaly detection using Variational Graph Autoencoders ​

Author: Georgios Anyfantis, Pere Barlet-Ros
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2511.17113v3 Announce Type: replace-cross Abstract: Network Intrusion Detection Systems (NIDS) are essential tools for detecting network attacks and intrusions. While extensive research has explored the use of supervised Machine Learning for attack detection and characterisation, these methods...

📖 Read original article


149. Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation ​

Author: Scott Merrill, Shashank Srivastava
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD

arXiv:2511.17813v3 Announce Type: replace-cross Abstract: LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anonymous labels such as $S...

📖 Read original article


Author: Changpeng He, Yang Lu, Yanqing Xu, Chong-Yung Chi, Arumugam Nallanathan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2511.20305v2 Announce Type: replace-cross Abstract: This paper investigates a reconfigurable intelligent surface (RIS)-assisted multi-waveguide pinching-antenna (PA) system (PASS) for multi-user downlink information transmission, motivated by the unknown impact of the integration of emerging P...

📖 Read original article


151. Data-Driven Learnability Transition of Measurement-Induced Entanglement ​

Author: Dongheng Qian, Jing Wang
Published: 7/13/2026, 4:00:00 AM
Categories: quant-ph, cond-mat.dis-nn, cs.AI

arXiv:2512.01317v3 Announce Type: replace-cross Abstract: Measurement-induced entanglement (MIE) captures how local measurements generate long-range quantum correlations and drive dynamical phase transitions in many-body systems. Yet estimating MIE experimentally remains challenging: direct evaluati...

📖 Read original article


152. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy ​

Author: Natalia Ponomareva, Zheng Xu, H. Brendan McMahan, Peter Kairouz, Lucas Rosenblatt, Vincent Cohen-Addad, Crist'obal Guzm'an, Ryan McKenna, Galen Andrew, Alex Bie, Da Yu, Alex Kurakin, Morteza Zadimoghaddam, Sergei Vassilvitskii, Andreas Terzis
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, stat.ML

arXiv:2512.03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated data will soon have been used. Additionally, publicly available ...

📖 Read original article


153. ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning ​

Author: Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett
Published: 7/13/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2512.16861v2 Announce Type: replace-cross Abstract: Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form an initial solution, ...

📖 Read original article


154. Transition Matching Distillation for Fast Video Generation ​

Author: Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, Arash Vahdat
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2601.09881v2 Announce Type: replace-cross Abstract: Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, w...

📖 Read original article


155. Principles of Lipschitz continuity in neural networks ​

Author: R'ois'in Luo
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2602.04078v2 Announce Type: replace-cross Abstract: Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence. Yet, despite these advances, critical challenges remain -- most notably, ens...

📖 Read original article


156. Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization ​

Author: Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain brittle in practice, fr...

📖 Read original article


157. Knowledge-Based Design Requirements for Generative Social Robots in Higher Education ​

Author: Stephan Vonschallen, Dominique Oberle, Theresa Schmiedel, Friederike Eyssel
Published: 7/13/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2602.12873v5 Announce Type: replace-cross Abstract: Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as misinformation, overreliance, and privacy violations. Existing frameworks for educational technologies ...

📖 Read original article


158. Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned ​

Author: Zirong Chen, Meiyi Ma
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2602.13241v3 Announce Type: replace-cross Abstract: Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a sustained training crisis: staffing shortages exceed 25% in many centers, and preparing a single new hir...

📖 Read original article


159. The LLMbda Calculus: AI Agents, Conversations, and Information Flow ​

Author: Zac Garby, Andrew D. Gordon, David Sands
Published: 7/13/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CR

arXiv:2602.20064v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as agents: they plan, call tools, read untrusted data, and act on the results. This exposes them to prompt injection: data meant only to be read is obeyed as an instruction. The most principled ...

📖 Read original article


160. SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution ​

Author: Gangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan, Yuhong Liu, Yuxin Yang, Dhruv Parikh, Rajgopal Kannan, Le Cong, Mengdi Wang, Qian Zhang, Viktor Prasanna, Xiangru Tang, Xingyao Wang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2603.13428v3 Announce Type: replace-cross Abstract: Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-running systems, are now entrusted to drive this evolution. Yet, existing benchmarks evaluate agents on i...

📖 Read original article


161. Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation ​

Author: Iakovos-Christos Zarkadis, Christos Douligeris
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, stat.AP, stat.ML

arXiv:2603.17717v4 Announce Type: replace-cross Abstract: Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivotal time for artificial intelligence (AI), with even more sophisticated attacks that utilize advanced te...

📖 Read original article


162. SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation ​

Author: Harshit Joshi, Priyank Shethia, Jadelynn Dao, Monica S. Lam
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.22294v2 Announce Type: replace-cross Abstract: Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual co...

📖 Read original article


163. Tuning Derivatives for Causal Fairness in Machine Learning ​

Author: Filip Edstr"om, Guilherme W. F. Barros, Tetiana Gorbach, Xavier de Luna
Published: 7/13/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.CY, cs.LG

arXiv:2605.05882v2 Announce Type: replace-cross Abstract: Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protected attributes such as race, gender, or age. Classical fairness notions, most notably Statistical Parity ...

📖 Read original article


164. Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue ​

Author: Vardhan Dongre, Dilek Hakkani-T"ur
Published: 7/13/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL

arXiv:2605.12920v3 Announce Type: replace-cross Abstract: Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's evolving understanding of the world. When agents can only partially observe their surrounding...

📖 Read original article


165. AnchorMoE: Interpretable Time Series Classification via Anchor-Routed MoE ​

Author: Tao Xie, Zexi Tan, Haoyi Xiao, Mengke Li, Yiqun Zhang, Yang Lu, Cuie Yang, Yiu-ming Cheung
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.03631v3 Announce Type: replace-cross Abstract: Multivariate time series classification (MTSC) is pivotal in high-stakes domains, such as clinical diagnosis and industrial fault detection, where safe deployment necessitates transparent decision-making. However, isolating the temporal segme...

📖 Read original article


166. Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories ​

Author: Ali Behrouz, Farnoosh Hashemi, Adel Javanmard, Vahab Mirrokni
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.03979v2 Announce Type: replace-cross Abstract: The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific shallow models to more general deep Large Language Models (LLMs). Despite showing promising results in ...

📖 Read original article


167. Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance ​

Author: Tianming Du, Peijie Yu, Sihan Shang, Danli Shi, My Linh Nguyen, Shengbo Gao, Guangyuan Li, Yinghong Yu, Yan Jiang, Qianlong Zhao, Behzad Bozorgtabar, Shaoxiong Ji, Jiazhen Pan, Daniel Rueckert, Jiancheng Yang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.18613v3 Announce Type: replace-cross Abstract: The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistanc...

📖 Read original article


168. ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ​

Author: Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.31650v2 Announce Type: replace-cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion,...

📖 Read original article


169. GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision ​

Author: Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.02360v4 Announce Type: replace-cross Abstract: Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This article presents GAP-GDRNet, a geometry-aware RGB framework built on GDR-Net for a single-target synthe...

📖 Read original article


170. Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language ​

Author: Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.03882v2 Announce Type: replace-cross Abstract: LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identic...

📖 Read original article


171. Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents ​

Author: Abhishek Kumar, Carsten Maple
Published: 7/13/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.03968v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine outputs over many turns. Yet their safety is still often evaluated as if they were chatbots: o...

📖 Read original article


172. SCOReD: Student-Aware CoT Optimization for Recommendation Distillation ​

Author: Haz Sameen Shahgir, Yufei Li, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Frank Shyu, Sandeep Pandey, Luke Simon, Yue Dong, Xi Liu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.05734v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually high reasoning uncert...

📖 Read original article


173. Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs ​

Author: Andrea Bacciu, Andrea Alfarano, Saab Mansour, Amin Mantrach, Marcello Federico
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.06327v2 Announce Type: replace-cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, m...

📖 Read original article


174. WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time ​

Author: Yusen Feng, Bingchen Han, Jiangran Lyu, Kai Liu, Yixin Zheng, Yuxuan Wan, Weiheng Liu, Sun Han, Ruiqin Li, Yulong Zhang, Fangfu Liu, Xuesong Shi, Libin Liu, Yizhou Wang, Zhizheng Zhang, He Wang
Published: 7/13/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.06988v2 Announce Type: replace-cross Abstract: Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a ...

📖 Read original article


175. Riemannian Geometry for Pre-trained Language Model Embeddings ​

Author: Szczepan Konior, Alexandre Quemy, Przemys{\l}aw Klocek, Bart{\l}omiej Sobieski, Gr'egoire Cattan
Published: 7/13/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.07047v2 Announce Type: replace-cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and p...

📖 Read original article


176. Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS-ANS Dynamics ​

Author: Zhoujie Hou, Song Wang, Kexin Lou, Mo Wang, Chen Wei, Quanying Liu
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.07720v2 Announce Type: replace-cross Abstract: Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected by multimodal polysomnography signals including EEG, EOG, EMG, ECG, and respiration. However, existing ...

📖 Read original article


177. Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE ​

Author: Haozhan Tang, Zerui Wang, Yuxian Gu, Song Han, Han Cai
Published: 7/13/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.07740v2 Announce Type: replace-cross Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitu...

📖 Read original article