Trafy

Research

Everything Trafy Intelligence has tracked in Research.

Learning When to Trust via Selective Context Preference OptimizationResearch

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, tr...

Papers with Code · Aug 6, 2026
1 min
Learning When to Trust via Selective Context Preference OptimizationResearch

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misle

arXiv (cs.AI) · Aug 6, 2026
4 min
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature EngineeringResearch

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubr

arXiv (cs.AI) · Aug 6, 2026
4 min
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of NigeriaResearch

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applications affects digital sovereignty, examined through platform transparency as a key indicator of user awareness and control. Using an interpretive approach, the research combines the forensic analysis of selected Android applications with contextual document analysis to identify AI features

arXiv (cs.AI) · Aug 6, 2026
3 min
An Optimal Agnostic PAC AlgorithmResearch

An Optimal Agnostic PAC Algorithm

Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size $n$, for every $0<δ\le 1/2$, with probability at least $1-δ$, \[ L(\widehat h) \le L^*+ 7\cdot10^8\left( \sqrt{\frac{L^*(d+\log(1/δ))}{n}} +\frac{d+\log(1/δ)}{n} \right). \] This settles the sample complexity of agnostic PAC learning up to universal constants at every fixed $L^*$, matching the lower bounds of Devroye, Györfi, and Lugosi [A Probabilistic

arXiv (cs.AI) · Aug 6, 2026
3 min
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information GamesResearch

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-informat

arXiv (cs.AI) · Aug 6, 2026
4 min
The Low Frequency Trap: Video Language Models Fail at Simple Event BookkeepingResearch

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and freque

arXiv (cs.AI) · Aug 6, 2026
4 min
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI AgentsResearch

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism seeks to establish the Safe AI paradigm that compute is an effective governance lever. We situate our work as a compliance or commons overlay on a deployer. One governance period is an extensive form game in which verified human stakeholders arrive sequentially and contribute, on a provision

arXiv (cs.AI) · Aug 6, 2026
4 min
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal TasksResearch

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated st

arXiv (cs.LG) · Aug 6, 2026
4 min
Challenges in Evaluating Explanation Methods for Static and Evolving DataResearch

Challenges in Evaluating Explanation Methods for Static and Evolving Data

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift. Experiences with adapting counterfactuals for this problem are discussed. Finally it is related to the challenges of tracking the co-evolution

arXiv (cs.AI) · Aug 6, 2026
3 min
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent TrajectoriesResearch

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with differe

arXiv (cs.AI) · Aug 6, 2026
4 min
Scalable estimation of VARMA modelsResearch

Scalable estimation of VARMA models

Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass over the entire series. Yet their moving-average term captures with a few parameters what a pure autoregression matches only with many lags. We introduce an estimation framework that removes this computational barrier: each optimization iteration is independent of the series length $T$. The framework combines a partial-autocorrelation reparametriza

arXiv (cs.LG) · Aug 6, 2026
4 min
Scalable estimation of VARMA modelsResearch

Scalable estimation of VARMA models

Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the param...

Papers with Code · Aug 6, 2026
2 min
Optimal Rates for Learning with Monotone AdversariesResearch

Optimal Rates for Learning with Monotone Adversaries

A monotone adversary observes an i.i.d. labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. The learner sees a uniform shuffle of the combined sample and is scored on the original distribution. Every example is correctly labeled, but the insertions depend on the clean sample, so the combined sample is not exchangeable. Larsen, Pabbaraju, and Shetty, who introduced this model, showed that empirical risk minimization attains expected error $O((d/n)\log(n/d))$ for classes of VC dimension $d$, and tha

arXiv (cs.LG) · Aug 6, 2026
4 min
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational DataResearch

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually written by hand. This is a knowledge-acquisition bottleneck that limits the scalability of analytic systems, keeps non-technical users dependent on experts, and is itself error-prone. We present TYTAN, a system for automatically constructing an analytic semantic schema from a relational databa

arXiv (cs.AI) · Aug 6, 2026
4 min
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational AgentsResearch

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by L

arXiv (cs.AI) · Aug 6, 2026
3 min
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard DocumentsResearch

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently ...

Papers with Code · Aug 6, 2026
2 min
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?Research

Does FLAIR super-resolution erase or hallucinate small white-matter lesions?

White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired with thick slices in clinical settings, giving it poor through-plane resolution. Super-resolution (SR) is a widely used method for recovering an isotropic volume from an anisotropic scan. Yet whether applying it prior to WMH segmentation preserves lesion content remains unknown: a model may erase small real lesions or hallucinate absent ones. We used 1-mm isotropic high-resolution (HR) F

arXiv (cs.AI) · Aug 6, 2026
4 min
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward ConstructionResearch

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models

arXiv (cs.LG) · Aug 6, 2026
4 min
UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate RegressionResearch

UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit map re...

Papers with Code · Aug 6, 2026
1 min
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic OperationsResearch

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands of near-identical figures compete in one embedding space, and a figure inherits its unit from a header a median of 13 lines above it -- so a chu

arXiv (cs.AI) · Aug 6, 2026
4 min
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationResearch

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchma

arXiv (cs.AI) · Aug 6, 2026
4 min
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation VectorsResearch

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attri

arXiv (cs.AI) · Aug 6, 2026
4 min
On-Policy Self-Distillation without Any SupervisionResearch

On-Policy Self-Distillation without Any Supervision

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We propose Unsupervised On-Policy Self-Distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseu

arXiv (cs.LG) · Aug 6, 2026
4 min
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality PredictionResearch

QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction

Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches ignore the temporal progression of physiological deterioration and recovery that unfolds throughout a patient's ICU stay. To address this limitation, we introduce QuanTiMedAI, a quantum-agentic framework developed for cardiac arrest mortality prediction using agentic AI guided quantum enhan

arXiv (cs.AI) · Aug 6, 2026
4 min
BaKron: Efficient Quantization with Kronecker-Factored HessiansResearch

BaKron: Efficient Quantization with Kronecker-Factored Hessians

We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian. GPTQ-style adaptive rounding typically uses one-sided information derived from input activations. Two-sided Kronecker-factored Hessian approximations can additionally capture correlations across output coordinates, but applying GPTQ directly in the vectorized weight domain is computationally expensive. Building on the two-sided adaptive-rounding formulation used by BoA and YAQA, we introduce BaKron, an efficient solver that combines a

arXiv (cs.AI) · Aug 6, 2026
3 min
Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival DataResearch

Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data

This work presents a novel attention-based framework for estimating the Individual Probability of Treatment Benefit (IPTB) in survival analysis contexts. The proposed model, called Surv-IPTB, directly quantifies the probability that a specific patient will experience extended survival time under treatment versus control. We reformulate IPTB estimation as a binary classification problem, leveraging pairwise patient comparisons across treatment and control cohorts. The framework incorporates a principled handling of right-censored observations through imprecise probability representati

arXiv (cs.LG) · Aug 6, 2026
4 min
The Tamed Subgradient Unadjusted Langevin Algorithm beyond ConvexityResearch

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked

arXiv (cs.LG) · Aug 6, 2026
3 min
Stochastic Dynamics on Persistence Diagram Space via Reinforcement LearningResearch

Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning

Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure. While substantial progress has been made in the statistical analysis of PDs, existing literature often treats diagrams as static objects and provide limited frameworks for probabilistic modeling and stochastic evolution on PD space. We introduce a reinforcement learning framework for stochastic dynamics on PD space, where diagrams evolve through topology aware local edit operations. The dynamics define controlled Markov processes on spaces of finite PDs with variable cardinality

arXiv (cs.LG) · Aug 6, 2026
3 min
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with ImagesResearch

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We th

arXiv (cs.AI) · Aug 6, 2026
4 min
Improving the Realism of Synthetic Clinical Benchmarks Under Utility ConstraintsResearch

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patie

arXiv (cs.AI) · Aug 6, 2026
4 min
OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and LocationsResearch

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. Data augmentation is a technique widely regarded as a way to improve model training. It includes simple transformations like spatial operations or intensity modifications, but also more advanced synthesis techniques. Their goal is to generate new realistic samples from an existing dataset to diversify the images used during training. Among them, several propose different mixing strategies

arXiv (cs.LG) · Aug 6, 2026
4 min
Hypothesis Testing with Conditional Queries: Learnability and the Value of InteractionResearch

Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We first ask which pairs of distribution classes can be reliably distinguished. We then ask how many additional queries are required to match an adaptive tester when all queried events must be fixed in advance. We show that learnability holds if and only if the two classes have positive separation in their pairwise conditional probabilities. When this separati

arXiv (cs.LG) · Aug 6, 2026
4 min
RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity PredictionResearch

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactions with complex substrates. We propose reaction contrastive learning foundation (RxnCLF), a self-supervised contrastive framework for reaction representation learning. RxnCLF is built on a condensed reaction graph (CRG

arXiv (cs.LG) · Aug 6, 2026
4 min
MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph constructionResearch

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledg

arXiv (cs.LG) · Aug 6, 2026
3 min
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based ModelResearch

Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model

Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets without expert verification and heavyweight pretrained backbones unsuited to on-device use. We introduce RSBdSL38, 10,874 expert-validated images spanning all 38 BdSL hand signs, representing the 51 letters of the Bangla alphabet, recorded from real signers at three special-needs schools across Bangladesh. We propose a lightweight attention bas

arXiv (cs.AI) · Aug 6, 2026
4 min
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture ClassificationResearch

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. In this work, we show that early stopping can overcome this suboptimality: in a Gaussian mixture model with label-flipping noise, GD stopped at an appropriate oracle time achieves minimax-optimal excess zero-one risk for covariance spectra with fast and continuous

arXiv (cs.LG) · Aug 6, 2026
3 min
A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI GovernanceResearch

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning. However, the literature remains fragmented across technique families, model classes, and deployment contexts, making it difficult to compare methods or describe how a trained model has been modified. This survey synthesizes the post-training adaptation literature and introduces a six-dimensional taxonomy organ

arXiv (cs.LG) · Aug 6, 2026
3 min
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning ModelsResearch

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence t

arXiv (cs.AI) · Aug 6, 2026
4 min
Timestep-Conditioned Transformers for Global Weather ForecastingResearch

Timestep-Conditioned Transformers for Global Weather Forecasting

Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorter timesteps (e.g. 1 to 6 hours) finely resolve atmospheric dynamics within the diurnal cycle but increase error accumulation for a given forecast horizon, while longer timesteps (e.g. 24 hours) reduce error accumulation but limit the usability of short-range forecasts where sub-daily predictability is high. In this work, we present GEM-3, a probabilistic global weather model that addresses this trade-off th

arXiv (cs.LG) · Aug 6, 2026
3 min
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image TranslationResearch

PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation

Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-based unpaired translators control preservation through a single global noise or guidance value applied across the image, which cannot separate content to keep from appearance to change. We present PRISM, a GAN-free flow-matching framework that replaces this global control with a learned per-feature gate. The gate's spatial prior is derived from each source feature's standardized distance to the target feature distribution, so features far from th

arXiv (cs.AI) · Aug 6, 2026
4 min
Depth-Guided Video Object Counting in Crowded ScenesResearch

Depth-Guided Video Object Counting in Crowded Scenes

Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline. By integrating depth cues with multi-scale RGB-D cross-attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occ

arXiv (cs.AI) · Aug 6, 2026
3 min
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over NetworksResearch

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots and vehicles operating under real-world physical laws. This stems from their inability to maintain reliable world models for long-horizon planning under uncertainty and generalize to unseen scenarios. In this context, wireless networks, through pervasive sensing and communication, can orchestrate physical intelligence. However, current architectures optimize throughput, latency, and reliab

arXiv (cs.AI) · Aug 6, 2026
4 min
TS-RAG: Retrieval Augmented Generation for Time Series ForecastingResearch

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in time series forecasting tasks. However, most time series models are constrained by limited training data, smaller parameter scales, and a lack of

arXiv (cs.AI) · Aug 6, 2026
3 min
TS-RAG: Retrieval Augmented Generation for Time Series ForecastingResearch

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of re...

Papers with Code · Aug 6, 2026
1 min
Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness EvaluationResearch

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. The resulting human-like robot motion is recognised as a key factor in building trust and enabling natural collaboration in human-robot interaction. This paper presents a framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation. A dataset of 3,142 handwritin

arXiv (cs.LG) · Aug 6, 2026
4 min
Muon on the Stiefel Manifold Admits an Exact Closed-Form UpdateResearch

Muon on the Stiefel Manifold Admits an Exact Closed-Form Update

We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold. This manifold consists of matrices with orthonormal columns and is ubiquitous in machine learning and scientific computing. Existing extensions of Muon to this manifold rely on heuristic, approximate, or iterative updates with varying computational efficiency. We show that the corresponding Stiefel Muon update admits an exact closed-form solution and use this result to develop Skewon, a practical algorithm for orthogonality-constrained optimization with an efficient implementa

arXiv (cs.LG) · Aug 6, 2026
3 min
AI model achieves breakthrough in forecasting cyclonesResearch

AI model achieves breakthrough in forecasting cyclones

WeatherNext enables accurate cyclone forecasts that can give an extra day of warning. Now we are open sourcing the model.

Google DeepMind · Aug 6, 2026
7 min
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic ShippingResearch

Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping

Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires rewa...

Papers with Code · Aug 6, 2026
2 min
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World ModelsResearch

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs an...

Papers with Code · Aug 6, 2026
2 min
TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact ConditionsResearch

TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions

In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged...

Papers with Code · Aug 6, 2026
1 min
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational HardeningResearch

Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening

We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the firs...

Papers with Code · Aug 6, 2026
2 min
BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cellsResearch

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual g...

Papers with Code · Aug 6, 2026
2 min
CodeGrep: An RL-Trained Retrieval Agent for LLM Coding AgentsResearch

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rath...

Papers with Code · Aug 6, 2026
2 min
MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion TransformersResearch

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot o...

Papers with Code · Aug 6, 2026
2 min
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence GroundingResearch

Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding

User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep researc...

Papers with Code · Aug 6, 2026
2 min
Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation TasksResearch

Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks

The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been dev...

Papers with Code · Aug 6, 2026
1 min
Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image EnhancementResearch

Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image Enhancement

Restoring high-fidelity remote sensing imagery from extreme low-light degradation is indispensable for reliable Earth observation and downstream machine visi...

Papers with Code · Aug 6, 2026
1 min
ChainClaw: A Layered Agent Framework for Reliable On-Chain ExecutionResearch

ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain e...

Papers with Code · Aug 6, 2026
1 min
Evidence-Driven Dynamic Visual Selector for Efficient Long Video UnderstandingResearch

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting que...

Papers with Code · Aug 6, 2026
2 min
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure ReadingResearch

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known b...

Papers with Code · Aug 6, 2026
2 min
Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone MappingResearch

Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping

We propose DOME-HDR, a dual-output multi-exposure HDR reconstruction framework that jointly produces a perceptually balanced SDR image and a consistent HDR i...

Papers with Code · Aug 6, 2026
1 min
The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare DecisionsResearch

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce m...

Papers with Code · Aug 6, 2026
1 min
Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended GenerationResearch

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable pl...

Papers with Code · Aug 6, 2026
2 min
APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-TuningResearch

APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning

Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning ...

Papers with Code · Aug 6, 2026
2 min
CoCo-IR: Contextual Composed Image RetrievalResearch

CoCo-IR: Contextual Composed Image Retrieval

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, r...

Papers with Code · Aug 5, 2026
1 min
Argus: A General-Purpose Agentic Runtime for Long-Horizon ReasoningResearch

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and,

arXiv (cs.AI) · Aug 5, 2026
4 min
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context ModelingResearch

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of depend

arXiv (cs.AI) · Aug 5, 2026
3 min
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon ReasoningResearch

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the

arXiv (cs.LG) · Aug 5, 2026
4 min
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon ReasoningResearch

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, ...

Papers with Code · Aug 5, 2026
2 min
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist DomainsResearch

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models o

arXiv (cs.AI) · Aug 5, 2026
4 min
The Loss Does Not See the Basis, but Adam DoesResearch

The Loss Does Not See the Basis, but Adam Does

Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mechanism is available to an optimizer only if that optimizer is gauge-equivariant, a condition necessary for the transfer but not sufficient for low-rank recovery. Gradient descent, momentum, "shared-scalar" Adam, Muon, and Shampoo satisfy it. Adam, RMSProp, and the other coordinate-wise meth

arXiv (cs.LG) · Aug 5, 2026
4 min
Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor EmbeddingsResearch

Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings

Predicting how a subcortical structure's shape will evolve from a few prior scans could support prognosis and clinical-trial enrichment. Existing longitudinal mesh predictors either extrapolate shape trajectories via high-dimensional embeddings or regress vertex deformations directly. We instead predict the surface's intrinsic geometry in continuous time: a single per-structure graph network predicts the future per-vertex first fundamental form (metric tensor) for an arbitrary causal multiple-visit history and an arbitrary prediction horizon, conditioned on a Fourier encoding of the

arXiv (cs.LG) · Aug 5, 2026
4 min
OPD-V: Visual On-Policy Self-Distillation with Modality BalanceResearch

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we con

arXiv (cs.AI) · Aug 5, 2026
3 min
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuantResearch

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propose Subsampled Stochastic TurboQuant (SSTQ), a framework that combines overcomplete equal-norm tight frames, coordinate subsampling, and privacy-aware one-dimensional quantization. SSTQ includes two variants: a Flat Randomized Response version and a Metric-Aware Laplace version, the latter b

arXiv (cs.AI) · Aug 5, 2026
3 min
Chained Recursive Language Models for Multi-Iteration ReasoningResearch

Chained Recursive Language Models for Multi-Iteration Reasoning

Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This becomes particularly difficult in tasks that require extraction, counting, ordering, or multi-hop reasoning, where an early mistake can propagate until the final response. In this work, we propose Chained Recursive Language Models (Chained RLM), an inference-time architecture, in which the same underlying model is called repeatedly as a seque

arXiv (cs.AI) · Aug 5, 2026
4 min
DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model DiscoveryResearch

DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery

Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an ite

arXiv (cs.LG) · Aug 5, 2026
4 min
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident RecognitionResearch

Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations. This setting remains underexplored, with limited benchmarks and few methods designed for the privacy, efficiency, and generalization demands of real-world deployment. We introduce a novel hybrid benchmark combining generative CCTV-style videos with real-world classroom pose data, and propose a lightweight, but robust motion-reasoning framework motivated by the observation that many incidents differ

arXiv (cs.AI) · Aug 5, 2026
4 min
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident RecognitionResearch

Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognitio...

Papers with Code · Aug 5, 2026
1 min
Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean ShiftResearch

Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean Shift

The Subspace Constrained Mean Shift (SCMS) algorithm is a popular nonparametric method for extracting density ridges, which serve as a low-dimensional representation of high-dimensional data. It is a widely held belief in the literature that SCMS trajectories converge to the classical density ridge, which we call the "static ridge", defined via the density gradient and the eigenvalues and eigenvectors of the density's Hessian. In this paper, we demonstrate that this assumption does not hold in general, as the static definition fails to account for the rotation of the trailing eigensp

arXiv (cs.LG) · Aug 5, 2026
4 min
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement LearningResearch

Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content

arXiv (cs.LG) · Aug 5, 2026
4 min
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depthResearch

Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth

Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the output distributions accessible to shallow unitary Born models. Introducing stochasticity into a unitary quantum Born model can improve the empirical generative performance of the resulting channel model and, for a restricted small-scale architecture, has been proven to represent a strictly larger family of distributions than its unitary counterpart. However, whether such randomness provides a provable separation at fixed shallow depth for

arXiv (cs.AI) · Aug 5, 2026
4 min
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument GraphsResearch

CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs

AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, functional, psychosocial, and environmental needs. However, many AI systems present recommendations as fixed outputs, limiting stakeholders' ability to inspect, challenge, and revise plans when they conflict with clinical judgment, patient values, or real-world feasibility. We present CoPlan - a Co-Intelligent and Contestable Interface for Human-AI Care Planning. CoPlan uses a multi-agent workflow in which specialized AI agents generate candidate interven

arXiv (cs.AI) · Aug 5, 2026
4 min
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket PruningResearch

BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art performance in natural language processing with a large pre-trained model such as BERT is expensive and time-consuming, carries a large carbon footprint, and is difficult to realize on machines with minimal computational capability. This creates a barrier to training complex models for resource-constrained languages such as Bengali. However, in a complex neural model, not all edges are equally impactful, and the contributions of some of

arXiv (cs.LG) · Aug 5, 2026
4 min
Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent Flow-matchingResearch

Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent Flow-matching

Data assimilation (DA) uses Bayesian inference to update the state of a numerical forecast model with observed data. In this study, we propose a fundamentally different, unified approach to atmospheric data assimilation. We use latent video flow-matching to sample temporally consistent trajectories from a prior trained using ERA5 reanalysis (69 variables over an 8-day window). We also use posterior sampling to assimilate real observation sources, such as those from the NOAA Integrated Global Radiosonde Archive and the Integrated Surface Database. Because the prior generates a continu

arXiv (cs.LG) · Aug 5, 2026
3 min
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit AssignmentResearch

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-l

arXiv (cs.AI) · Aug 5, 2026
4 min
Hierarchical Graph Memory for LLM Agents with Path-level Localization and RewriteResearch

Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite

Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedback continue to arrive. Recently, graph memory has been adopted to offer structural organization for multi-hop retrieval and reasoning. However, existing methods store all memories in a flat graph, and accumulated historical memories can introduce irrelevant contexts and increase the cost of evidence selection during retrieval. Moreover, they typically update memory units independently, requiring repeated unit-wise rewrite to cover related changes.

arXiv (cs.AI) · Aug 5, 2026
4 min
MALT: Lightweight Curvature-Aware Muon via Diagonal PreconditioningResearch

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon mitigates gradient anisotropy, it does not explicitly account for the curvature geometry of the loss landscape and may therefore remain sensitive to curvature anisotropy. We bridge this gap by proposing MALT (Muon Augmented by Lightweight Two-sided Preconditioning), which uses lightweight diagonal preconditioners to reduce the sensitivity of Muon to curvature anisotropy. Specifically, MALT uses two-sided diagon

arXiv (cs.LG) · Aug 5, 2026
3 min
Item Response Theory for AI SafetyResearch

Item Response Theory for AI Safety

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evalua

arXiv (cs.AI) · Aug 5, 2026
4 min
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment SelectionResearch

Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection

Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information gain per unit cost or a learned plausibility score. We identify a structural limitation of this approach. Some actions are constructive: they acquire an epistemic capability (an instrument, assay, pipeline, simulator, or abstraction) whose value lies not in the information returned immediately but in the future actions it makes available. W

arXiv (cs.AI) · Aug 5, 2026
4 min
Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous ControlResearch

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We introduce Prefix-Optimal Generative Policies (POGP), a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain. The prefix value function serves two purposes: it provides an auxiliary training objective that e

arXiv (cs.LG) · Aug 5, 2026
4 min
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention LearningResearch

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning

Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of rollouts to every task and trajectory state, even though some rollouts provide much more useful learning signals than others. Recent work has started to treat rollout generation as an adaptive decision, but two important limitations remain. First, intervention strategies are often based on fixed heuristics and therefore cannot adjust as the policy changes during training. Second, these methods usually decide

arXiv (cs.LG) · Aug 5, 2026
4 min
MultiPathFormer: Towards a Foundation Model for Multipath Wireless PropagationResearch

MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation

Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimation, beam prediction, and localization based on wireless signals. Existing wireless foundation models typically pretrain on channel tensors using masked reconstruction over subcarriers, antennas, or time but ignore the physical characteristics of wireless propagation. In this work, we propose to instead use multipath propagation as the fundamental pretraining object. We present MultiPathFormer, an autoregressive foundation model that represents ea

arXiv (cs.AI) · Aug 5, 2026
4 min
MultiPathFormer: Towards a Foundation Model for Multipath Wireless PropagationResearch

MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation

Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimation, beam predicti...

Papers with Code · Aug 5, 2026
2 min
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly DetectionResearch

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-based VAD, which focuses on motion dynamics rather than raw video data. However, existing pose-based approaches model human behavior in continuous latent spaces, limiting their ability to learn compact motion patterns necessary for robust behavior analysis. We address this by proposing Vecto

arXiv (cs.AI) · Aug 5, 2026
4 min
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language ModelsResearch

Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models

Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can support risk-controlled deferral, evaluating eleven instruction-tuned models from three families, 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA with 25,168 local predictions. Three theoretical results delimit what calibration can provide: strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC;

arXiv (cs.AI) · Aug 5, 2026
4 min
Hardware Design and Security in the Era of Chiplets and LLMsResearch

Hardware Design and Security in the Era of Chiplets and LLMs

The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) flows. While these paradigms offer unprecedented benefits in yield, modularity, design productivity, etc., they radically expand the hardware attack surface. This paper provides a unified analysis of these frontiers, ranging from attacks on chiplet systems (including hardware stacks for LLM acceleration) across architectural, logical, and physical levels, to various exploits against

arXiv (cs.AI) · Aug 5, 2026
3 min
RepairFormer: Automated Repair of Structured Inputs Using TransformersResearch

RepairFormer: Automated Repair of Structured Inputs Using Transformers

Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can cause parsers to reject otherwise useful data. Repairing such inputs is important because malformed configuration, program, and data files can interrupt testing, analysis, deployment, and downstream automation even when most of the original content remains intact. Existing repair techniques can produce structurally valid inputs, but they often rely on deletion or repeated search, which may lose original content and result in semantic incorrectness

arXiv (cs.AI) · Aug 5, 2026
3 min
MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary AtmospheresResearch

MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres

We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars. While GraphCast achieves state-of-the-art performance for terrestrial forecasting, its applicability to non-Earth environments remains unexplored. Using the Mars Climate Database (MCD), which provides global atmospheric fields across vertical altitude levels (similar to Earth pressure levels), we evaluate zero-shot and fine-tuned GraphCast predictions of Martian temperature and wind fields. Zero-shot forecasts produce

arXiv (cs.AI) · Aug 5, 2026
4 min
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR SimulationsResearch

The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the co

arXiv (cs.AI) · Aug 5, 2026
4 min
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR SimulationsResearch

The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depict...

Papers with Code · Aug 5, 2026
2 min
Gradient Immunity: Null-Space Resistance to Malicious Fine-TuningResearch

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, and therefore do not directly address the setting we study: a provider controlled partially protected open-weight (PPOW) release setting in which most weights remain trainable while a small safety-critical component is preserved at release. We propose a Unidirectional Safety Gate (USG), instantiated as a Null Space Cubic Layer together wi

arXiv (cs.AI) · Aug 5, 2026
4 min
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic SystemResearch

SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System

Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy. For the same SpMM on the same matrix, cuSPARSE exhibits a 350x performance gap between CSR and Blocked-ELL. Our study of multiple data formats, specialized systems, and sparse compilers shows that no single implementation consistently dominates across sparsity patterns and operators. This motivates a system that can adapt its representation, execution strategy, and hardware mapping to each wo

arXiv (cs.LG) · Aug 5, 2026
4 min
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score RerankingResearch

From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking

Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM co

arXiv (cs.AI) · Aug 5, 2026
4 min
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI AugmentationResearch

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images. Current AI-assisted annotation tools often lack assistance or rely on one-way workflows where experts have to perform extra manual calibrations to improve AI models, resulting in limited efficiency. To address this, we propose Bidirectional Human-AI Augmentation(BiHAA), a closed-loop framework in which skills and domain knowledge base evolve through re

arXiv (cs.AI) · Aug 5, 2026
4 min
Canonical Joint Energy-Based Model on CIFAR-10: failure modes and practical indistinguishability of Predictor-Corrector and SGLD samplersResearch

Canonical Joint Energy-Based Model on CIFAR-10: failure modes and practical indistinguishability of Predictor-Corrector and SGLD samplers

Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection. Canonical JEM training relies on stochastic gradient Langevin dynamics (SGLD); a theoretically motivated alternative, the Predictor-Corrector (PC) sampler, has not previously undergone a systematic replication test on the canonical model. We reproduce canonical JEM on WideResNet-28-10 without normalisation layers on two independent runs and test whether PC retains its theoretical advantage without an annealed noise schedule, across three protoco

arXiv (cs.LG) · Aug 5, 2026
4 min
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid loadResearch

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as critical. Determinism, reproducibility, and auditability are engineering requirements rather than optional extras. STLF is no longer purely an accuracy problem. It is also a software-engineering and compliance problem. This paper describes results from a 41-day live challenge that evaluated a complete STLF pipeline for the aggregated German transmission-grid load. The pipeline is based on the open-source Python library spotforecast2-s

arXiv (cs.AI) · Aug 5, 2026
4 min
Link prediction on multi-relational graphs from an influence propagation perspectiveResearch

Link prediction on multi-relational graphs from an influence propagation perspective

Predicting the existence and type of links (edges) between nodes in a multi-relational graph is key for applications from social interaction prediction to knowledge relationship identification. Enhancing local features with relevant global information is crucial for accurate link prediction, yet it remains challenging. We address this by modeling the relationship between node pairs as node influence. That is, whether the node influence can be propagated and what type of influence is propagated indicates where and what type the edge is, which will be the most relevant local and global

arXiv (cs.LG) · Aug 5, 2026
4 min
Revealed Rationality: Label-Free Evaluation and Regularization from Representation TheoremsResearch

Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems

Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implicat

arXiv (cs.AI) · Aug 5, 2026
3 min
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context RoutingResearch

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operati...

Papers with Code · Aug 5, 2026
2 min
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics ApplicationsResearch

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applicati...

Papers with Code · Aug 5, 2026
2 min
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical CoordinationResearch

A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination

Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses h...

Papers with Code · Aug 5, 2026
1 min
The Neural Echo: A Signal Processing Perspective for Understanding Neural NetworksResearch

The Neural Echo: A Signal Processing Perspective for Understanding Neural Networks

We introduce the neural echo as a tool for understanding the behavior of neural networks. It generalizes the model-based concepts of impulse responses, diffu...

Papers with Code · Aug 5, 2026
2 min
ContextWeave: A Real-World Workflow BenchmarkResearch

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or...

Papers with Code · Aug 5, 2026
1 min
Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop EvaluatorResearch

Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator

Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus...

Papers with Code · Aug 5, 2026
2 min
Segmentation Pre-training for Label-Efficient Lumbar Spine Degeneration GradingResearch

Segmentation Pre-training for Label-Efficient Lumbar Spine Degeneration Grading

Automated assessment of degenerative pathology in the lumbar spine on magnetic resonance imaging (MRI) requires access to large-scale datasets of expert-anno...

Papers with Code · Aug 5, 2026
1 min
MGSB: Manifold Gated Signature Branch Pressure-Domain Baseline Architecture for Two-Phase Pipeline Flows Under Distributional ShiftResearch

MGSB: Manifold Gated Signature Branch Pressure-Domain Baseline Architecture for Two-Phase Pipeline Flows Under Distributional Shift

Leak detection models for multiphase pipelines often degrade when deployed under flow regimes that differ from training. Existing evaluations typically asses...

Papers with Code · Aug 5, 2026
1 min
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary ToolsResearch

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model C...

Papers with Code · Aug 5, 2026
2 min
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label ClassificationResearch

Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification

Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize b...

Papers with Code · Aug 5, 2026
1 min
The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read DivergencesResearch

The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences

We study distributionally robust PAC learning for the $0$--$1$-loss, where adversarial perturbations of the data distribution are constrained by a Cressie--R...

Papers with Code · Aug 5, 2026
2 min
An entropic explanation of insistence on sameness in autismResearch

An entropic explanation of insistence on sameness in autism

An information theory-based framework is proposed in attempt to explain insistence on sameness in autism as an instance of a general behavior pattern in whic...

Papers with Code · Aug 5, 2026
2 min
Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State NetworksResearch

Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks

Echo State Networks (ESNs) offer an efficient framework for temporal prediction, but their randomly initialized reservoirs are often over-parameterized and d...

Papers with Code · Aug 5, 2026
1 min
What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent SkillsResearch

What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills

Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study skill...

Papers with Code · Aug 5, 2026
1 min
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery RobustnessResearch

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current un...

Papers with Code · Aug 5, 2026
1 min
RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty CareResearch

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course ...

Papers with Code · Aug 5, 2026
2 min
AFD-Ledger: Deployment Provisioning for Attention--FFN DisaggregationResearch

AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation

Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) language models. Whil...

Papers with Code · Aug 5, 2026
2 min
Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video PrefixesResearch

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, ...

Papers with Code · Aug 5, 2026
2 min
Generative Optimization for Incentivized Advertising with Global Level ConstraintsResearch

Generative Optimization for Incentivized Advertising with Global Level Constraints

Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes ...

Papers with Code · Aug 5, 2026
1 min
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance GapResearch

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle...

Papers with Code · Aug 5, 2026
2 min
EdgeLM: Edge Demonstrations for Language Models' Table UnderstandingResearch

EdgeLM: Edge Demonstrations for Language Models' Table Understanding

Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Existing r...

Papers with Code · Aug 5, 2026
1 min
Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-TuningResearch

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. H...

Papers with Code · Aug 5, 2026
2 min
An Analysis and Implementation of Seam Carving for Content-Aware Image ResizingResearch

An Analysis and Implementation of Seam Carving for Content-Aware Image Resizing

Seam carving is a classical content-aware image resizing operator that modifies the width or height of an image by repeatedly removing (or inserting) seams, ...

Papers with Code · Aug 5, 2026
1 min
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)Research

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclus...

Papers with Code · Aug 5, 2026
2 min
Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic EmbeddingsResearch

Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings

Translating wordplay across languages has long challenged both professional translators and machine translation systems. We investigate three approaches to t...

Papers with Code · Aug 5, 2026
1 min
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated ReasoningResearch

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the t

arXiv (cs.AI) · Aug 4, 2026
3 min
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and ReproducibilityResearch

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample completed candidates and aggregate them through voting or verification, or search over unfinished partial states. These algorithms differ in their statistical structure, compute accounting, and failure modes. Treating these procedures as interchangeable under a single scalar "budget," or reporting accuracy without the inference protocol that pr

arXiv (cs.AI) · Aug 4, 2026
4 min
Assessment of Conditional Diffusion Model for Synthetic Histopathology Image GenerationResearch

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This work investigates and addresses limitations in existing evaluation metrics, investigating an approach for assessing synthetic histopathology image quality through domain-specific metrics and downstream task validation. We show that conventional synthetic data evaluation metrics such as Frechet Inception Distance (FID) and Inception Score (IS) may h

arXiv (cs.LG) · Aug 4, 2026
4 min
Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?Research

Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs) can recover such semantics from heterogeneous C/C++ context and realize them as validated, contract-preserving artifacts. We introduce SeGaBench, an executable benchmark containing 100 synthetic and 20 source-backed cases spanning low-level assumptions, data-structure invariants, and high-level semantic lifting. Each case includes hidden enabling semantics, an oracle artifact, correctness and semantic valid

arXiv (cs.AI) · Aug 4, 2026
3 min
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch AgentResearch

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline w

arXiv (cs.AI) · Aug 4, 2026
4 min
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct ReasoningResearch

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed tr

arXiv (cs.AI) · Aug 4, 2026
4 min
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input PerturbationsResearch

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do they impact an LLM's performance? In this paper we present HIVE (Human Input-Variation Engine), a suite of voice transcription perturbations and QWERTY keyboard perturbations. We use HIVE to evaluate how robust models are to these perturbations. We present seven findings. (i) Voice transcription perturbations lower accuracy across every i

arXiv (cs.AI) · Aug 4, 2026
4 min
Information-Geometric Forward Policy Training in GFlowNetsResearch

Information-Geometric Forward Policy Training in GFlowNets

Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical

arXiv (cs.LG) · Aug 4, 2026
4 min
Separating quantum circuits from classical LLMsResearch

Separating quantum circuits from classical LLMs

Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional separations between low-depth quantum computation and the corresponding bounded-resource classical language-model architectures in both regimes. Concretely, we exhibit the following: 1. Distributional separation. We give a distribution that is sampleable by $\textsf{QNC}^0$ circuits (i.e., a family of constant-depth quantum circuits consisting of bounded fan-in gates) that no constant-round diffusion language m

arXiv (cs.AI) · Aug 4, 2026
3 min
Interpretable Adaptive Sampling for LLM Test-Time ScalingResearch

Interpretable Adaptive Sampling for LLM Test-Time Scaling

Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier

arXiv (cs.AI) · Aug 4, 2026
3 min
A game theory for foundation models shows new paths to rational cooperation through similarity inferenceResearch

A game theory for foundation models shows new paths to rational cooperation through similarity inference

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking findi

arXiv (cs.AI) · Aug 4, 2026
4 min
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English TutoringResearch

TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring

Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effective ESL tutoring, however, requires more than fluent response generation: a tutor must select an appropriate pedagogical action based on learner behavior and dialogue context. Human-tutoring research offers principles for adaptive support, but they are often task-specific and remain insufficiently integrated into LLM-based ESL tutor training and evaluation. We present TACT (Taxonomy-Aligned Conversational Tutor), a human-grounded framework for p

arXiv (cs.AI) · Aug 4, 2026
4 min
Muon Meets Mamba: Spectral Optimization for State Space ModelsResearch

Muon Meets Mamba: Spectral Optimization for State Space Models

Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. We compare Muon with AdamW on Mamba-2 130M under a controlled protocol that varies only which weight groups are trained with Muon. The benefit is localized. Muon on the output projection alone beats Muon on the input projection or on both. The advantage is mainly one of token efficiency. It holds

arXiv (cs.LG) · Aug 4, 2026
3 min
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and CompressibilityResearch

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to im

arXiv (cs.AI) · Aug 4, 2026
4 min
Latent Reward Registers for Diffusion Preference AlignmentResearch

Latent Reward Registers for Diffusion Preference Alignment

Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process. We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents by prepending learnable, position-free register tokens to the input sequence of a frozen Diffusion Transformer (DiT). This independent readout mechanism extracts latent reward evidence without altering the generator's hidden states or

arXiv (cs.LG) · Aug 4, 2026
4 min
Robust Low-Tubal-Rank Tensor Completion under Cross-Concentrated SamplingResearch

Robust Low-Tubal-Rank Tensor Completion under Cross-Concentrated Sampling

Tensor cross-concentrated sampling (t-CCS) bridges entrywise sampling and t-CUR slice-wise sampling by observing entries only within selected horizontal and lateral slices. Existing t-CCS completion methods, however, assume that the observations are free of gross corruption. In this work, we study robust recovery of a third-order low-tubal-rank tensor from partial t-CCS observations contaminated by sparse, arbitrarily large outliers. We propose Robust Iterative t-CUR (R-ItCUR), a tensor-native algorithm that partitions the sampled tensor cross into two exterior blocks and an intersec

arXiv (cs.LG) · Aug 4, 2026
3 min
A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle TissuesResearch

A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle Tissues

Engineered Skeletal Muscle Tissues (ESMs) have become a key structure for biomedical disease modeling and pharmacological screening, yet their functional characterization often relies on simplistic metrics like peak force, discarding critical kinetic information. This is partially due to the high level of mathematical complexity which mechanistic models introduce to capture these dynamics. Hence, exactly the complexity prevents scalable application and widespread adaptation in the field. Here we present a Physics-Flavored Neural Network (PFNN) that automates the kinetic phenotyping o

arXiv (cs.LG) · Aug 4, 2026
3 min
PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly DetectionResearch

PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection

Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings. While transforming time series into images has shown success in forecasting and classification, it remains unclear how multivariate, high-dimensional series should be mapped to multi-channel images and whether vision backbones can match time-domain baselines in TSAD. We introduce PRISM, a plug-and-play meta-workflow enabling systematic construction and evaluation of im

arXiv (cs.AI) · Aug 4, 2026
4 min
The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight InterconnectionsResearch

The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections

This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and apply prompt-dependent transformations whose parameters are generated during inference. We call this form of computation SIDPP: Sequence-level Interactive Dynamic Parallel Processing. The Transformer is interpreted as a system that transforms concepts by means of concepts. Token vectors are the concepts to be transformed; parameterized transfor

arXiv (cs.AI) · Aug 4, 2026
4 min
Equivariant Music TransformerResearch

Equivariant Music Transformer

Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models become progressively less equivariant as they scale in size or train longer. This suggests that in standard music transformers, additional model capacity is allocated to memorizing absolute patterns rather than capturing shared musical structures. In this paper, we propose t

arXiv (cs.AI) · Aug 4, 2026
3 min
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video UnderstandingResearch

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to se

arXiv (cs.AI) · Aug 4, 2026
4 min
Implementing Causal Perception: Competing SCMs and Situated FairnessResearch

Implementing Causal Perception: Competing SCMs and Situated Fairness

Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distributions, including the hypothetical distributions implied by each agent's SCM under the same set of interventions. It shapes how agents reason about the system and how they perceive its fairness. Causal perception is a promising probabilistic framework, but it has remained purely theoretical. This work provides the first implementation of the causal perception framework of Álvarez and Ruggieri (2025). We operationalize structural (agents disagree on

arXiv (cs.AI) · Aug 4, 2026
3 min
Trajectory inference via Acceleration MatchingResearch

Trajectory inference via Acceleration Matching

Trajectory inference is a fundamental problem in many scientific domains: given a collection of unpaired snapshots of observations at discrete time points, the goal is to generate smooth trajectories that best resemble and interpolate the data. Existing algorithms exhibit computational challenges: they either rely on preprocessing subroutines to enforce smoothness or on simulation-based training objectives, both of which can be expensive. In order to overcome these limitations, we propose a new algorithm called Acceleration Matching (\texttt{AM}). Our approach consists of lifting the

arXiv (cs.LG) · Aug 4, 2026
3 min
Sparse Weight Decomposition for Efficient Circuit ExtractionResearch

Sparse Weight Decomposition for Efficient Circuit Extraction

Dense pretrained transformers do not naturally expose interpretable units for circuit extraction. Existing approaches obtain such units by learning auxiliary sparse representations or training sparse models, incurring substantial additional computation while potentially introducing a fidelity gap between the representation being analyzed and the original pretrained model. We propose Sparse Weight Decomposition (SWD), which reparameterizes pretrained linear projections by factorizing each weight matrix into two sparse factors whose shared intermediate coordinates serve as individually

arXiv (cs.LG) · Aug 4, 2026
4 min
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social TheoryResearch

Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory

As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. Instead, systems must be able to recognize, represent, and respond to multiple legitimate perspectives. This has led to growing interest in pluralistic alignment, which seeks to move beyond one-size-fits-all models of appropriate behaviour. However, current approaches often lack a clear account of how values are socially organized, contested, and coordinated in practice. In this paper, we argue that social theory provides es

arXiv (cs.AI) · Aug 4, 2026
4 min
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV TrackingResearch

When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking

Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficiency. Adaptive Transformer Trackers, which leverage an input-dependent dynamic routing architecture, have emerged as a representative solution to this challenge. However, we reveal that behind this computation-on-demand flexibility hides a critical structural flaw: the Lipschitz singularity of computational path decisions, which has an unbounded local Lipschitz constant at discrete layer-skipping decision boundaries. This mathematical d

arXiv (cs.AI) · Aug 4, 2026
4 min
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill ReuseResearch

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B->32B, one source layer explains 56% of variance in the target's keys and 32% in values, rising to 7

arXiv (cs.LG) · Aug 4, 2026
4 min
Intertemporal Preference Steering in Qwen3 via Contrastive Activation AdditionResearch

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon direc

arXiv (cs.AI) · Aug 4, 2026
3 min
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented MeasurementResearch

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vision-Language Models (VLMs) treat these as separate problems, if they address them at all, leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its genera

arXiv (cs.AI) · Aug 4, 2026
4 min
Omega-S: A Functional Resilience Index for LLM Fine-TuningResearch

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existing training loop and adds under 4% to the cost of a step. Retention. On Llama-3-8B with LoRA, fine-tuned from code to prose and measured by HumanEval over ten seeds, Omega-S retains more of the original capability than no regularisation on 9 of 10 seeds (0.173 -> 0.238 absolute pass@1; sign test one-sided p=0.

arXiv (cs.LG) · Aug 4, 2026
4 min
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial ReasoningResearch

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function famil

arXiv (cs.AI) · Aug 4, 2026
3 min
Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint DistributionResearch

Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint Distribution

Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications. Recent synthetic grid generation methods have improved structural realism and operational feasibility by incorporating engineering knowledge through post-generation validation, optimization, or physics-aware generation. However, generated scenarios may still exhibit low AC feasibility and robustness, limiting their practical value for downstream power-system studies. This paper proposes a feasibility-aware distribution-learning framework t

arXiv (cs.LG) · Aug 4, 2026
4 min
Enhancing VLM Reward Models Through Structure-Aware Fine-TuningResearch

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineering. Although promising, these rewards are often noisy and unreliable, limiting their direct utility during deployment. We present Structure-Aware Fine-Tuning (SAFT), a simple, self-supervised method that refines these imperfect reward signals online without access to ground-truth supervision. SAFT leverages intrinsic structural priors to r

arXiv (cs.AI) · Aug 4, 2026
4 min
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?Research

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execu

arXiv (cs.AI) · Aug 4, 2026
3 min
GENESIS: Towards Explainable Causal DiscoveryResearch

GENESIS: Towards Explainable Causal Discovery

Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-wo

arXiv (cs.AI) · Aug 4, 2026
4 min
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM AdvisoriesResearch

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is supported by the available evidence, permitted under the stated authority and procedure, and acceptable under the plant-specific consequence checks encoded in the selected evaluation profile. In this report, \emph{safety-governed} means that eligibility is determined through explicit, non-compensatory checks derived from a versioned pla

arXiv (cs.AI) · Aug 4, 2026
3 min
CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical EvidenceResearch

CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence

Emergency triage requires reliable decisions within a short time period. However, the available electronic health record (EHR) data, including structured data and clinical text, are often incomplete, unreliable, and inconsistent. This makes machine learning (ML)-based triage prediction more challenging, as existing ML models typically rely on complete and reliable EHR data to accurately predict patients' acuity levels. To address this, we propose confidence- and reliability-aware selective triage (CRS-Triage) to predict patients' acuity levels with a confidence score. By comparing th

arXiv (cs.LG) · Aug 4, 2026
4 min
Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural LanguageResearch

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-training often causes models to overfit to chemical syntax, catastrophically forgetting their foundational semantic capabilities. To address this challenge, we introduce CheMatE, a chemistry-oriented embedding model that jointly captures molecular structure and domain-specific natural language within the same representation space. Built on a ModernBERT backbone, CheMatE

arXiv (cs.LG) · Aug 4, 2026
4 min
Quantization Effects on Biomedical LLM ReliabilityResearch

Quantization Effects on Biomedical LLM Reliability

When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables. We present a controlled evaluation of three Mistral-7B variants (Base, BioMistral, and Instruct) on PubMed RCT sentence classification (n=2000) under FP16, INT8, and INT4 precision using four answer-text prompt templates. Our primary finding is that the probability extraction protocol dominates apparent calibration. Switching from su

arXiv (cs.LG) · Aug 4, 2026
4 min
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANsResearch

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term

arXiv (cs.LG) · Aug 4, 2026
4 min
Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its ScalingResearch

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two propagation regimes - spike-and-suppress (Phi-3.5, Gemma-2-9B) and late-accumulation (Llama-3, Mistral, Qwen2.5-7B) - and on the two models meeting an

arXiv (cs.LG) · Aug 4, 2026
4 min
Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence LayersResearch

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. Five widely deployed agent workflow frameworks answer differently, none exposes a machine-checkable contract, and behavior violates even the fragments they state. The RESUME CONTRACT states six properties over the persistence API (prefix continuation, effect exactly-once, fork determinism, checkpoint validity, consume-once, recovery determinism), plus fork-intent and liveness obligations. A TLA+ model checks a reference s

arXiv (cs.LG) · Aug 4, 2026
4 min
Geo-Embed: Towards Unified Multimodal Embeddings for Urban UnderstandingResearch

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. Firs

arXiv (cs.LG) · Aug 4, 2026
3 min
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review QualityResearch

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understa...

Papers with Code · Aug 4, 2026
1 min
Towards Robust Tool Use in Agents via Experience-Driven Adaptive GuidanceResearch

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role...

Papers with Code · Aug 4, 2026
2 min
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark GainsResearch

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach...

Papers with Code · Aug 4, 2026
2 min
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language AnalysisResearch

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis

Plant root phenotyping is fundamental to understanding below-ground structures, optimizing crop management, and improving agricultural sustainability. This p...

Papers with Code · Aug 4, 2026
2 min
AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill CoachingResearch

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based ...

Papers with Code · Aug 4, 2026
1 min
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery DetectorsResearch

V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors

As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use vi...

Papers with Code · Aug 4, 2026
2 min
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM PruningResearch

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.1\% gain, while th...

Papers with Code · Aug 3, 2026
2 min
Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable FrameworkResearch

Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework

Artificial intelligence (AI) is increasingly central to power and energy systems, supporting modeling, forecasting, optimization, and control. Yet most existing works emphasize specialized applications and offer little reusable material for newcomers or interdisciplinary learners, who increasingly rely on large language models rather than building their own. This gap points to a need for engineering-grounded AI (EGAI), in which AI workflows follow established engineering and power-system domain rules rather than acting as task-agnostic black boxes. Motivated by a community survey of

arXiv (cs.AI) · Aug 3, 2026
4 min
onepot-Bench 0: towards lab-aware in silico chemistry benchmarksResearch

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evalu

arXiv (cs.LG) · Aug 3, 2026
3 min
The Condition-Number Barrier in Sparse Least SquaresResearch

The Condition-Number Barrier in Sparse Least Squares

In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be improved by a polynomial-time algorithm. We establish their conjectured lower bound for least-squares objectives, conditional on the randomized exact-volume Small-Set Expansion Hypothesis in the weighted regular-graph formulation of Raghavendra, Steurer, and Tulsiani [RST12]. Concretely, for every fixed $γ\in(0,1]$, there is no randomized polynomial-time algorithm that, with probability at least $2/3$, returns a vector $x$ such that, writ

arXiv (cs.LG) · Aug 3, 2026
3 min
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent ReasoningResearch

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shape subsequent reasoning. We introduce GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated conti

arXiv (cs.LG) · Aug 3, 2026
4 min
UEmbed: Unified Sparse and Dense Multimodal EmbeddingsResearch

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forw

arXiv (cs.AI) · Aug 3, 2026
4 min
CoWAM: Coordination Contracts for Selective Policy Intervention with WAMsResearch

CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs

World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future alone does not justify changing the action that a bimanual policy would execute. We present CoWAM, a selective intervention layer that expresses synchronization, role compatibility, and collision convergence as coordination contracts. Each contract combines typed admissibility checks with event-conditioned verification and calibrated intervention gates. CoWAM preserves the nominal action unless an alternative satisfies every active obligation and provides a clear, low-r

arXiv (cs.AI) · Aug 3, 2026
3 min
Smooth Reparameterizations of Functions on Simplicial Product Spaces: Applications to Probabilistic Tensor Decomposition and Functional Data RegistrationResearch

Smooth Reparameterizations of Functions on Simplicial Product Spaces: Applications to Probabilistic Tensor Decomposition and Functional Data Registration

We consider optimization problems defined on product spaces of simplices. Examples of this class of problems include learning low-rank discrete multivariate probability distributions via simplex constrained tensor decomposition and performing functional data registration under the Square Root Velocity Function (SRVF) representation. In this work, we demonstrate the feasibility of replacing the product simplex with a smooth, elementwise strictly convex reparameterization, resulting in an unconstrained optimization problem on a manifold. We show that performing such a reparameterizatio

arXiv (cs.LG) · Aug 3, 2026
3 min
Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation QualityResearch

Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality

Diffusion models rely on stochastic inputs, yet on finite-precision hardware, the "randomness" they consume is realized as deterministic numerical orbits generated by pseudorandom rules. Accessible orbit structure can become a learnable input and affect both training and generation because the realized loss and its gradient depend on the concrete pseudorandom values consumed at each optimization step. A small multilayer perceptron predicts the next value of an orbit from its recent history, measuring general sequence predictability. A diffusion probe replaces real images with online

arXiv (cs.LG) · Aug 3, 2026
4 min
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane PoliciesResearch

AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agentic AI promises to automate this search. Off the shelf, however, it falls short on three fronts. It is not formal: with no structured, searchable statement of the problem, the search has little structure to exploit and hard constraints are not guaranteed. It is not transferable: each task is solved from scratch, so nothing learned on one

arXiv (cs.AI) · Aug 3, 2026
4 min
Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State InjectionResearch

Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second cost by construction; we eliminate the first, collapsing prefill from $O(L_{context})$ to $O(1)$ per query. We introduce PRECOG (Pre-Computed Context Injection), a retrieval mechanism that exploits a property unique to SSMs: the fixed-size, position-agnostic recurrent hidden state is a complete summary of everything the model has read. PRECOG pre-encodes do

arXiv (cs.AI) · Aug 3, 2026
4 min
Benchmarking Sheaf Neural Networks for Inductive TasksResearch

Benchmarking Sheaf Neural Networks for Inductive Tasks

Sheaf Neural Networks (SNNs) generalize message passing by replacing scalar edge weights of standard Graph Neural Networks (GNNs) with learnable, edge-dependent restriction maps between node stalks. Despite their strong theoretical foundations and promising transductive results, SNNs have been evaluated almost exclusively on transductive node classification, leaving their behaviour under inductive protocols unknown. We address this gap through the first systematic benchmark of the sheaf design space, evaluating three diffusion mechanisms (neural sheaf diffusion, sheaf attention, and

arXiv (cs.LG) · Aug 3, 2026
4 min
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AIResearch

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents a taxonomy-driven survey of the major cognitive capability gaps that continue to constrain the development of Cognitive AI. The literature is o

arXiv (cs.AI) · Aug 3, 2026
4 min
Who Should Be Generated? Justifying Demographic Targets in Open-Ended GenerationResearch

Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation

Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates "a CEO in the United States," the prompt leaves demographic realization to the model. Existing group fairness definitions assume that sensitive attributes are given on the input side. Generative audits instead examine output-side demographic composition, yet the targets they compare it against are typically supplied rather than justified. The upstream question is what the target distribution should be. We formalize this missing-target problem for

arXiv (cs.AI) · Aug 3, 2026
4 min
A Simple Approximation to the Distribution of the Ridge Regression EstimatorResearch

A Simple Approximation to the Distribution of the Ridge Regression Estimator

We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage.

arXiv (cs.LG) · Aug 3, 2026
4 min
Interaction Is Not Necessary for Order-Optimal 1-Bit Mean EstimationResearch

Interaction Is Not Necessary for Order-Optimal 1-Bit Mean Estimation

This paper is concerned with one-bit mean estimation, where each independent sample is represented by a single binary message. We consider distributions on $\mathbb{R}$ with mean in $[-λ,λ]$ and absolute $k$-th central moment at most $σ^k$, where $k>1$ is fixed. For this class, previous work attained the optimal sample complexity for general queries using a two-stage protocol. The first stage localizes the mean. The second-stage queries are chosen after localization and refine the estimate around the decoded center. We show that this interaction can be avoided by constructing a rando

arXiv (cs.LG) · Aug 3, 2026
4 min
Optimal Unambiguous DNFs and Alon-Saks-SeymourResearch

Optimal Unambiguous DNFs and Alon-Saks-Seymour

We construct unambiguous DNFs having width $O(n)$ but $0$-certificate complexity $Ω(n^2)$. By utilizing the special structure of these DNFs, we prove a lifting theorem with a constant-sized gadget that lifts the DNF to a communication problem, while losslessly translating the separation in certificate complexity to a separation in communication complexity. This leads to an optimal refutation of the Alon-Saks-Seymour conjecture, as well as an optimal communication lower bound for the Clique versus Independent Set problem, improving the previous results of Balodis, Ben-David, Göös, Jai

arXiv (cs.LG) · Aug 3, 2026
3 min
Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA ExpertsResearch

Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA Experts

Mixtures of low-rank adaptation experts increase parameter-efficient capacity by routing each input through a subset of adapters. Recent dynamic routers activate more experts when the router or prediction is uncertain. This rule silently equates uncertainty with useful additional computation: an uncertain example may contain complementary, unqueried expert evidence, but it may instead remain ambiguous after every expert agrees. We formulate routing as certified value-of-information allocation. VI-MoLE learns the counterfactual risk remaining after each expert prefix, converts these p

arXiv (cs.LG) · Aug 3, 2026
3 min
Analytic Planning under Uncertainty with Moment ClosureResearch

Analytic Planning under Uncertainty with Moment Closure

Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is p

arXiv (cs.AI) · Aug 3, 2026
3 min
Analytic Planning under Uncertainty with Moment ClosureResearch

Analytic Planning under Uncertainty with Moment Closure

Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state di...

Papers with Code · Aug 3, 2026
1 min
Magnet: Detecting Cross-Session AI Misuse Through Capability AccumulationResearch

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a critical gap: an attacker can decompose a harmful goal into innocuous-looking units and execute each in isolated agentic sessions. The agent is sta

arXiv (cs.AI) · Aug 3, 2026
4 min
LiveMem: Maintaining Memory State Continuity in Long-Running LLM InferenceResearch

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the full lifecycle when working context changes. We formulate this missing inference capability as \emph{state continuity under context turnover}: carrying computation forward through a fixed-capacity memory state whose lifetime is independent of the active context. We introduce an intrinsic memory method, \textbf{LiveMem}, which augments a pretrained

arXiv (cs.LG) · Aug 3, 2026
4 min
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of PoliciesResearch

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) r...

Papers with Code · Aug 3, 2026
2 min
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of PoliciesResearch

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards. Optimizing a single policy across all possible MDPs may sacrifice performance, while preparing an individually optimized policy for every MDP may violate operational, regulatory, or interpretability constraints on the number of policies that can be prepared and deployed. We consider se

arXiv (cs.AI) · Aug 3, 2026
4 min
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesResearch

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility sp

arXiv (cs.LG) · Aug 3, 2026
4 min
Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic RegressionResearch

Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression

We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-δ)$ quantile over all fixed collections of design vectors and all target parameters: \[ d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}δ\right). \] This is a nonasymptotic analogue of the Wilks $χ^2_d$ phenomenon and requires no regularity assumptions on the design. The low dimensional cases exhibit unusual behavior. The worst case qua

arXiv (cs.LG) · Aug 3, 2026
3 min
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis GenerationResearch

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis gene...

Papers with Code · Aug 3, 2026
2 min
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis GenerationResearch

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world. We defend a narrower claim: online embodiment is not necessary for every abductive scientific act. Our focus is identity abduction: the inference that two independently developed structures are one object under an explicit correspondence, reached through representational grounding rather than bodily interaction. An agent may acquire new inferential affordances n

arXiv (cs.AI) · Aug 3, 2026
4 min
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum OrthogonalizationResearch

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse functionally distinct weights (e.g., within AdaLN and QKV layers) into unified tensors for computational efficiency. Applying Muon to these fus

arXiv (cs.AI) · Aug 3, 2026
3 min
SWE-Touch: Benchmarking Coding Agents When Users Touch the CodeResearch

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to task-relevant code that conflict with task completion. SWE-Touch mines task-critical regio

arXiv (cs.AI) · Aug 3, 2026
4 min
DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label DisambiguationResearch

DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation

Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cues. Previous works typically attempt to construct discriminative representation of the small objects. However, the wide range frequency domain noises and label ambiguities have been greatly overlooked, which significantly hinders the accurate localization. To address these issues, we propose a novel small object detection (SOD) detector termed DyFrDet, which is able to precisely localize the small object by dynamically suppressing the

arXiv (cs.AI) · Aug 3, 2026
4 min
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI InteractionsResearch

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of h

arXiv (cs.AI) · Aug 3, 2026
3 min
Computational and Statistical Guarantees of the \textit{c}-Rectified flowResearch

Computational and Statistical Guarantees of the \textit{c}-Rectified flow

Recently, rectified flow has emerged as a fundamental framework for large-scale image generation, powering state-of-the-art systems such as FLUX.1 and Stable Diffusion 3. Despite its remarkable empirical success, the computational and statistical guarantees of iterative rectified flow have remained largely unexplored. We address this problem by studying \textit{c}-rectified flow, a cost-aware class of rectified flow that projects velocity fields onto a gradient class while preserving endpoint marginals. The ordinary rectified flow can fail to recover the optimal transport coupling: i

arXiv (cs.LG) · Aug 3, 2026
4 min
Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMsResearch

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-

arXiv (cs.LG) · Aug 3, 2026
3 min
Private Generative Bootstrap via BlockingResearch

Private Generative Bootstrap via Blocking

With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers. To this end, we adopt a Bayesian likelihood-free framework and make simulation from the posterior private. In particular, we propose a new private instantiation of the Bayesian bootstrap using a blocking strategy. Rather than assigning idiosyncratic random weights to each individual, we randomly group individuals and assign a single weight to each group. By concealing ind

arXiv (cs.LG) · Aug 3, 2026
4 min
CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMsResearch

CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMs

Trust is fundamental in modern regulatory ecosystems, and compliance checking plays a critical role in fostering that trust. Regulatory compliance verificati...

Papers with Code · Aug 3, 2026
2 min
Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgeryResearch

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical intent is highly specialized and tacit. Here, we introduce DiffeoAfford, an action-grounded tissue affordance framework that retrospectively derives visual attention supervision from completed surgical procedures. By combining diffeomorphism-constrained tissue tracking with instrument trajectory analysis, DiffeoAfford generates affordance hotspot labels without manual per-frame annotation. A real-time predicti

arXiv (cs.AI) · Aug 3, 2026
3 min
Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage AssessmentResearch

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the context of automated vehicle damage assessment, where fine-grained defects such as scratches and hairline cracks occupy few pixels, produce weak gradient signal, and are easily confused with reflections and surface texture. We show that a state-of-the-art VLM (Qwen-VL) achieves strong semantic classification accuracy (87.3%) on this task but is system

arXiv (cs.AI) · Aug 3, 2026
4 min
Real-Time Detection and Repair of LLM Agent FailuresResearch

Real-Time Detection and Repair of LLM Agent Failures

LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself. We ask how much detection is achievable from observable step telemetry alone, using monitors costing microseconds per step and trained only on healthy runs. On 2,823 committed agent episodes across three frameworks, three local models (qwen2.5 7b/3b, llama3.1 8b) and a commercial API (gemini-2.5-flash), a one-class echo-state-network ensemble with CUSUM

arXiv (cs.AI) · Aug 3, 2026
4 min
Syntax Meets Semantics: Understanding Scientific FormulaeResearch

Syntax Meets Semantics: Understanding Scientific Formulae

Scientific formulae are a fundamental component of scholarly communication, yet their dual nature -- as structured syntax and carriers of semantics -- remains underexplored in scholarly information retrieval. Although prior studies show that jointly modeling syntactic and semantic modalities improves retrieval performance, the relationship between their underlying representations has not been systematically investigated. In this work, we empirically study cross-modal correspondence between formula syntax and semantics. We find that their native representation spaces exhibit extremely

arXiv (cs.AI) · Aug 3, 2026
3 min
Aggregate-then-Calibrate for Human-centered Assessment with Theoretical GuaranteesResearch

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranki

arXiv (cs.LG) · Aug 3, 2026
3 min
Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+Research

Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+

Linear Temporal Logic (LTL) is one of the most widely adopted languages for specifying temporal extended objectives in AI, with applications ranging from reactive synthesis to stochastic planning in Markov decision processes and reinforcement learning. Traditionally, solving any of these problems requires translating the LTL specification to a nondeterministic automata on infinite words and then determinizing it, a step that is notoriously difficult in theory and in practice. Recent work has introduced LTLf+, which lifts the finite-trace logic LTLf to infinite traces. LTLf+ has the s

arXiv (cs.AI) · Aug 3, 2026
4 min
Advancing Relevance Measurement with Vision-Language Models for Web-Scale SearchResearch

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs ca

arXiv (cs.LG) · Aug 3, 2026
3 min
ParEvalLayer: When Partial LLM-Agent Evaluations Support a DecisionResearch

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

LLM-agent evaluations often produce task outcomes long before the full benchmark run is complete. A partial score is tempting to report, but it does not show whether the observed tasks support the same conclusion as the completed evaluation. Early tasks can omit important parts of a benchmark, running cheaper tasks first can distort the observed sample, and a rule that decides only easy pairs can appear accurate while leaving many comparisons unresolved. We introduce ParEvalLayer, a decision layer that reads paired outcomes for two agent systems and a comparison policy chosen in adva

arXiv (cs.AI) · Aug 3, 2026
4 min
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science BenchmarksResearch

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking

arXiv (cs.AI) · Aug 3, 2026
4 min
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe CommerceResearch

Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld

arXiv (cs.AI) · Aug 3, 2026
3 min
Intention Inference Under Execution Noise: Separating Aleatoric and Epistemic Uncertainty in Social DilemmasResearch

Intention Inference Under Execution Noise: Separating Aleatoric and Epistemic Uncertainty in Social Dilemmas

In noisy social dilemmas, intended actions are stochastically corrupted before execution, so an observed defection may reflect hostile intent or action error. Standard Markov Decision Process (MDP) formulations treat executed actions as states, structurally precluding this distinction and causing systematic over-retaliation. We introduce a Partially Observable MDP (POMDP) formulation encoding opponent intentions as latent states and executed actions as noisy observations, solved within the active inference (AIF) framework with a cost function that decomposes into epistemic and pragma

arXiv (cs.LG) · Aug 3, 2026
3 min
xPress: Parallel Refinement for Diffusion Drafters in Speculative DecodingResearch

xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising process involves using the logit distribution at each position to sample conditionally independent tokens. The resulting draft is thus a set of per-position marginals, rather than a joint distribution: no draft token is guaranteed to depend on its predecessors. Such independently sampled marginals tend to produce sequences with tokens th

arXiv (cs.AI) · Aug 3, 2026
4 min
Foundations of Reinforcement Learning and Control:Connections and New PerspectivesResearch

Foundations of Reinforcement Learning and Control:Connections and New Perspectives

Reinforcement learning and control theory are two adjacent scientific fields that focus on optimizing the controller of unknown dynamical systems using feedback. While both fields have common roots in dynamic programming, they have evolved with distinct methodologies, goals, and cultures. Despite decades of mutual influence, a significant gap persists between the two communities. This tutorial introduces adaptive control, actor-critic reinforcement algorithms, and a new way to combine these two paradigms for data-driven decision making on a classical locomotion control problem. Our a

arXiv (cs.LG) · Aug 3, 2026
3 min
Agentic Incident Response through Digital Twin-Enhanced Multiscale PlanningResearch

Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning

Incident response is currently managed by security operators using predefined playbooks, resulting in slow, labor-intensive security decision-making processes. Consequently, there is a growing need for automated incident response planning. Decision-theoretic approaches based on control, optimization, and reinforcement learning have been proposed to automate such planning tasks with well-grounded approaches, yet most of which, while guaranteeing strong performance, are limited to abstract models and cannot be directly applied to operational systems. A promising approach to mitigate th

arXiv (cs.AI) · Aug 3, 2026
4 min
SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event ReasoningResearch

SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning

ANN-based All-in-One image restoration (AiOIR) unifies diverse degradation handling but incurs high computational costs, limiting its real-time deployment. W...

Papers with Code · Aug 3, 2026
1 min
Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral SignaturesResearch

Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures

Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a significan...

Papers with Code · Aug 3, 2026
1 min
Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task LossResearch

Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss

Agents that act on a compressed representation of their history face a structural risk: if the representation aliases histories with different optimal action...

Papers with Code · Aug 3, 2026
1 min
An Accessible Solution for Deformable Image Registration Compared with Learning-Based ApproachesResearch

An Accessible Solution for Deformable Image Registration Compared with Learning-Based Approaches

Deformable image registration (DIR) is a core problem in medical image analysis; but, unlike labeling decision problems such as classification and segmentati...

Papers with Code · Aug 3, 2026
1 min
Empowering Credit Risk Detection in Weixin Pay with Billion-Scale Deep Graph LearningResearch

Empowering Credit Risk Detection in Weixin Pay with Billion-Scale Deep Graph Learning

Credit risk detection, particularly mitigating individual fraud, is crucial for maintaining the stability of digital financial ecosystems. Accurately identif...

Papers with Code · Aug 3, 2026
2 min
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMsResearch

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physi...

Papers with Code · Aug 3, 2026
2 min
CoRe-GNN: Multilevel Message passing on Coarsened graphsResearch

CoRe-GNN: Multilevel Message passing on Coarsened graphs

Training Graph Neural Networks on large graphs is challenged by the memory cost of storing all node representations across layers. We show that several exist...

Papers with Code · Aug 3, 2026
2 min
Cross-Domain Hybrid OPD for Generalizable Search AgentsResearch

Cross-Domain Hybrid OPD for Generalizable Search Agents

Recent advances in Reinforcement Learning (RL) have substantially improved the capabilities of autonomous search agents, enabling sophisticated planning, and...

Papers with Code · Aug 3, 2026
1 min
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-DistillationResearch

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation

Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action...

Papers with Code · Aug 3, 2026
1 min
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional CompanionshipResearch

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted...

Papers with Code · Aug 3, 2026
2 min
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and TeachersResearch

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inap...

Papers with Code · Aug 3, 2026
1 min
Finite-Time Analysis of Discounted Exponential-Utility Reinforcement LearningResearch

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinfor...

Papers with Code · Aug 3, 2026
2 min
CRISP: Critical Step Perception for Training Efficient Deep Search AgentsResearch

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external sear...

Papers with Code · Aug 3, 2026
2 min
Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Structure Reconstruction of the Ground-State ITG BranchResearch

Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Structure Reconstruction of the Ground-State ITG Branch

Physics-informed neural networks (PINNs) combine sparse observations with physical equations, providing an important approach for modeling complex plasma pro...

Papers with Code · Aug 3, 2026
1 min
CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM AgentsResearch

CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents

As memories accumulate across tasks and sessions, the performance of long-term LLM agents depends jointly on query-specific retrieval and continual memory re...

Papers with Code · Aug 3, 2026
2 min
MNC: Scope-Bound Semantic Declassification for Private LLM-Agent CommunicationResearch

MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication

Multi-agent large language model (LLM) systems can expose protected state through internal messages, tool arguments, logs, and persistent memory even when th...

Papers with Code · Aug 3, 2026
2 min
STC-Net: Electroluminescence-Based Solar Cell Crack Segmentation for Power Loss EstimationResearch

STC-Net: Electroluminescence-Based Solar Cell Crack Segmentation for Power Loss Estimation

Accurate crack assessment in electroluminescence (EL) images is important for photovoltaic (PV) reliability analysis, yet existing segmentation methods often...

Papers with Code · Aug 3, 2026
1 min
Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?Research

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for ...

Papers with Code · Aug 3, 2026
2 min
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea GenerationResearch

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open questi...

Papers with Code · Aug 3, 2026
2 min
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language ModelsResearch

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill ...

Papers with Code · Aug 3, 2026
2 min
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor LearningResearch

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tu...

Papers with Code · Aug 3, 2026
1 min
D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian SplattingResearch

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rende...

Papers with Code · Aug 3, 2026
2 min
Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standardResearch

Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard

Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth f...

Papers with Code · Aug 3, 2026
2 min
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation MetricsResearch

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with...

Papers with Code · Aug 2, 2026
2 min
Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+Research

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languag...

Papers with Code · Aug 2, 2026
2 min
Perspectives on Tsallis Statistics for Artificial IntelligenceResearch

Perspectives on Tsallis Statistics for Artificial Intelligence

Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and freque...

Papers with Code · Aug 2, 2026
2 min
ExtractBench: A Benchmark for Schema-Guided Enterprise Document ExtractionResearch

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The evaluation system contains 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types, with clear tags differentiating th

arXiv (cs.AI) · Jul 31, 2026
3 min
Differentially Private Nonparametric Modal Learning with Applications to Regression and ClusteringResearch

Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering

Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. We study differentially private recovery of density modes for multivariate distributions under local smoothness, curvature, and separation conditions. We propose DP-GRAMS, a mean-shift inspired method that performs noisy ascent on a differentially private score estimator. Assuming the density belongs locally to a Hölder class with smoothness parameter $β> 2$, our score estimator uses bias-reducing hig

arXiv (cs.LG) · Jul 31, 2026
4 min
Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error FeedbackResearch

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. It outperforms SignSGD in practice, yet it can ascend even on a linear function. Signing the gradient before the Linear Minimization Oracle (LMO), rather than after, does not repair this: we construct a small explicit instance on which sign-before (MuonUSign) and sign-on-both-sides (MuonSign) ascend as well, so no placement of the sign around the oracle descends in general. Error feedbac

arXiv (cs.LG) · Jul 31, 2026
4 min
Freeze, Then Select: Structured Field Adapters and Stability-Validated Weak Selection for PDE Discovery from Sparse ObservationsResearch

Freeze, Then Select: Structured Field Adapters and Stability-Validated Weak Selection for PDE Discovery from Sparse Observations

PDE discovery from sparse observations requires reconstructing a continuous field and selecting the correct differential terms. Our analysis of optimization paths in coupled neural PDE discovery reveals three behaviors: the exact support can persist to the end of training, appear only transiently, or fail to emerge. To decouple equation selection from neural optimization, we develop a freeze-then-select method combining a structured field adapter with Stability-Validated Weak Selection (SVWS). Trained from observations without a PDE residual, the adapter factorizes the field into lea

arXiv (cs.LG) · Jul 31, 2026
4 min
GQ-FSL: Green Quantized Federated Split LearningResearch

GQ-FSL: Green Quantized Federated Split Learning

Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device computation by offloading workloads to an edge server, this may introduce systemic overheads, while the continuous exchange of cut-layer data, and submodels still incurs significant energy consumption (EC). To address this, we propose a green quantized FSL (GQ-FSL) framework that incorporates stochastic quantization for both local collaborative training and wireles

arXiv (cs.LG) · Jul 31, 2026
4 min
Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and DiagnosticsResearch

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that can bridge heterogeneous data sources, diverse equipment types, and varied diagnostic outputs. Limited data interpretability and interoperability within the FDD domain have led to fragmented information silos, hindering the implementation of FDD and related applications, such as the digital twin-enabled FDD frameworks and artificial inte

arXiv (cs.AI) · Jul 31, 2026
4 min
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter OptimizersResearch

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a vali

arXiv (cs.AI) · Jul 31, 2026
3 min
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational ExaminationsResearch

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time pr

arXiv (cs.AI) · Jul 31, 2026
3 min
CENDRe: Concept Extraction with Natural Domain RepresentationsResearch

CENDRe: Concept Extraction with Natural Domain Representations

Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. Concept extraction (CE) methods identify such patterns by analyzing representations within the models' latent space. However, existing time-series CE methods have three limitations: they operate only in the time domain and overlook frequency features, predefine the number of concepts, and produce localizations misaligned with the regions the model uses. We address these lim

arXiv (cs.AI) · Jul 31, 2026
4 min
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation LearningResearch

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical, e.g., in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner's own trajectories, and using value function estimation en route

arXiv (cs.AI) · Jul 31, 2026
4 min
A Human-Centered Validation of the Explainability-Performance CoefficientResearch

A Human-Centered Validation of the Explainability-Performance Coefficient

The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligence (XAI). However, objectively evaluating explanation fidelity and aligning XAI metrics with human-centered understanding remain critical open challenges. In this work, we propose a model-agnostic metric, the EPC score, which is an extension of the Explainability-Performance Coefficient (EPC), that quantifies explanation quality by explicitly balancing the trade-off between feature selection sparsity and preserved model performance. Through an empi

arXiv (cs.AI) · Jul 31, 2026
3 min
QASP: Query-Adaptive Robust Vector Search PolicyResearch

QASP: Query-Adaptive Robust Vector Search Policy

A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. We introduce QASP (Query-Adaptive robust vector Search Policy), which predicts the complete recall progression curve per query via a single upfront supervised regression, from which a search policy is derived for any recall target; this avoids iterative model invocations during search or separate predictors per

arXiv (cs.LG) · Jul 31, 2026
4 min
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language ModelsResearch

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach

arXiv (cs.AI) · Jul 31, 2026
3 min
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMsResearch

The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading to poor transfer and catastrophic forgetting. Existing approaches mainly improve adapter expressiveness by increasing parameter capacity or composing multiple adapters, yet they still rely on a shared optimization path. In this paper, we propose an optimization-path organization framework for parameter-efficient fine-tuning of large language models,

arXiv (cs.LG) · Jul 31, 2026
4 min
Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion EnvironmentResearch

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the constant learning rate is below a time-invariant threshold; and the regret bound has order

arXiv (cs.LG) · Jul 31, 2026
3 min
TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual LearnersResearch

TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly important though less well-studied aspect of CL systems is their ability to distinguish inputs that are unlikely to come from within the set of tasks the system has already encountered, often called out-of-distribution (OOD) detection. This paper presents several findings related to the dynamics of OOD detection in CL systems, causes of performance degradation over time which we call OOD forgetting (OODF), and proposed mitigation strategie

arXiv (cs.LG) · Jul 31, 2026
4 min
TraceViT: Grounded Trace Supervision for Visual Abstract ReasoningResearch

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional training constrains only the final output, leaving intermediate refinements unconstrained. We propose that these refinements should instead follow the transformation step by step. We introduce TraceViT, a looped visual reasoner trained with semantically monotonic transformation chains. We obtain these chains by rewriting and verifying progra

arXiv (cs.AI) · Jul 31, 2026
3 min
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons CombatResearch

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At ea

arXiv (cs.AI) · Jul 31, 2026
4 min
MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language ModelsResearch

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. While recent Large Language Model (LLM) based approaches show promise, they face two limitations. First, they lack data analysis mechanisms for uncovering variable dependencies, which reduces the efficiency of equation discovery. Second, most methods rely on single-objective evaluation focused solely on fitting error. This neglect of structural complexity and generalization often causes models to converge prematurely to local optima, limiting their a

arXiv (cs.AI) · Jul 31, 2026
4 min
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference FeedbackResearch

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference

arXiv (cs.AI) · Jul 31, 2026
4 min
Pyramidal Width Can Increase Under Vertex InsertionResearch

Pyramidal Width Can Increase Under Vertex Insertion

Lacoste-Julien and Jaggi conjectured in 2015 that the pyramidal width of a polytope cannot increase when a vertex is added, provided that every old point remains a vertex. We give an exact counterexample with six integer points in $\R^3$. For \[ P=\conv\{v_0,\ldots,v_4\},\qquad Q=\conv\{v_0,\ldots,v_5\}, \] where \[ \begin{aligned} v_0&=(-1,-3,-1), & v_1&=(3,2,-2), & v_2&=(0,2,1),\\ v_3&=(-1,-3,3), & v_4&=(-2,0,1), & v_5&=(-1,0,-2), \end{aligned} \] all five vertices of $P$ remain vertices of $Q$, but \[ \PWidth(P)^2=\frac{48}{353} \quad\text{and}\quad \PWidth(Q)^2=\frac{36}{133}. \]

arXiv (cs.LG) · Jul 31, 2026
3 min
COntExt: Towards Context-Aware Ontology Extension from Operational MetricsResearch

COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and relationships, that often extends what is captured in formal ontologies. Yet the connection between operational metric catalogues and ontological knowledge remains manual, ad-hoc, and labor-intensive. We present COntExt, a framework for context-aware ontology extension that takes structured metric definitions as input and suggests how refer

arXiv (cs.AI) · Jul 31, 2026
3 min
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-CorrectionResearch

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language reflection or assist verification by directly generating verification programs; the former may not reliably support exact computation, whereas the latter prematurely couples mathematical modeling with low-level implementation. We propose AMTFV (Agentic Mathematical Tool-Flow Verification). By introducing Mathematical Tool Flow (MTF) as an interrupt--e

arXiv (cs.AI) · Jul 31, 2026
4 min
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector EvaluationResearch

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B,

arXiv (cs.AI) · Jul 31, 2026
4 min
A Neurosymbolic Approach for Explainable Early Diagnosis of Alzheimer's DiseaseResearch

A Neurosymbolic Approach for Explainable Early Diagnosis of Alzheimer's Disease

Identifying reliable Alzheimer's disease (AD) markers typically requires manual, labor-intensive transcription and expert analysis, limiting its scale. We introduce an automated pipeline that extracts qualitative knowledge about potential AD progression indicators directly from audio recordings of verbal fluency tests. Our method uses pretrained foundation models to process raw audio and extract clinically relevant variables to construct a Bayesian Network (BN); this BN is used to reason about the AD progression markers and infer their qualitative relationships. Our system successful

arXiv (cs.LG) · Jul 31, 2026
3 min
TerraNova: A Foundation Model for the AnthropoceneResearch

TerraNova: A Foundation Model for the Anthropocene

A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geom

arXiv (cs.AI) · Jul 31, 2026
4 min
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at ScaleResearch

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance. We present ARCTIC, an AI-powered Code Critique system that reframes code review around three capabilities: intent prediction, which infers why a change was made from conversation logs and metadata; drift detection, which measures divergence between the developer

arXiv (cs.AI) · Jul 31, 2026
4 min
Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxidesResearch

Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxides

High-entropy perovskite oxides (HEPOs) represent a chemically complex class of materials with promising functional properties, yet their vast compositional space and, chemical/structural disorder pose significant challenge for accurate property prediction. Graph neural networks (GNNs) enable rapid exploration of materials space but are often limited by the availability of representative training data. Here, we investigate ordered-to-disordered transfer learning using GNNs for formation-energy and HOMO-LUMO gap prediction in HEPOs by transferring knowledge learned from chemically orde

arXiv (cs.LG) · Jul 31, 2026
4 min
Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic SegmentationResearch

Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance, which results from relatively lower proportions of small and limitedly exposed structures. Recent works on laparoscopic multi-organ segmentation focus on learning structure-specific features through class-specific decoder architectures and report favorable results. This work extends the decoder-focused architectures to investigate knowledge sharing in the cross-surgical domain. We utilize two datasets representing different surgical

arXiv (cs.LG) · Jul 31, 2026
4 min
The Grokked Illusion: True Equilibrium Mitigates Catastrophic ForgettingResearch

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown that solutions occupying larger volumes in parameter space, as quantified by Boltzmann entropy, often exhibit superior generalizability compared to those reached by conventional optimization, a phenomenon known as the high entropy advantage. Here we ask whether this advantage persists beyond generalization. Specifically, we investigate models' robustness, the ability to retain the learned knowledge when the

arXiv (cs.LG) · Jul 31, 2026
4 min
Transcript-Managed Transformers: Monotone Multi-Agent Collapse and Universality with Two Pop-Enabled TranscriptsResearch

Transcript-Managed Transformers: Monotone Multi-Agent Collapse and Universality with Two Pop-Enabled Transcripts

We study transcript management for fixed, finite-precision causal Transformers. A transcript is partitioned into channels of bounded blocks. Each transition consults a fixed visible suffix and may append one block, leaving the model, weights, and token protocol unchanged. The operation $P_c:=\PopContext(c)$ deletes the newest block on channel $c$ and exposes its predecessor. We model the layer by the Transcript-Managed Transducer $\TMTn{k}$: one finite controller, $k$ channels, and per-round actions from stay, push, and pop under a caller-driven status map. Fixed visible windows enco

arXiv (cs.LG) · Jul 31, 2026
4 min
Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy DistillationResearch

Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation

On-policy distillation (OPD) provides dense teacher supervision along student-generated trajectories, but its online rollout process incurs substantial computational cost, particularly when a few long responses delay batch completion. Existing acceleration methods typically control rollout length using fixed budgets or absolute teacher--student agreement thresholds, which may not reflect learning progress across different models and training stages. We propose Adaptive FastOPD, a progress-aware strategy that expands the rollout horizon only when learning near the current boundary reg

arXiv (cs.LG) · Jul 31, 2026
3 min
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture SearchResearch

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. We introduce DreamQAS, a model-based RL framework that preserves these exact circuit dynamics and learns only the expensive post-VQE feedback. A recurrent randomized-prior ensemble predicts an oracle-free score relative to an empirical energy frontier and supports multi-step imagined policy learning over explicit legal circuits. Ranking-based activation,

arXiv (cs.AI) · Jul 31, 2026
4 min
Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?Research

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fails instructively. In Simpson's-paradox worlds, where the two have systematically opposite signs, increasing the fraction of interventional samples in pretraining does not improve causal direction: the magnitude of the model's do()-response grows monotonically, yet its sign is copied from the observational context. What governs whether interventional

arXiv (cs.LG) · Jul 31, 2026
4 min
MolGVR: A Chemistry-Grounded Framework for Text-to-Molecule GenerationResearch

MolGVR: A Chemistry-Grounded Framework for Text-to-Molecule Generation

Text-to-molecule generation is typically formulated as a one-shot sequence generation problem, where a model directly maps target descriptions to molecular representations. However, molecular descriptions often contain informative structural constraints, and violating such constraints can change the molecular identity. This makes chemical verification and error correction important but underexplored. To fill this gap, we propose MolGVR, a chemistry-grounded Generator--Verifier--Refiner framework. The Generator infers structural evidence and generates candidate molecules. The Verifier

arXiv (cs.LG) · Jul 31, 2026
3 min
Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder ModuleResearch

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects of the problem. On one hand, affordance segmentation requires high-level abstraction capabilities, that typically involve large-size models. On the other hand, computing resources hosted on wearable robots prevent to run large-size models in real-time. The paper presents an analysis of the role of the segmentation head in the trade-off between generalization performance and compute cost. The obtained models outperform modern baseline s

arXiv (cs.LG) · Jul 31, 2026
3 min
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and RememberResearch

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural experience but are typically learned from fixed task distributions. We introduce \textbf{SESA} (Self-Evolving Skill-Augmented Agent), which makes procedural memory an evolving state of tool-augmented search self-play. A challenger poses problems, while a separately parameterized solver alone retrieves skills. Informative failures are distil

arXiv (cs.AI) · Jul 31, 2026
4 min
TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate FusionResearch

TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion

Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predictive maintenance. While recent time series foundation models exhibit strong generalization, they rely on static parametric knowledge and lack dynamic access to external historical patterns during inference. Retrieval-Augmented Generation (RAG) offers a potential remedy, yet its application to time series forecasting is challenged by magnitude variations across heterogeneous sources and the mismatch between historical similarity and future consist

arXiv (cs.AI) · Jul 31, 2026
3 min
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language ModelsResearch

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot therma

arXiv (cs.AI) · Jul 31, 2026
4 min
Beyond Retrieval: Analytic Memory for Multimodal AgentsResearch

Beyond Retrieval: Analytic Memory for Multimodal Agents

Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a complementary abstraction that organizes recurring multimodal observations into queryable structures supporting filtering, aggregation, ranking

arXiv (cs.AI) · Jul 31, 2026
3 min
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization ModelsResearch

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-spac

arXiv (cs.AI) · Jul 31, 2026
4 min
AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability RepairResearch

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair

Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair. However, vulnerability repair demands richer program context than general bug repair - context that security engineers routinely assemble in practice but that existing agentic approaches do not engineer. We identify three critical gaps: code-structure context capturing cross-file data flows and memory operation patterns, runtime-execution context revealing crash sem

arXiv (cs.AI) · Jul 31, 2026
4 min
Explore Beyond the Boundary Using Entropic InformationResearch

Explore Beyond the Boundary Using Entropic Information

In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process. Addressing this issue requires extensive exploration in the state space to discover valuable reward signals. In this paper, we propose Entropic Information for Exploration (ENTINEX), a novel method that enhances exploration by incentivizing agents to explore beyond the boundaries of the state distribution. ENTINEX achieves this by assigning intrinsic rewards to these boundaries, leveraging entropic information

arXiv (cs.AI) · Jul 31, 2026
3 min
Learning to Trace Seiberg DualitiesResearch

Learning to Trace Seiberg Dualities

Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when all of the "rules of the game" are well-known. Said differently, when confronted with two systems, how can one efficiently establish that they are in fact dual? In this paper we use machine learning methods to address this question for Seiberg dualities of supersymmetric quiver gauge theories. Mathematically, this involves establishing mutations o

arXiv (cs.AI) · Jul 30, 2026
4 min
ReToken: One Token to Improve Vision-Language Models for Visual RetrievalResearch

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4

arXiv (cs.AI) · Jul 30, 2026
3 min
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid DodgeballResearch

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and

arXiv (cs.AI) · Jul 30, 2026
3 min
AskChem: Claim-Centered Infrastructure for Chemistry Literature SynthesisResearch

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemble cross-paper answers manually. We present AskChem, a claim-centered infrastructure for cross-paper chemistry search. AskChem changes the unit of retrieval from the paper to the provenance-carrying claim: each paper is converted into atomic, typed claims, each grounded by a source DOI and

arXiv (cs.AI) · Jul 30, 2026
3 min
AISPA: User-Centric System Prompt Auditing for Large Language Model ApplicationsResearch

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users

arXiv (cs.AI) · Jul 30, 2026
4 min
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsResearch

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduc

arXiv (cs.AI) · Jul 30, 2026
4 min
KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk ModelsResearch

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determi

arXiv (cs.LG) · Jul 30, 2026
4 min
Change2Task: From Repository Changes to Executable Coding Agent Tasks and EnvironmentsResearch

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the

arXiv (cs.LG) · Jul 30, 2026
4 min
MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision TransformersResearch

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers

Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However, existing PTQ methods typically employ uniform bit-widths across transformer components, overlooking their heterogeneous sensitivity to quantization and leading to inefficient precision allocation. In this paper, we propose {MixFrag, a fragility-guided mixed-precision PTQ framework for Vision Transformers. MixFrag first estimates component-level quantization fragility by measuring the Kullback--Leibler (KL) divergence between full-pre

arXiv (cs.LG) · Jul 30, 2026
4 min
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like BenchmarksResearch

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the PR description; the issue description is used as the problem statement, and the PR patch serves as the test oracle. However, due to the inherent complexity of developing and maintaining large repositories, such PR-Issue pairings are often misaligned in practice. In this work, we systematically study SWE-bench Verified instances, finding that 13.6%

arXiv (cs.AI) · Jul 30, 2026
3 min
$β$-OPSD: Deriving with Policy Optimization, Training with Self-DistillationResearch

$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the $β=1$ member of a broader policy-optimization family, where $β$ weights the KL penalty anchoring the student to a reference policy. This equivalence turns $β$ from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity t

arXiv (cs.LG) · Jul 30, 2026
4 min
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented GenerationResearch

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation

While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex multi-hop reasoning tasks. Existing methods primarily focus on independent instance-level matching, which often fails to capture explicit relationships across modalities and documents. Although Graph-enhanced methods introduce structural modeling, they face a fundamental challenge in multimodal scenarios: incorporating fine-grained visual features leads to rapid graph expansion and retrieval noise, whereas coarse-grained representations cause the discarding of critica

arXiv (cs.AI) · Jul 30, 2026
4 min
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7BResearch

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought. Because generating more text raises accuracy by itself, a gain over one chain of thought does not show the method's idea is what helped. Wang et al. (2024) reported that a simple baseline, sampling the same question repeatedly and keeping the most common answer, often wins once budgets are comparable, but gave point estimates with no confidence in

arXiv (cs.AI) · Jul 30, 2026
4 min
Algorithms for Structured Elections under Thiele Voting RulesResearch

Algorithms for Structured Elections under Thiele Voting Rules

We study the computational complexity of winner determination problems in approval-based committee elections under Thiele voting rules. These form a class of rules parameterized by a fixed weight vector that specifies how a voter's satisfaction depends on the number of approved candidates elected. We first analyze the structure of optimal solutions based on the sets of voters who approve each candidate---that is, how voters' approval ballots induce dependencies between candidates---revealing constraints on a winning committee under any fixed Thiele voting rule. Using this, we design

arXiv (cs.AI) · Jul 30, 2026
4 min
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute TradeoffsResearch

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs

Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constraints remains challenging. While recent studies show that inference-time scaling can improve frontier computer-use agents through additional computation during execution, its effectiveness for resource-constrained local models remains poorly understood. We present a systematic empirical study of inference-time scaling in local CUAs across contextual, temporal, structural, and parallel dimen

arXiv (cs.AI) · Jul 30, 2026
4 min
Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular HistoriesResearch

Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional an

arXiv (cs.LG) · Jul 30, 2026
4 min
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic SystemsResearch

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that

arXiv (cs.AI) · Jul 30, 2026
4 min
ORCA-bench: How Ready Are Language Model Agents for Oncall?Research

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and

arXiv (cs.AI) · Jul 30, 2026
4 min
ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature ProgramsResearch

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable featur

arXiv (cs.LG) · Jul 30, 2026
4 min
Graph Neural Network Force Fields for Spin Dynamics in Metallic MagnetsResearch

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically require repeated solutions of an underlying electronic problem throughout the time evolution, creating a major computational bottleneck. Here we introduce a graph neural network (GNN) magnetic force-field framework that learns the effective magnetic energy functional governing itinerant spin dynamics directly from electronic calculations. Conceptually analogous to machine-learned interatomic potentials, the proposed framework enables effi

arXiv (cs.LG) · Jul 30, 2026
3 min
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent SystemsResearch

MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems

Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent specialization, information exchange, and intermediate validation. However, existing systems typically treat communication topology as a fixed design choice or an offline optimization target. We introduce MANTA, a framework for Multi-Agent Network Topology Adaptation that enables communication structures to self-evolve at inference time. Before execution, MANTA initializes a task-conditioned topology from prior structural experience. During deployment, it monitors collabor

arXiv (cs.AI) · Jul 30, 2026
3 min
What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image RestorationResearch

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonly encode heterogeneous degradation conditions in a shared latent space, where degradation-related cues and scene content can remain entangled. We characterize the resulting challenge as dual ambiguity: semantic ambiguity in channel-wise modulation and spatial ambiguity in restoration responses, which can lead to content corruption and residual artifacts. To mitigate this issue, we propose DAR-Net, a Dual-Ambiguity Rectification Network for all-in-one image restoration.

arXiv (cs.AI) · Jul 30, 2026
4 min
Same Graph Cross-Task Transfer in GNNs: Protocols and PredictorsResearch

Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors

Many real-world graphs support multiple predictive tasks over the same underlying structure, creating an opportunity to reuse supervision across node classification (NC) and link prediction (LP). However, existing evaluations often rely on incompatible splits, observed-graph assumptions, and negative sampling rules, making conclusions about same-graph cross-task transfer unreliable. We formalize same-graph NC-LP transfer and propose a leakage-free protocol that fixes node and edge splits, uses a shared message-passing graph that excludes evaluated edges, and employs fixed negatives f

arXiv (cs.LG) · Jul 30, 2026
4 min
Selective Credibility-Limited Belief UpdateResearch

Selective Credibility-Limited Belief Update

Belief update concerns changes in an agent's beliefs induced by changes in the underlying world. Standard Katsuno-Mendelzon update assumes that an epistemic input can be incorporated from every initially possible world, whereas credibility-limited belief update restricts, for each source world, the successor worlds regarded as credible or reachable. Nevertheless, existing credibility-limited approaches treat the epistemic input as an indivisible whole, and therefore cannot represent cases in which only part of a compound epistemic input can be realized. We introduce selective credibi

arXiv (cs.AI) · Jul 30, 2026
4 min
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent ExploitationResearch

Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation

An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emph{budget-constrained confidence-scheduled restricted responses} (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actu

arXiv (cs.AI) · Jul 30, 2026
4 min
InfoOps Bench: A live information operations safety benchmarkResearch

InfoOps Bench: A live information operations safety benchmark

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live monitoring pipeline which tracks Russian, Chinese and Iranian state-backed information assets. Alongside this paper, we release a companion website that tracks the most prominent claims spread by state-backed media outlets, updated weekly, available from: pattrn.ai/research/infoopsbench. The dynamic nature of the benchmark makes it resistant

arXiv (cs.AI) · Jul 30, 2026
4 min
TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration RetrievalResearch

TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval

Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Existing SIR methods rank papers by topical similarity and do not explicitly represent how a candidate inspiration transfers to a target problem. This is especially limiting for remote inspirations, whose value often lies in reusable problem-solving principles rather than topical overlap. Motivated by how humans abstract transferable aspects of a source and remap them to a new target, we reformulate SIR as target-conditioned abstraction (TC

arXiv (cs.AI) · Jul 30, 2026
3 min
The Role of Causality in Algorithmic RecourseResearch

The Role of Causality in Algorithmic Recourse

Algorithmic recourse aims to provide individuals with actionable changes to improve their predicted outcomes in high-stakes classification settings, such as loan and mortgage applications. However, most existing approaches focus only on flipping a model's prediction, without accounting for whether the recommended changes lead to genuine improvement in an individual's true qualifications or merely enable strategic gaming of the classifier. Consequently, deployed recourse policies can induce behavioral responses that degrade predictive accuracy and become ineffective after model retrai

arXiv (cs.LG) · Jul 30, 2026
4 min
Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache TransplantationResearch

Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation

Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage boundary in a Qwen2.5-derived system. A matched 200-item experiment compares retained live cache with one-shot prefill of identical integer tokens and places an exact replica on both sides. In BF16, replicas remain exact while the constructions differ on 166 suffixes and 20 correctness labels; the accuracy difference is only one point (paired 95% CI [-3.5, +5

arXiv (cs.LG) · Jul 30, 2026
4 min
SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End CoordinationResearch

SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination

Can supply-chain AI move beyond isolated decision modules toward unified operational planning? A complete replenishment plan specifies which products each location carries, which upstream facility supplies it, how often it is replenished, and how deliveries are routed. These decisions are operationally coupled: the selected assortment changes the demand and load passed to later stages; source assignment and replenishment frequency reshape the delivery requests; and route feasibility and cost, in turn, determine the system value of the earlier choices. Yet in modern supply chains, the

arXiv (cs.AI) · Jul 30, 2026
4 min
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer NetworksResearch

A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks

Standard automated sewer pipe severity assessment relies on direct image classification, creating a "black box" where the link between visual defects and final severity scores remains implicit. This study introduces a modular, fuzzy rule-based neuro-symbolic framework that bridges this gap by decoupling neural perception from symbolic reasoning. The perception module utilizes a Swin Transformer to predict 14 multilabel inspection CODE degrees directly from images. For reasoning, a DT, specifically Weka's J48, algorithm is trained on ground-truth CODEs and severity labels, and its pat

arXiv (cs.AI) · Jul 30, 2026
4 min
Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data AugmentationResearch

Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation

Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conventional approach still transmits terabytes of raw imagery to the ground for processing, and open satellite datasets for aircraft are scarce and severely class-imbalanced. These limitations either delay timely decision-making or prevent standard detectors from learning robust representations of rare aircraft classes. In this paper, a workflow that combines on-board inference with generative data augmentation is proposed to address both l

arXiv (cs.AI) · Jul 30, 2026
4 min
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reportsResearch

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and s

arXiv (cs.AI) · Jul 30, 2026
3 min
Cybersecurity Detection Classification with Reasoning-enabled Language ModelsResearch

Cybersecurity Detection Classification with Reasoning-enabled Language Models

A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards. We find that CoT reasoning al

arXiv (cs.LG) · Jul 30, 2026
4 min
LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in LeanResearch

LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean

Constraint programming is a core technology for solving complex combinatorial problems in scheduling, planning, configuration, and verification. Trusting its results therefore demands guarantees at two levels: that reformulations applied beforehand are semantics-preserving, and that solvers produce correct answers. In this work, we introduce a framework that addresses both verification levels in the Lean theorem prover: it can be used to prove formulation-level properties, such as equivalence, equisatisfiability, and the correctness of symmetry-breaking constraints, parametrically fo

arXiv (cs.AI) · Jul 30, 2026
3 min
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time ComputeResearch

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues

arXiv (cs.AI) · Jul 30, 2026
4 min
Graph Neural Multilevel Preconditioners for Iterative SolversResearch

Graph Neural Multilevel Preconditioners for Iterative Solvers

Solving large, sparse linear systems is a core task in scientific computing, and efficient iterative solvers rely critically on effective and robust preconditioning. While classical methods such as algebraic multigrid (AMG) are highly scalable, their robustness can degrade on indefinite or nonsymmetric systems where heuristics originally developed for elliptic PDEs are less reliable. Recently, Graph Neural Networks (GNNs) have emerged as data-driven preconditioners; yet, the practical impact of imposing an AMG-style hierarchy remains underexplored for general sparse matrices. In this

arXiv (cs.LG) · Jul 30, 2026
4 min
Oracle-Budgeted Molecular Optimization with Short-Term Graph MemoryResearch

Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory

Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to generate. We introduce short-term graph memory, a plug-in module that preserves the generator architecture and native update rule while learning from previously evaluated molecules to prioritize subsequent oracle queries. The module maintains an online graph neural surrogate that pre-screens each round's candidate pool, so the fixed oracle budget is spent on molecules with higher predicted utility. Applied to a fragment-based generator on

arXiv (cs.LG) · Jul 30, 2026
4 min
Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image ClassificationResearch

Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

We introduce Kohn--Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral embedding evaluated at the Nishimori temperature of an associated Random-Bond Ising Model. By mapping pre-trained features onto quasi-cyclic low-density parity-check graphs and constructing a regularized Laplacian acting as a Kohn--Sham Hamiltonian, we solve $D$ independent channel spectral problems in $\mathcal{O}(N\log N + k^2_{\text{mode}} N)$ time via FFT on circulant blocks (leveraging Pontryagin self-duality of $\mathbb{Z}/p\mathb

arXiv (cs.LG) · Jul 30, 2026
4 min
Negative controls reveal volume-driven confounding in radiomics and imaging foundation model featuresResearch

Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features

Radiomics and imaging foundation models promise non-invasive biomarkers of tumour biology, yet predictive signatures may reflect tumour volume or acquisition artifacts rather than meaningful image structure. We introduce READII-2-ROQC, an open-source framework that uses volume-preserving negative controls to assess whether radiomic and deep imaging features capture independent spatial signals. READII-2-ROQC generates voxel-perturbed images across tumour, background and whole-image regions using configurable randomization strategies, then compares feature behaviour and model performan

arXiv (cs.LG) · Jul 30, 2026
4 min
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error CorrectionResearch

QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction

Fault-tolerant quantum computing (FTQC) relies on quantum error correction to suppress physical errors and preserve logical information at scale. In practice, however, performance is constrained not only by physical noise but also by the latency of classical decoders processing rapidly generated syndrome data. This challenge is exacerbated by hardware noise that is strong, heterogeneous, and nonstationary, as well as by the simulation-to-hardware distribution shift that can substantially degrade fixed neural decoders. We present QAdapt, a noise-adaptive neural pre-decoding framework

arXiv (cs.LG) · Jul 30, 2026
4 min
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width PruningResearch

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver practical throughput gains, but their input-agnostic computation allocation often causes substantial accuracy degradation under aggressive sparsity. Recent dynamic sparsity methods improve quality retention by adapting computation to individual inputs, yet they remain largely limited to coarse-grained structural decisions and their practical acceleration under real-world inference scenarios remains challenging. To address these challen

arXiv (cs.LG) · Jul 30, 2026
4 min
QQWorld: Quantile-Quantile Matching for World Model RegularizationResearch

QQWorld: Quantile-Quantile Matching for World Model Regularization

Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns pr

arXiv (cs.LG) · Jul 30, 2026
3 min
Gemini Robotics 2 brings whole body intelligence to robotsResearch

Gemini Robotics 2 brings whole body intelligence to robots

From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.

Google DeepMind · Jul 30, 2026
11 min
Windowed thinning and query complexity for the bouncy particle and Zigzag samplersResearch

Windowed thinning and query complexity for the bouncy particle and Zigzag samplers

Let $μ(d x)\propto e^{-U(x)} d x$ on $\R^d$, where $U$ is $m$-strongly convex and $L$-smooth, and denote by $κ=L/m$ the condition number. We consider windowed thinning, an exact simulation method for the bouncy particle sampler and the coordinate Zigzag process. The method divides a trajectory into deterministic windows and uses a gradient evaluation at the beginning of each window to construct a tractable local envelope for the event rate. Combining this construction with quantitative mixing estimates and finite-time bounds on the expected numbers of bounces and flips yields query c

arXiv (cs.LG) · Jul 30, 2026
3 min
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy TreesResearch

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify...

Papers with Code · Jul 30, 2026
2 min
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and EvaluationResearch

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reac...

Papers with Code · Jul 30, 2026
1 min
Scaling Vision-Language Models Is Not Enough to Mitigate BiasResearch

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understoo...

Papers with Code · Jul 30, 2026
1 min
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection LogsResearch

ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs

Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existing retrieval-augme...

Papers with Code · Jul 30, 2026
1 min
Distilling Answer Set Programming Theories from Large Language ModelsResearch

Distilling Answer Set Programming Theories from Large Language Models

Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model ...

Papers with Code · Jul 30, 2026
2 min
Contrastive Reinforced Policy Optimization via Privileged Self-DistillationResearch

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Di...

Papers with Code · Jul 30, 2026
1 min
Memory Decoder at Scale: A Pretrained, Parametric Long-Term MemoryResearch

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. M...

Papers with Code · Jul 30, 2026
2 min
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano PerformancesResearch

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual note...

Papers with Code · Jul 30, 2026
1 min
SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark AttackResearch

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capt...

Papers with Code · Jul 30, 2026
2 min
Introducing Gemini Robotics ER 2Research

Introducing Gemini Robotics ER 2

Gemini Robotics ER 2 is a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.

Google DeepMind · Jul 30, 2026
7 min
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic MemesResearch

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural...

Papers with Code · Jul 29, 2026
1 min
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAMResearch

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation sp...

Papers with Code · Jul 29, 2026
2 min
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?Research

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function. In this paper, we systematically study whether pretraining the Q-function actually helps when fine-tuning on

arXiv (cs.LG) · Jul 29, 2026
4 min
From Classification to Regression: Using a Fruitfly to Solve EquationsResearch

From Classification to Regression: Using a Fruitfly to Solve Equations

We present a novel approach to regression tasks using classification which is motivated by the mechanism used by fruitflies to sense their environment. Specifically, we formulate a general framework for learning nonlinear input-output relationships by replacing complex global surrogate models with a finite library of representative local patterns. Since scientific data often occupy limited and recurring regions of the input space, we generate predictions by measuring similarities between a query and stored patterns, then combining their associated responses through weighted reconstru

arXiv (cs.LG) · Jul 29, 2026
3 min
Can AI agents conduct open-ended AI research? Early evidence from two case studiesResearch

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors gr

arXiv (cs.AI) · Jul 29, 2026
4 min
APEX-AccountingResearch

APEX-Accounting

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mea

arXiv (cs.AI) · Jul 29, 2026
3 min
Inverse Learning of Latent Risk-Neutral Densities from Irregular Option QuotesResearch

Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes

Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only held-out market prices. A two-component lognormal mixture has the lowest aggregate price, $L^1$, Wasserstein, and fixed-tail errors on the synthetic benchmark. Learned operators retain narrower strengths: DeepONet reduces 1% quantile and variance error by 39.0% and 34.6% relative to the mixture, and a quote t

arXiv (cs.LG) · Jul 29, 2026
4 min
Inverse Learning of Latent Risk-Neutral Densities from Irregular Option QuotesResearch

Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes

Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A cont...

Papers with Code · Jul 29, 2026
1 min
The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-MakingResearch

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate

arXiv (cs.AI) · Jul 29, 2026
4 min
Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc TeamworkResearch

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint plan

arXiv (cs.AI) · Jul 29, 2026
4 min
Improving Item Discoverability in e-Commerce Search via Related Intent GenerationResearch

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items. In this paper, we present a scalable system for discovery-augmented search that leverages intent-conditioned recall expansion. Our approach generates implicit user intents to expand candidate recall while maintaining relevance. The sys

arXiv (cs.AI) · Jul 29, 2026
4 min
Improving Item Discoverability in e-Commerce Search via Related Intent GenerationResearch

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces ...

Papers with Code · Jul 29, 2026
2 min
When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic SystemsResearch

When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic Systems

Solving a continuous algebraic constraint system requires two decisions: which values satisfy the constraints, and which structural augmentation renders an unsolvable system solvable. Classical solvers answer the first well and the second only by enumeration. On that discrete decision, a candidate-conditioned repair ranker choosing among K augmentations reaches the exhaustive-search ceiling at a fraction of the calls, outperforming random (0.997 vs 0.236 balanced nonlinear menu accuracy; p < 10^-70; 0.982 +/- 0.006 across seeds) and beating a budget-matched per-candidate probe on acc

arXiv (cs.LG) · Jul 29, 2026
4 min
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from ScratchResearch

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamental...

Papers with Code · Jul 29, 2026
2 min
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingResearch

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to com

arXiv (cs.AI) · Jul 29, 2026
4 min
Anatomy Contextualized Adaption of CT Foundation ModelsResearch

Anatomy Contextualized Adaption of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight frame

arXiv (cs.AI) · Jul 29, 2026
4 min
Skillful forecasting of offshore winds from satellite scatterometer constellationsResearch

Skillful forecasting of offshore winds from satellite scatterometer constellations

Accurate intraday forecasts of offshore wind are becoming increasingly important for power system operation and the integration of growing shares of offshore wind energy. Operational forecasts rely predominantly on numerical weather prediction (NWP), which is not optimized for lead times of minutes to hours, where initial-condition accuracy dominates forecast skill. Although satellite scatterometer observations are routinely assimilated into NWP, they have not previously been used directly for forecasting. Here we present WindCastNet, the first satellite-based nowcasting framework fo

arXiv (cs.LG) · Jul 29, 2026
4 min
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program SynthesisResearch

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap,

arXiv (cs.LG) · Jul 29, 2026
4 min
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical ApplicationsResearch

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, th...

Papers with Code · Jul 29, 2026
2 min
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain BenchmarkResearch

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and cost-

arXiv (cs.AI) · Jul 29, 2026
4 min
Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devicesResearch

Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devices

This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic. With the target workload of branch prediction (BP) for multistage pipelined central pro-cessing unit (CPU) cores. For this, a novel memristor based RC design framework was developed within the context of the workload requirements. This was then implemented in simulation using industry standard modelling languages of System Verilog (SV) and Verilog-AMS (VAMS).The developed RC design framework was subsequently verified using a bas

arXiv (cs.LG) · Jul 29, 2026
4 min
DLAM: Distributional Latent Actions with Temporal ConstraintsResearch

DLAM: Distributional Latent Actions with Temporal Constraints

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional

arXiv (cs.AI) · Jul 29, 2026
4 min
Linguistic Monoculture in LLM-Assisted Language UseResearch

Linguistic Monoculture in LLM-Assisted Language Use

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shar

arXiv (cs.AI) · Jul 29, 2026
4 min
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision ProcessesResearch

Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In particular, we construct the stable quotient, the coarsest observation-wise abstraction preserving one-step rewards and quotient successors, and prove

arXiv (cs.LG) · Jul 29, 2026
4 min
AgentMap: Joint Equivalence and Subsumption Discovery for Ontology MatchingResearch

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented by a series of interdependent semantic decisions. Given a concept in the source ontology,

arXiv (cs.AI) · Jul 29, 2026
3 min
Voronoi Histograms for Adaptive Vectorization of Expected Persistence DiagramsResearch

Voronoi Histograms for Adaptive Vectorization of Expected Persistence Diagrams

Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves as a distribution of topological features. Existing EPD vectorizations often rely on predefined point transformations, such as Gaussian or landscape functions. We study an alternative discretization based on Voronoi histograms, which trades smooth functional approximation for adaptive partition-based counting.

arXiv (cs.LG) · Jul 29, 2026
3 min
Voronoi Histograms for Adaptive Vectorization of Expected Persistence DiagramsResearch

Voronoi Histograms for Adaptive Vectorization of Expected Persistence Diagrams

Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. Expected Persistence Diagram (EP...

Papers with Code · Jul 29, 2026
1 min
MMAC: A Massive Multi-dimensional Benchmark for Audio CaptioningResearch

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model

arXiv (cs.AI) · Jul 29, 2026
3 min
Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department ForecastingResearch

Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting

Emergency Departments (EDs) are critical access points in healthcare systems, yet they face persistent pressure from unpredictable patient demand, seasonal surges, and non-urgent visits. Effective ED planning requires forecasts at multiple decision-making levels: hospitals need local demand estimates for staffing and bed management, regions require forecasts to coordinate healthcare units, and national authorities need system-wide projections for capacity planning. However, most existing approaches forecast ED demand independently at a single level, ignoring the hierarchy linking hos

arXiv (cs.LG) · Jul 29, 2026
4 min
Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approachResearch

Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approach

Most existing seizure detection algorithms require extensive pre-processing of the data and rely on heuristic or currently unexplainable machine learning approaches. These approaches often struggle with balancing detection sensitivity and specificity in the presence of variable seizure morphologies, interictal epileptiform discharges, and artefacts. Here, we consider an alternative approach: our seizure detection algorithm, which is based on the concept of critical transitions and overcomes the aforementioned limitations. Specifically, we perform a receiver-operating-characteristic a

arXiv (cs.LG) · Jul 29, 2026
4 min
Sky sphere representation in language modelsResearch

Sky sphere representation in language models

We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. We find that most of the considered open-source models do have such a representation, and it often even surfaces to the top principal components on prompts that ask questions like ``what is close to this object in the night sky''. In all but one model this representation showed significant scores in LOO testing, containing up to 65-85% of variance ($R^2$-score) and having median angular error down to $12^\circ-21^\circ$. We verify that our represen

arXiv (cs.LG) · Jul 29, 2026
3 min
InferScale: GPU-Native KV Injection for Personalized LLM ServingResearch

InferScale: GPU-Native KV Injection for Personalized LLM Serving

Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, tha...

Papers with Code · Jul 29, 2026
2 min
InferScale: GPU-Native KV Injection for Personalized LLM ServingResearch

InferScale: GPU-Native KV Injection for Personalized LLM Serving

Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. As the retrieval budget grows, time-to-first-token (TTFT) increases even though the underlying memory is reused across requests. We present InferScale, a GPU-native LLM memory system that replac

arXiv (cs.LG) · Jul 29, 2026
4 min
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript ContextResearch

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-te

arXiv (cs.AI) · Jul 29, 2026
4 min
Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM AgentsResearch

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches leave acquisition under heterogeneous costs unaddressed. We formulate this decision as cost-aware marginal decision-focused stopping (CAM-DF) over

arXiv (cs.AI) · Jul 29, 2026
4 min
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust RealignmentResearch

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susc

arXiv (cs.AI) · Jul 29, 2026
4 min
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and RepairResearch

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory security, yet few trace the same malicious semantics across persistence, downstream consequences, and selective repair under diverse memory-backend comparisons. To address this gap, we introduce MemSecBench, a task-grounded benchmark for the lifecycle security of agent

arXiv (cs.AI) · Jul 29, 2026
4 min
Field Codes for Distributed Coupling Samplers and Certified Empirical TransportResearch

Field Codes for Distributed Coupling Samplers and Certified Empirical Transport

In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and scalar value-certified sampling. Our main result is a field-code compiler: any communicated transport field approximating an optimal empirical Monge map to error $η$ can be completed by sparse target-cell residuals into an exact-marginal value-certified sampler with scalar certificate $W_1(μ,ν)\leq U\leq W_1(μ,ν)+2Δ$, where $Δ$ is the public target-partition diameter. The certificate accuracy is controlled by $Δ$ alone. The field er

arXiv (cs.LG) · Jul 29, 2026
4 min
Field Codes for Distributed Coupling Samplers and Certified Empirical TransportResearch

Field Codes for Distributed Coupling Samplers and Certified Empirical Transport

In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and sca...

Papers with Code · Jul 29, 2026
2 min
Equilibrium Training of Energy-Based Models with Parallel Trajectory TemperingResearch

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning. This enables stable and fast training on highly multimodal and data-scarce scientific datasets. Combined with reservoir sampling and adaptive optimization, PTT has a computational cost comparable to Persistent Contrastive

arXiv (cs.LG) · Jul 29, 2026
3 min
Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning FrameworkResearch

Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework

Continuous cuffless blood pressure (BP) monitoring remains challenging due to motion artifacts, physiological variability, and the limited robustness of conventional pulse transit time (PTT) models under dynamic conditions. Many prior approaches rely on multi-second windows to stabilize estimation, an assumption that is frequently violated during real-world monitoring with intermittent signal corruption. Here, we show that discriminative BP-related information is preserved at the single-beat level and present a lightweight multi-modal wearable framework for continuous BP estimation.

arXiv (cs.LG) · Jul 29, 2026
4 min
Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed NoiseResearch

Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret in a parameter-free manner remains an open challenge. We resolve this by proposing \textbf{HT-PAder}, a parameter-free algorithm combining restarted AdaGrad experts over a geometric pool of block lengths with a pathwise meta-algorithm, \textbf{AdaGrad-Hedge}, which requires no moment conditions on meta-losses.

arXiv (cs.AI) · Jul 29, 2026
3 min
Visual Credit Audit for Multimodal Spatial ReasoningResearch

Visual Credit Audit for Multimodal Spatial Reasoning

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-al

arXiv (cs.AI) · Jul 29, 2026
4 min
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript EvidenceResearch

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript

arXiv (cs.AI) · Jul 29, 2026
4 min
ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch DetectionResearch

ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection

While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale annotated synthetic training data using BlenderProc, with configurable material appearance, camera modes, and domain randomization, producing automatic COCO-format annotations. To show the potential of our approach, we evaluate four training strategies, namely synthetic-only, real-only, mixed, and fine-tuning from

arXiv (cs.AI) · Jul 29, 2026
4 min
PIKS: Universal Physics-Informed Kernel MethodsResearch

PIKS: Universal Physics-Informed Kernel Methods

Physics-informed machine learning incorporates physical principles --often expressed via differential operators-- into data-driven models. While physics-informed neural networks (PINNs) dominate empirical applications, the complexity of neural network architectures and optimization landscapes hinders the development of a corresponding learning theory. In turn, kernel methods offer an appealing alternative with closed-form solutions and analytical tractability, yet existing guarantees primarily cover the well-specified setting where the target belongs to the native Reproducing Kernel

arXiv (cs.LG) · Jul 29, 2026
3 min
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous DataResearch

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with h

arXiv (cs.AI) · Jul 29, 2026
4 min
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge DistillationResearch

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The ...

Papers with Code · Jul 29, 2026
2 min
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge DistillationResearch

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invaria

arXiv (cs.AI) · Jul 29, 2026
4 min
GPTQ-2D: Cubic-Time Two-Sided Adaptive RoundingResearch

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. They process the entries in a fixed order, one at a time, propagating each rounding error to the entries not yet processed through a triangular feedback matrix. We study the two-sided version of this task, in which fixed nonsingular basis matrices act on both the left and the right of the residual; the familiar one-sided case is the special case of an identity right basis. Vectorizing the matrix turns the two-sided objective into a quadrat

arXiv (cs.LG) · Jul 29, 2026
3 min
Mitigating Compounding Error via Video Representation RegularizationResearch

Mitigating Compounding Error via Video Representation Regularization

Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional

arXiv (cs.LG) · Jul 29, 2026
4 min
BayesAME: Bayesian Active Model EvaluationResearch

BayesAME: Bayesian Active Model Evaluation

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically ta

arXiv (cs.AI) · Jul 29, 2026
4 min
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual PerceptionResearch

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invariant transformations that yield a geometric grid structure. The state is modelled with ego-

arXiv (cs.AI) · Jul 29, 2026
4 min
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentResearch

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic...

Papers with Code · Jul 29, 2026
2 min
Progressive Multimodal Alignment for Continual Instruction TuningResearch

Progressive Multimodal Alignment for Continual Instruction Tuning

Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving instruction semantics cause this shared projector to drift, leading to projector-level forgetting, an issue largely overlooked by methods that focus primarily on the LLM backbone. We introduce Progressive Multimodal Alignment (PMA), a framework that enables the projector to adapt continually while preserving p

arXiv (cs.AI) · Jul 29, 2026
3 min
Surrogate assisted diversity estimation in neural ensemble searchResearch

Surrogate assisted diversity estimation in neural ensemble search

Ensembles are a standard way to improve the performance and robustness of deep neural networks, but their effectiveness crucially depends on both the quality...

Papers with Code · Jul 29, 2026
1 min
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth ObservationResearch

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks,...

Papers with Code · Jul 29, 2026
1 min
ToxScreen: Detecting Whether an LLM Has Been PoisonedResearch

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly...

Papers with Code · Jul 29, 2026
2 min
Budget-Aware LLM Discovery via Cost-Calibrated Frontier UtilityResearch

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive di...

Papers with Code · Jul 29, 2026
2 min
AI as Friction for Reflection Support in IdeationResearch

AI as Friction for Reflection Support in Ideation

Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output t...

Papers with Code · Jul 29, 2026
1 min
FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated LearningResearch

FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architecture...

Papers with Code · Jul 29, 2026
1 min
PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classificationResearch

PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification

Breast DCE-MRI AI is increasingly being explored for breast-level classification of no-lesion, benign, and malignant findings, beyond conventional lesion-cen...

Papers with Code · Jul 29, 2026
2 min
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill EvolutionResearch

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning tre...

Papers with Code · Jul 29, 2026
2 min
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLMResearch

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater...

Papers with Code · Jul 29, 2026
2 min
MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software EntitiesResearch

MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities

Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the trade-off between t...

Papers with Code · Jul 29, 2026
1 min
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill OverfittingResearch

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. ...

Papers with Code · Jul 29, 2026
1 min
RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge DeploymentResearch

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, ex...

Papers with Code · Jul 29, 2026
2 min
Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure ModesResearch

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified i...

Papers with Code · Jul 29, 2026
1 min
Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting MethodsResearch

Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assum...

Papers with Code · Jul 29, 2026
1 min
ContactFlow: A video action conditioning that transfers across embodimentsResearch

ContactFlow: A video action conditioning that transfers across embodiments

World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, cu...

Papers with Code · Jul 29, 2026
1 min
ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless PlatformResearch

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittent...

Papers with Code · Jul 29, 2026
2 min
A Graph-Native Bitemporal Memory Store for Conversational AI AgentsResearch

A Graph-Native Bitemporal Memory Store for Conversational AI Agents

Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window, or de...

Papers with Code · Jul 29, 2026
2 min
From Interface to Inference: Eliciting Any-Order Inference from Any-Order ModelsResearch

From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models

Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process...

Papers with Code · Jul 29, 2026
2 min
Learning Dynamic User Personas from Implicit Interaction Streams via Iterative RefinementResearch

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit...

Papers with Code · Jul 29, 2026
2 min
MultivationBench: A Benchmark for Multimodal Sequential Motivation ReasoningResearch

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequenti...

Papers with Code · Jul 29, 2026
1 min
CG-World: A Large-Scale World-State Dataset and Protocol for World ModelsResearch

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually captur...

Papers with Code · Jul 29, 2026
2 min
Reinforcement Learning on Cost-Constrained Quadrupedal HardwareResearch

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-t...

Papers with Code · Jul 29, 2026
1 min
PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak SupervisionResearch

PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision

Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at...

Papers with Code · Jul 29, 2026
2 min
Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion SegmentationResearch

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially ...

Papers with Code · Jul 29, 2026
2 min
Q-Steer: Action-Value Guidance for Molecular Policy OptimizationResearch

Q-Steer: Action-Value Guidance for Molecular Policy Optimization

Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. ...

Papers with Code · Jul 29, 2026
2 min
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-TuningResearch

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a r...

Papers with Code · Jul 29, 2026
2 min
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMsResearch

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to pe...

Papers with Code · Jul 29, 2026
1 min
We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control.Research

We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control.

Our newest music generation model, Lyria 3.5, delivers significant advancements across musicality, lyrics, and vocal quality, empowering you to craft richer tracks. We’r…

Google DeepMind · Jul 29, 2026
1 min
Pramana: A Composable, Domain-Specific Backend for Empirical Networking ResearchResearch

Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research

Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a hypoth...

Papers with Code · Jul 28, 2026
2 min
Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck CancerResearch

Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment outcome...

Papers with Code · Jul 28, 2026
2 min
Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point FreightResearch

Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight

Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address only o...

Papers with Code · Jul 28, 2026
2 min
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response RewritingResearch

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate t...

Papers with Code · Jul 28, 2026
2 min
Position: Evaluation Scores Are Perishable Knowledge ClaimsResearch

Position: Evaluation Scores Are Perishable Knowledge Claims

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and ...

Papers with Code · Jul 28, 2026
2 min
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage ClosureResearch

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly ...

Papers with Code · Jul 28, 2026
2 min
Pass the Baton: Trajectory-Relayed On-Policy DistillationResearch

Pass the Baton: Trajectory-Relayed On-Policy Distillation

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). Dur

arXiv (cs.AI) · Jul 28, 2026
4 min
$π\mathbf{R}^2$: Reactive Real-time Flow PoliciesResearch

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $π\mathbf{R}^2$, which makes thes

arXiv (cs.AI) · Jul 28, 2026
4 min
Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRAResearch

Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA

Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncertainty signal: peaked mass indicates confidence, while a flat distribution indicates ambiguity. We introduce CARE (Confidence-Adaptive Routing of Experts), which admits experts in a nucleus fashion. Experts are activated in decreasing router weight until their cumulative

arXiv (cs.LG) · Jul 28, 2026
4 min
Re-thinking Mammography Transfer Learning: The Dataset-Informed Transfer Learning (DITL) Framework for Breast Cancer Screening and Lesion DiagnosisResearch

Re-thinking Mammography Transfer Learning: The Dataset-Informed Transfer Learning (DITL) Framework for Breast Cancer Screening and Lesion Diagnosis

Enhancing classification performance in mammography remains a persistent challenge across both small curated datasets and large-scale clinical cohorts. Conventional transfer learning approaches often neglect dataset-specific characteristics, while recent neighborhood-informed methods have been restricted to narrow tasks with rigid formulations, limiting their scalability to population-level datasets. To address these challenges, we propose the Dataset-Informed Transfer Learning (DITL) framework, which integrates dataset-derived difficulty signals with neighborhood-based triplet super

arXiv (cs.LG) · Jul 28, 2026
4 min
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease ScreeningResearch

VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with optional symptom descriptions, to a server-hosted vision-language model for zero-shot disease classification. The system separates agent interaction from workflow orchestration: OpenClaw provides scheduling, tool access, user interaction, and notification services on the edge device, while LangGraph manages the stateful screening workflow, including input validation, image transmission, mode

arXiv (cs.LG) · Jul 28, 2026
3 min
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?Research

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carr

arXiv (cs.AI) · Jul 28, 2026
4 min
Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent GuidanceResearch

Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance

Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learni

arXiv (cs.LG) · Jul 28, 2026
3 min
Falling Behind Drives Unsafe Development in an Idealised AI Race ExperimentResearch

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but

arXiv (cs.AI) · Jul 28, 2026
4 min
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot TransferResearch

CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately, zero-shot transfer on multimodal graphs remains underexplored. Existing GNN-based graph foundation models typically require downstream adaptation, wher

arXiv (cs.AI) · Jul 28, 2026
4 min
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent CrossbarResearch

MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar

Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expensive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifically, MDTransformer performs complex m

arXiv (cs.AI) · Jul 28, 2026
4 min
Pictura: Perspective-View Self-Play at Scale for DrivingResearch

Pictura: Perspective-View Self-Play at Scale for Driving

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play

arXiv (cs.AI) · Jul 28, 2026
4 min
Parallel Decoding Distillation for Fast Image and Video GenerationResearch

Parallel Decoding Distillation for Fast Image and Video Generation

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In this paper, we introduce Parallel Decoding Distillation (PDD), a simplified and scalable traj

arXiv (cs.LG) · Jul 28, 2026
3 min
Sharpness-Aware Minimization and Muon: Robustness under the Spectral NormResearch

Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study

arXiv (cs.LG) · Jul 28, 2026
3 min
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation ModelsResearch

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained and evaluated on independent and identically distributed data, but this assumption changes in real-world scenarios due to distribution shifts, which compromise the robustness of models. Limited research has been conducted of TFMs under distribution shifts. We present an empirical evaluation of Out-Of-Distribution (OOD) performance of nine TFMs, spanning diverse pre-training strategies and

arXiv (cs.AI) · Jul 28, 2026
4 min
Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?Research

Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?

Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that large language models (LLMs) can automate cluster security remediation, generating configuration patches from Kubernetes Security Posture Management (KSPM) findings without human authoring. Such systems, however, prompt the model with each finding in isolation from the live service call graph, assuming general hardening knowledge suffices. This assumption breaks down whenever a patch must preserve a runtime service dependency invisible to the model: an otherwise compl

arXiv (cs.AI) · Jul 28, 2026
4 min
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based AgentsResearch

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically adopt a coarse-grained (utility-agnostic) manner that treats heterogeneous user-LLM interaction records uniformly, leading to redundant and low-impact records persisting in the memory repository. To address this challenge, we present MemLens, a value-aware memory management system that takes memory records as first-class data objects. MemLens provides an end-to-end i

arXiv (cs.AI) · Jul 28, 2026
3 min
Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving NetworksResearch

Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving Networks

The vision of self-driving networks that monitor, reason, and act upon themselves with minimal human intervention relies on tightly coupled monitoring, analytics, and actuation functions. In this work, we treat these functions as three operational macro-intents: continuous telemetry, real-time analytics, and programmatic actuation, and formalize the health of each function as an intent that the network must continuously satisfy. A critical, yet underexplored, challenge stems from the causal coupling among these intents, where a singular fault within one macro-intent propagates as a c

arXiv (cs.LG) · Jul 28, 2026
4 min
Generator-Aligned Representation Interfaces for Diagnostic Soft EquivarianceResearch

Generator-Aligned Representation Interfaces for Diagnostic Soft Equivariance

Exact-equivariant architectures typically encode prescribed group actions in specialized operators, which can complicate their reuse with generic backbones and across data modalities. We introduce the Generator-Aligned Representation Interface (GARI), a representation-level design principle that exposes selected transformation generators to a generic sequence backbone through aligned canonical and generator-induced views. We formalize the resulting behavior using a probe-specific soft-equivariance residual defined over declared data and transformation distributions. This framework di

arXiv (cs.LG) · Jul 28, 2026
4 min
Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator DynamicsResearch

Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator Dynamics

Unmanned aerial vehicles (UAVs), particularly quadcopters, present unique challenges for autonomous control due to their underactuated dynamics: only four available control inputs must govern six degrees of freedom. This paper investigates a physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques $(T, τ_x, τ_y, τ_z)$, and closes the loop through a high-fidelity Simulink environment. Our simulator integrates a 12-state rigid-body model (MATLAB Level-2 S-Function) with (i) an Action2RPM allocation b

arXiv (cs.LG) · Jul 28, 2026
4 min
Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene KinematicsResearch

Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics

Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a probabilistic model of scene kinematics that learns a structured spatio-temporal latent representation of the distribution over possible futures given an im

arXiv (cs.LG) · Jul 28, 2026
3 min
Reinforcement Learning for Code OptimizationResearch

Reinforcement Learning for Code Optimization

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnable through three stages: (1) how code is tested, by building DMC-Optim with large optimiza

arXiv (cs.AI) · Jul 28, 2026
4 min
Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imagingResearch

Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

Singular Value Decomposition (SVD) underlies matrix factorisation tasks across computational imaging, with medical applications increasingly demanding real-time processing. Yet SVD algorithms are inherently sequential, constraining real-time GPU throughput and limit online deployment in clinical pipelines. This study introduces Quasi-SVD, a differentiable, fully parallelized matrix factorization framework for GPUs. Rather than enforcing orthogonality on both factors, it guarantees exact orthogonality for a single Lie-parameterized factor while recovering the remaining components thro

arXiv (cs.LG) · Jul 28, 2026
4 min
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy RecognitionResearch

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognition of A/H at the video level is difficult, since the signal arises from disagreement across and within facial, vocal, linguistic, and bodily modalities, and manifests differently across individuals. The proposed PRISM-AH (Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition), is a framework that treats A/H as a multimodal conflict that unfolds over time. Frozen vision, audio, and text encoders are aligned

arXiv (cs.AI) · Jul 28, 2026
4 min
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge GraphsResearch

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a practical question: when these modalities disagree, how can we detect and explain the conflict? We study this problem as \emph{modality-level inconsistency detection}. We first introduce a taxonomy of cross-modal knowledge inconsistencies, covering information granularity differences, direct conflicts, temporal changes, and KG incompleteness. We then present \text

arXiv (cs.AI) · Jul 28, 2026
3 min
Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory AllocationResearch

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility. We study this issue as instance-wise operations research (OR) formulation selection, where each allocation instance is assigned to a solver-executable formulation from a candidate OR expert library. We propose a solver-guided large language model (LLM) framework for OR formulat

arXiv (cs.AI) · Jul 28, 2026
4 min
MODUS: Decoder-Only Any-to-Any Modeling of Diverse ModalitiesResearch

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitr

arXiv (cs.AI) · Jul 28, 2026
4 min
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time SeriesResearch

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture s

arXiv (cs.AI) · Jul 28, 2026
4 min
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical CasesResearch

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multim

arXiv (cs.AI) · Jul 28, 2026
4 min
Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?Research

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Deep generative models (DGMs) are widely used for complex high-dimensional data and increasingly applied to spatial and spatio-temporal modeling. Their generated samples implicitly represent the learned data distribution and associated uncertainty. However, for real-world data, assessing whether DGMs have learned the underlying process is difficult because the ground truth is unknown and evaluation often relies on observations alone. We evaluate representative DGMs, flow matching (FM), DDPM, score-SDE, and VAE, on a known non-stationary Gaussian random field. This paper provides comp

arXiv (cs.LG) · Jul 28, 2026
3 min
Face De-Identification: A Domain-Centric Survey from Capture to ProcessingResearch

Face De-Identification: A Domain-Centric Survey from Capture to Processing

Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks. With the rising emphasis on data privacy and responsible AI, face De-ID has emerged as an active research area spanning computer vision and privacy-preserving communities. Early approaches, and many contemporary ones, operate in the digital domain by modifying pixel-level or appearance-level features through post-capture processing. Recent advances extend face De-ID beyond post-processing by

arXiv (cs.AI) · Jul 28, 2026
4 min
dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision TreesResearch

dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees

Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comprehensible. Unfortunately, reducing the size of the decision tree is not straightforward, as missing just a single crucial case might result in an incorrect controller. We tackle this issue in the setting of Markov decision processes, extending dtControl2 by "$\varepsilon$" functionality:

arXiv (cs.AI) · Jul 28, 2026
3 min
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QAResearch

Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini, GPT, Qwen, Gemma, Llama, and Ministral) under a zero-shot prompting setting and analyse

arXiv (cs.AI) · Jul 28, 2026
4 min
Penelope: Localized Latent Recurrence for Efficient Structured ReasoningResearch

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-condit

arXiv (cs.AI) · Jul 28, 2026
4 min
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous NetworksResearch

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool trust management. The model comprises: a formally defined trust state machine with provable graduated enforcement, damped cascade propagation w

arXiv (cs.AI) · Jul 28, 2026
3 min
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action ModelsResearch

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object

arXiv (cs.AI) · Jul 28, 2026
4 min
AnnoBench: A Benchmark for Visualization Annotation GenerationResearch

AnnoBench: A Benchmark for Visualization Annotation Generation

Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for v

arXiv (cs.AI) · Jul 28, 2026
3 min
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language ModelsResearch

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regu

arXiv (cs.AI) · Jul 28, 2026
4 min
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneResearch

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relati

arXiv (cs.LG) · Jul 28, 2026
4 min
A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional FieldsResearch

A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields

In this paper, we present an automated data-driven workflow using Machine Learning (ML) for gas lift optimization in unconventional fields. This workflow integrates a ML model that accurately forecasts the Gas Lift Performance Curve, and a Bayesian Optimization Framework to solve for the optimal gas injection rates under the constraints of facility capacity. The ML model leverages the historical production time series data without requiring downhole gauges or multi-rate well tests. We piloted this workflow on 30 wells across 5 well pads in Bakken and obtained >5% production uplift on

arXiv (cs.LG) · Jul 28, 2026
4 min
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor NetworksResearch

A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks

Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies typically assume fixed sensor graphs, overlooking the continuous evolution of real-world traffic networks, e.g., ongoing road network construction and evolving human mobility patterns. These dynamic changes can substantially degrade conventional forecasting models, motivating test-time adaptation (TTA) to efficiently adapt pretrained models during deployment. However, applying TTA to evolving traffic sensor networks remains challenging in two aspe

arXiv (cs.LG) · Jul 28, 2026
4 min
VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge DeploymentResearch

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment

Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent layers, or non-causal post-processing. We propose kiloVAD, designed for embedded inference using standard Mel features, CNN-only layers, and tunable context/spectral parameters. We introduce per-layer structured pruning with self-distillation and angle-based quantization-aware training (QAT) that outperforms sta

arXiv (cs.LG) · Jul 28, 2026
3 min
DRIFT: Direct-Recursive Intervention-Conditioned Forecasting of ICU Physiological TrajectoriesResearch

DRIFT: Direct-Recursive Intervention-Conditioned Forecasting of ICU Physiological Trajectories

Many time-series forecasts depend not only on prior observations but also on actions specified during the forecast period. In intensive care units (ICUs), future vital signs and laboratory values are influenced by treatments such as vasopressors. However, models that predict the full future sequence all at once make little use of these treatments, whereas autoregressive models can accumulate errors. We introduce DRIFT, a hybrid framework in which a direct model produces the primary forecast and a recursive, action-conditioned model contributes constrained corrections. We evaluate DRI

arXiv (cs.LG) · Jul 28, 2026
4 min
Prototype Adaptation for Zero-Shot sEMG Movement ClassificationResearch

Prototype Adaptation for Zero-Shot sEMG Movement Classification

Surface electromyography (sEMG) enables the control of prostheses, allowing upper-limb amputees to re-gain some hand function. Most current research focuses on recognizing basic movements for prosthesis control. However, in most daily activities, such as opening a door, combined movements are essential. However, collecting training data for all possible combined movements is time-consuming and requires re-training of the model for any new combination. We propose two novel recognition approaches, Compositional Prototype Interpolation (CPI) and Synthetic Adaptation for Prototypes (SAP)

arXiv (cs.LG) · Jul 28, 2026
3 min
SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam DynamicsResearch

SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics

This paper proposes a novel physics-guided spectral deep operator network, termed SpectONet, for solving Euler-Bernoulli beam (EBB) vibration problems. The proposed framework integrates the operator-learning capability of DeepONet with physics-informed constraints and Chebyshev-Gauss-Lobatto (CGL) sensor placement. Unlike conventional DeepONet frameworks, which commonly employ uniformly distributed sensors, SpectONet uses nonuniform spectral sensor locations with a higher concentration of points near the domain boundaries. This sampling strategy improves the finite-dimensional repres

arXiv (cs.LG) · Jul 28, 2026
4 min
Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite OptimizationResearch

Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite Optimization

We study stochastic composite nonconvex optimization over a compact convex set when gradient samples arrive along a single trajectory of a fixed ergodic Markov chain. Existing single-trajectory variance-reduction theory covers smooth unconstrained objectives; we address the projection-free composite setting using the generalized Frank-Wolfe gap. We propose MC-ALFCG, which combines a momentum conditional-gradient method with coupled capped multilevel Monte Carlo estimation and per-iteration clipping. The deepest nested average uses consecutive states from the same trajectory, yielding

arXiv (cs.LG) · Jul 28, 2026
3 min
Rashomon AlignmentResearch

Rashomon Alignment

We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measures are distributio...

Papers with Code · Jul 28, 2026
1 min
Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess PositionsResearch

Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that hum...

Papers with Code · Jul 28, 2026
2 min
SPARC Segmentation to Prediction via Affine Regression and CounterfactualsResearch

SPARC Segmentation to Prediction via Affine Regression and Counterfactuals

Transaction propensity prediction in B2B e commerce presents unique challenges distinct from B2C contexts, primarily due to the heterogeneous procurement beh...

Papers with Code · Jul 28, 2026
2 min
Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive RoboticsResearch

Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics

In the heart of Pittsburgh, a quiet revolution is underway — one that promises to redefine what independence means for people with disabilities. At the center of this movement is the Human Engineering Research Laboratories (HERL), a pioneering institute at the University of Pittsburgh, now leading an initiative with up to $41.5 million in funding from the Advanced Research Projects Agency for Health (ARPA-H), a US research funding agency established to support transformative biomedical and healt

Meta AI · Jul 28, 2026
8 min
PanoLess: Environment Reconstruction from Partial Reflective ViewsResearch

PanoLess: Environment Reconstruction from Partial Reflective Views

Reflections from shiny objects and glass facades naturally extend the field of view of a camera, capturing the surrounding environment without the need to pa...

Papers with Code · Jul 28, 2026
1 min
ScoreShield: Differentially Private Release of Similarity ScoresResearch

ScoreShield: Differentially Private Release of Similarity Scores

A growing number of applications, such as biometrics and retrieval-augmented generation (RAG), rely on cosine similarity scores computed between vector embed...

Papers with Code · Jul 27, 2026
2 min
Localized Anomaly Detection via Differentiable D-vine CopulasResearch

Localized Anomaly Detection via Differentiable D-vine Copulas

Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. F...

Papers with Code · Jul 27, 2026
2 min
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical UnderstandingResearch

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional an

arXiv (cs.AI) · Jul 27, 2026
4 min
Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal TransportResearch

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present TemporalSinkhorn, a parallel-in-time executor that batches future candidates and their repairs without making output accuracy speculative. A centered, row-sharded certificate accepts only a deterministic safe prefix. The remaining candidates share packed Sinkhorn updates; an online projective forgetting rate places audit milestones, while a p

arXiv (cs.LG) · Jul 27, 2026
4 min
Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal TransportResearch

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed S...

Papers with Code · Jul 27, 2026
2 min
Learning Distributions from Multiple Data ProvidersResearch

Learning Distributions from Multiple Data Providers

Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution $p$ on a finite domain $[n]$. The learner is given a fixed family of queryable sets $\mathscr{S} \subseteq 2^{[n]}$, and each query to $S \in \mathscr{S}$ returns an independent sample from the conditional distribution $p(\cdot \mid S)$. Learnability is governed by the co-occurrence graph associated with $\mathscr{S}$: two domain elements are adjacent if they appear together in som

arXiv (cs.LG) · Jul 27, 2026
4 min
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical ExplainabilityResearch

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual model. With the emergence of Kolmogorov-Arnold Networks (KANs), whose spline-based components provide inherently interpretable functional units, we investigate whether this architectural transpa

arXiv (cs.AI) · Jul 27, 2026
4 min
Rethinking Classifier-Free Guidance in On-Policy Diffusion DistillationResearch

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases,

arXiv (cs.AI) · Jul 27, 2026
4 min
Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEsResearch

Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural network is trained to approximate the PDE solution by using (stochastic) gradient descent to minimize the PDE residual of the neural network. Due to the non-convexity of the PDE residual objective function, the trained neural network may, in principle, only converge to a local minimizer of the objective function (which would not be a solution

arXiv (cs.LG) · Jul 27, 2026
3 min
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationResearch

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explic

arXiv (cs.AI) · Jul 27, 2026
4 min
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining DataResearch

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose DataOrchestra, a framework that unifies different processing operations and orchestrates an example-specific pipeline for each example. Given a chunk of pretraining data, an orchestrator decides whether to drop, untouch, or clean it. For a chunk to be cleaned, it selects one or more downs

arXiv (cs.AI) · Jul 27, 2026
3 min
Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion ArchitecturesResearch

Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures

Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit movements within a given architecture. We present the first study in which a single frontier large language model (LLM), Claude Opus 4.7, generates and iteratively refines the full Python code of shuttling compilers from written specifications. We start with a compiler for (i) a linear segmented trap, extend it to (ii) a trap with junctions, and finally achieve efficient compilation for (iii) a broad class of connected trap graphs. The compilers for the more general c

arXiv (cs.AI) · Jul 27, 2026
4 min
ERUnderstand: Evaluating Vision-Language Models on Structured ER DiagramsResearch

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a stand

arXiv (cs.AI) · Jul 27, 2026
3 min
Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference PipelinesResearch

Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions. Across several application domains, we abstract this inference architecture as a fast path, a slow path, and a coordination layer with two functions: a router that invokes the slow path and a merger that decides whether to incorporate its returned predictions. In this work, we show

arXiv (cs.AI) · Jul 27, 2026
4 min
Beyond Scale and Generation: Understanding Language Model-based Entity MatchingResearch

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 famil

arXiv (cs.LG) · Jul 27, 2026
4 min
Stacking the Deck: Tunable Trainability in Stacked LCUsResearch

Stacking the Deck: Tunable Trainability in Stacked LCUs

Variational quantum circuits have been central to many proposed near-term applications of quantum computing, but a growing body of evidence suggests that trainability and quantum advantage are fundamentally at odds: ansätze expressive enough to resist efficient classical simulation tend to exhibit barren plateaus, while structures that provably rule out barren plateaus typically render them classically simulable. We propose a stacked linear combination of unitaries (S-LCU) as a variational ansatz which provides a tunable trade-off between barren plateaus and classical simulability. U

arXiv (cs.LG) · Jul 27, 2026
3 min
Co-Learning for Missing Arbitrary Modalities in Multi-modal ClassificationResearch

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle missing modalities, prior studies have mainly covered bimodal data setups and focused on designing robust fusion processes. Instead, we adopt a multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion. Speci

arXiv (cs.AI) · Jul 27, 2026
3 min
Causal-TS: A Python Library for Causal Discovery in High-Dimensional and Nonstationary Time SeriesResearch

Causal-TS: A Python Library for Causal Discovery in High-Dimensional and Nonstationary Time Series

We describe Causal-TS, an open-source Python library for causal discovery in high-dimensional and nonstationary multivariate time series. Causal-TS provides four specialized algorithms-CDNOTS, CDNOTS+, CEDAR, and GRACE-along with wrappers for GES, Granger, LASSO-VAR, and LGES, all sharing a unified conditional independence (CI) test layer with GPU acceleration via PyTorch. A regime discovery pipeline detects structural breaks via pluggable changepoint detectors and runs discovery per regime with regime-specific parameters. A command-line interface, synthetic data generators, and opti

arXiv (cs.LG) · Jul 27, 2026
3 min
Explainable Reinforcement Learning via Physics-Aware Policy DistillationResearch

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the black-box nature of deep neural networks. This lack of transparency poses significant challenges for regulatory compliance and human-agent trust. This paper presents an experimental study aimed at making high-performance continuous control DRL systems interpretable. A policy distillation framework is implemented using the classic Inverted Pendulum benchmark. A high-performance Twin Delayed DDPG (TD3) agent serves as an opaque, continuous

arXiv (cs.LG) · Jul 27, 2026
3 min
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats AccumulatingResearch

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of

arXiv (cs.AI) · Jul 27, 2026
4 min
MMOE: Modernizing Diffusion Transformers with Efficient Expert DesignResearch

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly enlarge total parameter counts and sparsity ratios without importing the efficiency mechanisms that made LLM scaling practical, so generation quality is seldom balanced against training and deployment cost. This raises a natural question: can the architectural principles behind efficient

arXiv (cs.LG) · Jul 27, 2026
4 min
A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facilityResearch

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded evaluation. The retrieval engine fuses dense, sparse, and knowledge-graph (KG) channels wi

arXiv (cs.AI) · Jul 27, 2026
4 min
When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based CorrectionResearch

When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction

Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-matching loss decomposes exactly, with no independence assumption, into an irreducible entropy plus a divergence whose derivative along the training path is positive precisely for structures rare during training and common at deployment, diverging as their training probability goes to zero. Empirically the trade-off is a power law with expon

arXiv (cs.LG) · Jul 27, 2026
4 min
When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based CorrectionResearch

When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction

Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show t...

Papers with Code · Jul 27, 2026
2 min
Kimi K3: Open Frontier IntelligenceResearch

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning ac

arXiv (cs.LG) · Jul 27, 2026
8 min
Reason-Mediated Behavioral Models for Auditing LLM Social SimulatorsResearch

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states $Z$, where positive signs support adoption and negative signs

arXiv (cs.AI) · Jul 27, 2026
3 min
Efficiency Matters in Autonomous ResearchResearch

Efficiency Matters in Autonomous Research

AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as

arXiv (cs.AI) · Jul 27, 2026
4 min
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature EffectsResearch

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce

arXiv (cs.AI) · Jul 27, 2026
4 min
Agentic Permissions Policy Algebra for Taint Confinement in LLM AgentsResearch

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context branching and prospective acquisition enforcement. Before data acquisition occurs, APPA p

arXiv (cs.AI) · Jul 27, 2026
4 min
A Model for Imbalanced Label Aggregation: A Focus on Minority-Class DetectionResearch

A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection

We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest operational importance are also the rarest ones. In this setting, annotators may be reliable on both classes, unreliable on both classes, majority-class specialists, or minority-class specialists. Existing models only partially address this problem: they either capture class-dependent errors but ignore item difficulty, or they model item

arXiv (cs.LG) · Jul 27, 2026
4 min
A Model for Imbalanced Label Aggregation: A Focus on Minority-Class DetectionResearch

A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection

We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively undere...

Papers with Code · Jul 27, 2026
2 min
Attribution and Uncertainty Behavior of Learned Residual Gyro Correction for Gyro-Stellar EstimationResearch

Attribution and Uncertainty Behavior of Learned Residual Gyro Correction for Gyro-Stellar Estimation

This work investigates uncertainty decomposition and explainability in a deep learning-based framework for gyroscope bias correction. A 1-D Convolutional Neural Network is trained to predict residual angular rate corrections from multi-sensor inputs, including gyroscope and star tracker measurements. The bias corrections are sent to a flight-representative Gyro-Stellar Estimator. The network produces both mean corrections and input-dependent (heteroscedastic) aleatoric uncertainty, while epistemic uncertainty is estimated via an ensemble of independently trained models. The proposed

arXiv (cs.LG) · Jul 27, 2026
4 min
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code RepairResearch

Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair

Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified

arXiv (cs.AI) · Jul 27, 2026
4 min
Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewResearch

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of Explainable AI (XAI) in code review and its impact on trust remain underexplored. Objective: We study the influence of XAI on developer trust in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants, comparing three LLM-based code review systems w

arXiv (cs.AI) · Jul 27, 2026
4 min
Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future DirectionsResearch

Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The ri...

Papers with Code · Jul 27, 2026
2 min
Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future DirectionsResearch

Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions

The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation. Given AI's increasing prominence and role within IE, the paper analyzes this new form, examining both AI's unique contributions to IE and its pote

arXiv (cs.AI) · Jul 27, 2026
4 min
SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM AgentsResearch

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent processes required for operational extreme-weather early warning. To bridge this gap, this

arXiv (cs.AI) · Jul 27, 2026
3 min
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language ModelsResearch

D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models

Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score,

arXiv (cs.AI) · Jul 27, 2026
4 min
PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAMResearch

PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM

Large scale Bayesian nonparametrics (BNP) learner such as Stochastic Variational Inference (SVI) can handle datasets with large class number and large training size at fractional cost. Like its predecessor, SVI rely on the assumption of conjugate variational posterior to approximate the true posterior. A more challenging problem is to consider large scale learning on non-conjugate posterior. Recent works in this direction are mostly associated with using Monte Carlo methods for approximating the learner. However, these works are usually demonstrated on non-BNP related task and less c

arXiv (cs.LG) · Jul 27, 2026
4 min
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video UnderstandingResearch

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames an

arXiv (cs.AI) · Jul 27, 2026
3 min
Evaluating Fuzz Testing for Reinforcement Learning AgentsResearch

Evaluating Fuzz Testing for Reinforcement Learning Agents

Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as a promising method for exploring the vast state spaces of RL agents and exposing crashes. Although numerous RL fuzzing methods have been proposed, existing studies often differ in evaluation settings, baselines, and metrics, making it difficult to draw reliable conclusions about their relative effectiveness and practical usefulness. To

arXiv (cs.LG) · Jul 27, 2026
4 min
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in SportsResearch

LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benc

arXiv (cs.AI) · Jul 27, 2026
4 min
The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video GroundingResearch

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal

arXiv (cs.AI) · Jul 27, 2026
4 min
The balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flowsResearch

The balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flows

Model-based active flow control requires predictive models that are accurate, stable, and fast enough for real-time optimisation. In controlled wake flows, this is often achieved through Reduced-Order Models (ROMs) that first compress high-dimensional velocity snapshots into a latent space and then learn a time- stepping predictor for the dynamics in the latent space. Here, we study how the choice of the spatial encoder affects the predictability of the resulting latent coordinates for wake flows under control inputs. Using two actuated 2D wake configurations, a simplified truck wake

arXiv (cs.LG) · Jul 27, 2026
4 min
Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation ClassifierResearch

Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation Classifier

Learned feature reweighting can improve automatic modulation classification (AMC) in software, but the same operation introduces additional arithmetic and latency when implemented on an FPGA. This work measures that trade-off in a compact fixed-point classifier using 24 sparse DFT-energy features, 8 phase/statistical features, and a 32-to-128-to-11 multilayer perceptron. A second architecture inserts a learned 32-element, 8-bit, input-dependent gate before the classifier. Gated and ungated models are trained using post-training quantization (PTQ) and quantization-aware training (QAT)

arXiv (cs.LG) · Jul 27, 2026
4 min
DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic HashingResearch

DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and int

arXiv (cs.AI) · Jul 27, 2026
4 min
DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic HashingResearch

DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing

Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have...

Papers with Code · Jul 27, 2026
2 min
TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge GraphsResearch

TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted. We present TRACE- CTI, a post-extraction claim-governance framework that preserves run-level Predictions, aggregates them into configuration-level GraphAssertions, materializes setup-deduplicated corroboration as ConsensusAssertions, and exposes only GraphAssertions backed by policy-com

arXiv (cs.AI) · Jul 27, 2026
4 min
EgoPlay: Event-Triggered Video Editing for Egocentric StreamsResearch

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on even...

Papers with Code · Jul 27, 2026
2 min
BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient LeakageResearch

BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage

Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split placement can lead to severe privacy leakage through intermediate representations. In this work, we propose a topology-guided framework for privacy-aware split learning based on the persistent Betti complexity of smashed activations. Through comprehensive layer-wise analysis, we show that privacy risk in split learning is highly non-uniform across layers and exhibits sharp transition regions that are not captured by architectural depth alone. In part

arXiv (cs.LG) · Jul 27, 2026
4 min
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context DecodingResearch

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions that a page's own compact basis retains. LOCKS gives every page its own spectral summary (resident, about a tenth the cache's size), reconstructs within-page logits, estimates each page's attention mass by log-sum-exp, and attends only the top pages; selection itself reads no candidate keys or values. Selecting on this summary alone st

arXiv (cs.LG) · Jul 27, 2026
3 min
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac FindingsResearch

EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared--Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality,

arXiv (cs.LG) · Jul 27, 2026
4 min
The K-SCAN Clustering AlgorithmResearch

The K-SCAN Clustering Algorithm

In the Big Data era, the scalability of clustering algorithms constitutes a key challenge. Traditional density-based methods (e.g., DBSCAN) offer robustness to noise and the ability to detect non-linear clusters, yet their quadratic time complexity $O(N^2)$ drastically limits their applicability. Conversely, partitional algorithms (e.g., K-Means), with their linear complexity $O(N)$, impose sphericity on the resulting groups and fail in the presence of outliers. This paper presents K-SCAN -- a novel hybrid algorithm that optimizes this trade-off. The method integrates preliminary vec

arXiv (cs.LG) · Jul 27, 2026
4 min
A Modern ConvNet for Solar Filament DetectionResearch

A Modern ConvNet for Solar Filament Detection

Automated solar filament detection using deep learning faces several challenges. Semantic segmentation of solar filaments is a complicated multiscale feature...

Papers with Code · Jul 27, 2026
2 min
LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation ModelsResearch

LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models

Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for ev...

Papers with Code · Jul 27, 2026
2 min
Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative ControlsResearch

Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls

Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative controls r...

Papers with Code · Jul 27, 2026
2 min
Evaluating RAG for French immigration law: a benchmark and baseline studyResearch

Evaluating RAG for French immigration law: a benchmark and baseline study

International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available b...

Papers with Code · Jul 27, 2026
1 min
Context Is King: How In-Context Specification Shapes the Geometry of ConceptsResearch

Context Is King: How In-Context Specification Shapes the Geometry of Concepts

Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed...

Papers with Code · Jul 27, 2026
2 min
Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian EigenmodesResearch

Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian Eigenmodes

Finite-time driving of stochastic systems generates excess dissipation, causing the evolving probability distribution to lag behind the instantaneous equilib...

Papers with Code · Jul 27, 2026
2 min
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool CallsResearch

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but shoul...

Papers with Code · Jul 27, 2026
2 min
Rethinking the Generation Order of Block Diffusion Language ModelsResearch

Rethinking the Generation Order of Block Diffusion Language Models

Diffusion language models enable flexible arbitrary-order generation, but existing sampling methods are mostly designed for early masked diffusion models (MD...

Papers with Code · Jul 27, 2026
1 min
Catalyst Diffusion Transformer: Generative Inverse Design of Heterogeneous CatalystsResearch

Catalyst Diffusion Transformer: Generative Inverse Design of Heterogeneous Catalysts

The vast chemical design space and complex, interdependent design variables make catalyst discovery for targeted properties highly labor- and resource-intens...

Papers with Code · Jul 27, 2026
1 min
SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth EstimationResearch

SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sens...

Papers with Code · Jul 27, 2026
2 min
LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target DetectionResearch

LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in ...

Papers with Code · Jul 27, 2026
2 min
Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck LocalizationResearch

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may pr...

Papers with Code · Jul 27, 2026
2 min
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learningResearch

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-La...

Papers with Code · Jul 27, 2026
1 min
UniGen-AR: Unifying Visual Generation with Auto-Regressive ModelingResearch

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled by se...

Papers with Code · Jul 27, 2026
2 min
BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible FusionResearch

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sen...

Papers with Code · Jul 27, 2026
2 min
Towards High-Level Semantic IntelligenceResearch

Towards High-Level Semantic Intelligence

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of A...

Papers with Code · Jul 27, 2026
2 min
A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine InterfacesResearch

A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces

Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmentation, and human-...

Papers with Code · Jul 27, 2026
2 min
Structural Loss Metrics for Tensor Approximation via Matrix Low-Rank ApproximationResearch

Structural Loss Metrics for Tensor Approximation via Matrix Low-Rank Approximation

Matricized low-rank approximation via SVD is a standard surrogate for tensor decompositions, but entry-wise reconstruction error fails to capture multiway ge...

Papers with Code · Jul 27, 2026
1 min
Adaptive Data Admission and Retention for Streaming Federated LearningResearch

Adaptive Data Admission and Retention for Streaming Federated Learning

We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selective...

Papers with Code · Jul 27, 2026
2 min
Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AIResearch

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mai...

Papers with Code · Jul 27, 2026
2 min
RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold RepresentationResearch

RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embedded i...

Papers with Code · Jul 27, 2026
1 min
DECAF: De-Clustering for Adaptive Representational UnlearningResearch

DECAF: De-Clustering for Adaptive Representational Unlearning

Machine unlearning, which aims to remove the influence of specific training data from a trained model, is a key requirement for privacy, accountability, and ...

Papers with Code · Jul 27, 2026
1 min
Understanding Tone-Dependent Inference Cost in Large Language ModelsResearch

Understanding Tone-Dependent Inference Cost in Large Language Models

We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed t...

Papers with Code · Jul 27, 2026
1 min
Research Report on Noise-Shaped One-Bit Coefficients in Discrete Polynomial Fourier ExtensionResearch

Research Report on Noise-Shaped One-Bit Coefficients in Discrete Polynomial Fourier Extension

This report studies noise-shaped one-bit coefficients in normalized discrete polynomial Fourier extension. For first-order Sigma-Delta quantization, the erro...

Papers with Code · Jul 26, 2026
1 min
OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion ModelsResearch

OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models

High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative qualit...

Papers with Code · Jul 26, 2026
1 min
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-ImprovementResearch

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale o...

Papers with Code · Jul 26, 2026
2 min
Maximum Satisfiability of Simple Temporal ProblemsResearch

Maximum Satisfiability of Simple Temporal Problems

The Simple Temporal Problem (STP) is a core framework for quantitative temporal constraints. As STP data can be inconsistent, we study MAXSTP: compute a maxi...

Papers with Code · Jul 26, 2026
2 min
Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement LearningResearch

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-G...

Papers with Code · Jul 26, 2026
1 min
PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPsResearch

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progres...

Papers with Code · Jul 24, 2026
1 min
SM4RT: Learning Structured Motion Geometry for 4D ReconstructionResearch

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body t

arXiv (cs.AI) · Jul 24, 2026
4 min
Explainable Reinforcement Learning for assisting Air Traffic ControllersResearch

Explainable Reinforcement Learning for assisting Air Traffic Controllers

To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviation--and to advance toward higher levels of automation and seamless human-AI collaboration--building trust in AI-driven solutions is essential. Trust, in turn, is closely linked to the explainability of AI systems. The rapid advancements in AI across various domains have underscored the challenges of establishing trust, raising increasing interest in AI explainability even more when applied to deep learning. In this context, the present work aims to explore the applica

arXiv (cs.AI) · Jul 24, 2026
4 min
The Regression Tax: Decomposing Why Skills Help and Hurt LLM AgentsResearch

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enoug

arXiv (cs.AI) · Jul 24, 2026
4 min
PinEqualizer: Full Funnel Content Exploration and Debiasing System at PinterestResearch

PinEqualizer: Full Funnel Content Exploration and Debiasing System at Pinterest

In this paper, we propose a new solution for addressing the content cold-start problem in industry-scale search and recommender systems. Compared to prior approaches, we have made the following new contributions: 1) our solution spans the entire multi-stage funnel and generalizes well for both search and recommendation surfaces, 2) our solution reduces bias favoring existing content, allowing more accurate model prediction across content types and reducing short-term tradeoffs associated with high volumes of explicit content exploration, 3) our solution is evaluated with a scalable m

arXiv (cs.LG) · Jul 24, 2026
3 min
Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency SupportResearch

Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support

A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through spectral values and spectral subspaces; however, common coordinate-wise rotation-gate data-encoding unitaries used in most quantum machine learning models do not explicitly construct such a matrix-level representation. We introduce Quantum Spectral Models (QSMs), in which we construct the generator of the data-encoding unitary directly from

arXiv (cs.AI) · Jul 24, 2026
4 min
Dysphagia Risk Stratification in Head and Neck Cancer via Two-Stage PRO-Clinical StackingResearch

Dysphagia Risk Stratification in Head and Neck Cancer via Two-Stage PRO-Clinical Stacking

Dysphagia is a debilitating late effect of head and neck cancer (HNC) treatment, yet timely identification of at-risk patients remains challenging in survivorship care. Definitive assessment relies on videofluoroscopic imaging, as captured by the Dynamic Imaging Grade of Swallowing Toxicity (CTCAE-DIGEST), which, while validated, requires specialized equipment, trained personnel, and significant patient burden, limiting its routine use in surveillance. Patient-reported outcomes (PROs), by contrast, are low-cost, scalable, and easily collected at any clinical encounter, making them an

arXiv (cs.LG) · Jul 24, 2026
4 min
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-ScienceResearch

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which score

arXiv (cs.AI) · Jul 24, 2026
4 min
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal InferenceResearch

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers remain empirically unreliable: they may accept fabricated papers and detect them at rates close to chance (Bad Scientist, 2025). We present CausalForge, a framework for automated theoretical research in causal inference grounded in the Lean proof assistant. CausalForge combines Causalean, a foundational Lean library for causal inference

arXiv (cs.AI) · Jul 24, 2026
4 min
Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimesResearch

Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes

Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks. Predefined features carry a strong bias, since they fix in advance what counts as informative, while deep neural networks and foundation models are hard to interpret and need large amounts of data and compute. We present bag-of-waves, an interpretable framework that learns a small dictionary of recurring EEG waveform templates, called atoms, using shift-invariant k-means without labels. The continuous EEG is t

arXiv (cs.LG) · Jul 24, 2026
4 min
Susceptible Reservoir Architectures for Regime-Conditional Volatility ForecastingResearch

Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting

Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, and its two concrete implementations, based on complex-valued open-chain and periodic reservoirs and regime-conditioned experts to interpret reservoir features across calm, onset, recovery, and persistent-stress states. We also implement open-system $q$-qubit counterparts in Qiskit while retaining a common AR-Ridge anchor and a bounded residu

arXiv (cs.LG) · Jul 24, 2026
3 min
\k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth UpdatingResearch

\k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating

Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardless of their actual contribution to adaptation. This cost is especially prohibitive for large-scale models with billions of parameters and for resource-constrained settings such as edge deployment and on-device fine-tuning. We show for the first time that not all LoRA matrices are equally worth tuning: matrices with smaller condition nu

arXiv (cs.AI) · Jul 24, 2026
4 min
Singular value soft-thresholding via the polar decompositionResearch

Singular value soft-thresholding via the polar decomposition

Singular value soft-thresholding can be computed via a reduction to the matrix polar decomposition, which allows one to exploit GPU-friendly algorithms for computing the polar decomposition. Empirically, there is a significant speed-up on GPUs compared to the standard approach using the SVD. We leave the investigation of robustness to future work, but note that due to the discontinuous nature of the sign function, the reduction to the polar decomposition is likely only suitable for low-accuracy applications.

arXiv (cs.LG) · Jul 24, 2026
3 min
Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient DescentResearch

Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones. Early-stopped negative-shifted gradient descent escapes this constraint. Its filter is smooth at the would-be pole and mixed-sign-capable: above-ridgeless directions form a leading prefix, with

arXiv (cs.LG) · Jul 24, 2026
4 min
MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual ValidationResearch

MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual Validation

Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effectively when only natural-language requirements are available. Although recent work enables automatic test generation, it often overlooks the inherent stochasticity of LLMs, leading to two key defects: faulty tests generate misleading feedback that distorts code optimization, while mixed-quality test cases produce conflicting evaluation signals that hinder reliable code selection. To addre

arXiv (cs.AI) · Jul 24, 2026
4 min
Learning to Prepare Molecular Ground States with Transformer ModelsResearch

Learning to Prepare Molecular Ground States with Transformer Models

Quantum state preparation is a key component of many quantum algorithms. Performing this step efficiently is essential for realizing practical quantum advantage in quantum chemistry applications. Iterative algorithms like ADAPT-VQE can produce shallow ground-state preparation circuits, but become computationally prohibitive for the larger molecules relevant to materials science and pharmaceutical development. Here, we introduce ADAPT-GQE, a generative AI framework that learns to synthesize ground-state preparation circuits for electronic structure calculations. We first use ADAPT-VQE

arXiv (cs.AI) · Jul 24, 2026
4 min
Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver IteratesResearch

Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates

Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve. We therefore study a $k$-neighborhood data collection strategy that augments datasets of converged solutions with intermediate solver iterates, increasing the amount of training data without additional solver runs. To understand the benefits of this approach, we derive a generalization bound based on Rademacher complexity that reveals the role of the $k$-neighborhoods and related parameters. To ach

arXiv (cs.LG) · Jul 24, 2026
3 min
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AIResearch

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-horizon workflows whose quality is determined only by a delayed, task-level outcome. This mismatch prevents per-call routers from correctly attributing feedback to individual routing decisions. Towards mitigating this, we present TRACE-Router, a task-level routing framework that aligns routing with the unit of su

arXiv (cs.AI) · Jul 24, 2026
4 min
Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming EducationResearch

Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education

Generative AI is reshaping programming education, yet educators often infer students' AI-supported learning from classroom observations alone. This experience report presents a trio-ethnography involving two computing educators with different teaching philosophies and one undergraduate computer science student to examine how these interpretations evolve through dialogue. Across three conversations, the educators reflected on students' AI use, discussed changes to programming pedagogy, and revisited their assumptions after engaging with the student's lived experiences. Rather than sim

arXiv (cs.AI) · Jul 24, 2026
3 min
Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretrainingResearch

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from species vocalizations. If the geometry of the embedding space tracks the tree of life, the representation is picking up something deeper than the labels the model was optimized for. We run Mantel tests across two independent radiations. In 32 marine mammal species (1,754 recordings from the Watkins Marine Mammal Sound Database) the foun

arXiv (cs.AI) · Jul 24, 2026
4 min
Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission ArchitectureResearch

Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture

Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping must follow a dynamic least-privilege principle and be treated as a prevention mechanism before a detection one. A credential that does not exist in an agent's context cannot be misused regardless of the agent's reasoning or evasion sophistication. We outline a three-source architecture instantiating this principle: role-based ceilings, a

arXiv (cs.AI) · Jul 24, 2026
4 min
Hyperball May Not Be a Free LunchResearch

Hyperball May Not Be a Free Lunch

For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Starting from the angular displacement between consecutive parameter states, we derive an angular effective learning rate that accounts for the parameter-update angle, parameter norm, and update norm. We also show that the conventional norm-based measure is a special case under parameter-update orthogonality. We then decompose optimizer update

arXiv (cs.AI) · Jul 24, 2026
4 min
OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge IntelligenceResearch

OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence

We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automatically...

Papers with Code · Jul 24, 2026
1 min
Graph-Based Correlation Matrix Generation: A Convex Optimization ApproachResearch

Graph-Based Correlation Matrix Generation: A Convex Optimization Approach

This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures. We propose a novel convex optimization framework in which an initial matrix is projected onto an elliptope under a positive semidefiniteness constraint. Several numerical schemes are implemented and compared. The problem falls within the broader class of matrix completion, where off-diagonal entries corresponding to absent edges are fixed to zero and diagonal entries are fixed to one. Beyond this structural constraint, the approach offers greater fl

arXiv (cs.LG) · Jul 24, 2026
4 min
Robot Learning to Communicate through Projected Visual AbstractionsResearch

Robot Learning to Communicate through Projected Visual Abstractions

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through such projected visual abstractions requires reasoning not only about bodily motion but also about how that motion is transformed into an external representation perceived by an observer. Among these abstractions, shadows provide a particularly compelling example because they emerge directly from the robot's embodiment while remaining visu

arXiv (cs.AI) · Jul 24, 2026
4 min
On the Identifiability of Controlled World ModelsResearch

On the Identifiability of Controlled World Models

Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework for learning such models in representation space. Recent action-conditioned extensions perform promisingly in visual control and latent-space planning, but leave a fundamental question unresolved: when does controlled latent prediction identify both the underlying state and the controlled dynamics? This is challenging under nonlinear obse

arXiv (cs.LG) · Jul 24, 2026
4 min
Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based ExplainabilityResearch

Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability

Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque endtoend tools limiting this kind of material engagement We argue that even large models can function as creative materials when their internal structure is made visible and manipulable To support this we propose a handson approach to explainability centred on experimentation and intervention We instantiate this approach with a model bend

arXiv (cs.AI) · Jul 24, 2026
3 min
PRIMS: Physics-guided Representation for Fluid Identification in Multimodal SensingResearch

PRIMS: Physics-guided Representation for Fluid Identification in Multimodal Sensing

Accurate on-device fluid identification is essential for microfluidic applications, yet maintaining reliability under varying flow, pressure, and temperature remains a key challenge. Existing learning-based methods often treat sensor signals as domain-agnostic features, neglecting the underlying physical relationships that govern fluid behavior, thereby limiting generalization and interpretability. To address this, we propose PRIMS, a physics-aware multimodal Transformer that integrates physical knowledge into representation learning and attention mechanisms through three dedicated m

arXiv (cs.AI) · Jul 24, 2026
4 min
Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based CompositionResearch

Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition

Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision. We present Reflector, an interactive audio workstation that tracks harmonic combinations as they accumulate on the composer's timeline and adapts retrieval as the arrangement develops. The system is organized around a fixed interval-class oracle: a hand-designed table of weights that scores how pitch-class content combines between sources. An encoder train

arXiv (cs.LG) · Jul 24, 2026
4 min
LunarFM: A Shared Multimodal Representation of the Moon's SurfaceResearch

LunarFM: A Shared Multimodal Representation of the Moon's Surface

The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast quantities of orbital remote-sensing data have been collected, scientific analysis and resource mapping remain fragmented by heterogeneous multiinstrument observations, sparse labels, and bespoke task-specific modelling workflows. Here we introduce LunarFM, a multimodal foundation model that learns a general representation of the lunar

arXiv (cs.LG) · Jul 24, 2026
4 min
A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource AllocationResearch

A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation

Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational reliability of these agentic AI systems is fundamentally challenged by the absence of reliable ground truth in open-ended environments and the risk of increasing operational drift over time. To address this challenge, we propose and experimentally evaluate an agentic AI framework, designed to enforce autonomous integrity within LLM-driven systems.

arXiv (cs.AI) · Jul 24, 2026
4 min
Learning Ergodic Dynamical Systems from a Finite TrajectoryResearch

Learning Ergodic Dynamical Systems from a Finite Trajectory

We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system. More precisely, we study discrete-time autonomous stochastic systems defining time-homogeneous Markov processes. We first focus on estimating the optimal one-step prediction function by nonlinear least squares, and derive high-probability guarantees measured with respect to the invariant measure of the process. These results make explicit how the non-independent and non-identically distributed nature of trajectory data modifies the classical statistical learning analysis. We

arXiv (cs.LG) · Jul 24, 2026
3 min
Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation ModelsResearch

Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation Models

Frozen EEG encoders proliferate; per-model fine-tune defaults do not scale. We present Nimbus Personalizer: one contract encode to Bayesian head to BrainState (optional affine mid-tier) that sits on heterogeneous frozen trunks without a new personalization stack per architecture. Thesis (systems): the contribution is the trunk-agnostic API - not LDA-on-embeddings as an ML novelty - so OEMs integrate once and swap trunks. Evidence: the same surface runs on five classical trunks EEGNet, Shallow, Deep, Conformer, ATCNet x four MI datasets (18 cells) and on a foundation encoder (REVE) un

arXiv (cs.LG) · Jul 24, 2026
4 min
SceneActBench: Can Agents Act on the 3D Scenes They See?Research

SceneActBench: Can Agents Act on the 3D Scenes They See?

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geom

arXiv (cs.AI) · Jul 24, 2026
3 min
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM DecodingResearch

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding

With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that exploits KV cache redundancy through hierarchical importance awareness. Algorithmically, HiKV compresses the KV cache at two granularities: Stage I evicts unimportant tokens within a fixed budget, and Stage II further loads only the significant elements of each retained token, reaching compression ratios unattainable at a single granularity

arXiv (cs.AI) · Jul 24, 2026
4 min
Agentic Root Cause Analysis through Evidence-Grounded ReasoningResearch

Agentic Root Cause Analysis through Evidence-Grounded Reasoning

Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing data-driven approaches aim to automate this, two critical limitations restrict their deployment: their operate as black boxes unable to justify their diagnosis, and they require scarce labeled examples of faulty operation. To address this gap, we introduce AgentRCA, a zero-shot agentic framework for evidence-grounded root cause analysi

arXiv (cs.AI) · Jul 24, 2026
3 min
Local-Global Geometric Insights for Graph Neural Networks via Entropic CurvatureResearch

Local-Global Geometric Insights for Graph Neural Networks via Entropic Curvature

Curvature notions on graphs, particularly Ollivier-Ricci and Forman, have emerged as powerful tools for addressing fundamental issues in Graph Neural Networks (GNNs) such as oversmoothing and oversquashing, but rely almost exclusively on local edge-level comparisons and therefore fail to certify how information actually propagates over long distances. We introduce Entropic Curvature, a global, transport-based curvature obtained by extending the Lott-Sturm-Villani framework to graphs through the displacement convexity of entropy along Wasserstein geodesics. We define a tractable Weak

arXiv (cs.LG) · Jul 24, 2026
3 min
IDEAgent: Agentic Quality-Diversity Search for Research Idea GenerationResearch

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives and framed as a Quality-Diversity (QD) search. In line with this perspective, we introduce I

arXiv (cs.AI) · Jul 24, 2026
4 min
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AIResearch

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across bench

arXiv (cs.AI) · Jul 24, 2026
3 min
Interior interpretability with attention rollout: contraction and propagation profiles in TransformersResearch

Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator encoding attention-mediated propagation between feature tokens. By applying classical Doeblin--Dobrushin contraction theory, we show that a rollout o

arXiv (cs.AI) · Jul 24, 2026
4 min
Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal ReasoningResearch

Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning

High-complexity operational environments require methods that detect and anticipate temporally distributed patterns rather than classify isolated events. This paper introduces TRACTA (Temporal Reasoning and Capability-Trajectory Analysis), a controlled synthetic benchmark for temporal structural reasoning in high-complexity event-driven systems, instantiated through Multi-Domain Operations (MDO)-like scenarios. The benchmark includes three tasks: early_warning, pattern_detection, and run_classification, and compares raw-event neural models, a contract-lite semantic baseline, and a ne

arXiv (cs.AI) · Jul 24, 2026
3 min
Indexing: the Beginning and the EndResearch

Indexing: the Beginning and the End

We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits and one integer $i$ from $1$ to $n$ called the index, and the output equals the value of the $i$-th bit. We introduce causal complexity for masked architectures. We show that architectures with low causal complexity cannot solve the indexing primitive in any constant number of layers when the index appears at the end of the input. In pa

arXiv (cs.LG) · Jul 24, 2026
4 min
Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice CloningResearch

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinic...

Papers with Code · Jul 24, 2026
1 min
3D-Aware VLMs with Implicit and Explicit GeometriesResearch

3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode det

arXiv (cs.AI) · Jul 23, 2026
3 min
Expanding Flow MapsResearch

Expanding Flow Maps

Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which define flows between distributions of increasing dimensionality along an expanding interpolant that grows the state by augmenting it with conditional noise. Building on this construction, we propose Expanding Flow Maps (EFMs), a new class of flow maps that distill the expanding interpolant

arXiv (cs.LG) · Jul 23, 2026
3 min
GraphVid: Interactive Graph-Controllable Video GenerationResearch

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structu

arXiv (cs.AI) · Jul 23, 2026
3 min
Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$Research

Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$

Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorly understood. In particular, a central unresolved question is whether BB converges superlinearly for almost every strictly convex quadratic problem and initialization. We provide a negative answer to this question. Specifically, for every finite dimension $n\geq4$, we construct a nonempty open, hence positive-Lebesgue-measure, family of strictly convex quadratic problems and initial points for which the long Barzilai--Borwein method (BB1) converge

arXiv (cs.AI) · Jul 23, 2026
4 min
Synthetic data generation framework for quality control automation in gravure printingResearch

Synthetic data generation framework for quality control automation in gravure printing

Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards in rotogravure printing. Deep learning models give prospects for automation. However, training robust deep learning models, such as YOLO or Vision Transformers, is heavily hindered by the extreme scarcity of real-world industrial defects images. To overcome this limitation, this paper introduces a novel synthetic data generation framework tailored for rotogravure printi

arXiv (cs.AI) · Jul 23, 2026
4 min
Beyond Sufficiency: Time Series Explanation with Counterfactual NecessityResearch

Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that support the prediction without being essential to the model's decision. We introduce \textbf{TimePNS}, a necessity-aware framework for time-series explanation. Inspired by Pearl's counterfactual notion of necessity, TimePNS assesses whether a temporal factor is necessary by intervening on it and

arXiv (cs.AI) · Jul 23, 2026
3 min
Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble ModelingResearch

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordin

arXiv (cs.LG) · Jul 23, 2026
3 min
Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in GhanaResearch

Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patterns. Anomalies were highly structured in space and time. Ashanti and Northern Regions accounted for most recurrent anomalies, with persistent hotspots at Tamale, Kumasi, and Accra. A key finding was the spatial distinction between anomaly burden (cumulative cases during anomalous periods) and anomaly frequency (persistence of unusual behaviour). Tamale had the highest burden during anomalies, whereas the highest anomaly rates clustered

arXiv (cs.AI) · Jul 23, 2026
3 min
Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral ReasoningResearch

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the m

arXiv (cs.AI) · Jul 23, 2026
3 min
OpenForgeRL: Train Harness-native Agents in Any EnvironmentResearch

OpenForgeRL: Train Harness-native Agents in Any Environment

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls w

arXiv (cs.AI) · Jul 23, 2026
4 min
Visual Contrastive Self-DistillationResearch

Visual Contrastive Self-Distillation

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content re

arXiv (cs.AI) · Jul 23, 2026
4 min
MIRROR: Learning from the Other View for Multi-Modal ReasoningResearch

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To s

arXiv (cs.AI) · Jul 23, 2026
3 min
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy AlignmentResearch

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using ma

arXiv (cs.LG) · Jul 23, 2026
3 min
Neural solutions of coupled ghost and gluon Dyson--Schwinger equations in Landau gaugeResearch

Neural solutions of coupled ghost and gluon Dyson--Schwinger equations in Landau gauge

The coupled ghost and gluon Dyson--Schwinger equations (DSEs) of four-dimensional Landau-gauge Yang--Mills (YM) theory are solved with a neural representation trained only from renormalized equation residuals. The neural and fixed-point solutions agree at the percent level and remain stable under changes of initialization, network size, integration grid, and infrared boundary condition. Variations of the three-gluon vertex model produce substantially larger effects than the neural error. The MiniMOM ultraviolet running and the sign change of the gluon Schwinger function are also repr

arXiv (cs.LG) · Jul 23, 2026
3 min
The Boundaries of Automation: A Theory of Persistent Human ParticipationResearch

The Boundaries of Automation: A Theory of Persistent Human Participation

The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper challenges that assumption. Rather than asking how far automation can extend, we ask where its conceptual limits lie and argue that human participation may persist even with highly capable AI systems for three distinct reasons. Technical or complementarity grounds arise when humans contribut

arXiv (cs.AI) · Jul 23, 2026
4 min
UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical ImagingResearch

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, ye...

Papers with Code · Jul 23, 2026
1 min
Zero-Flow Two-Sample TestsResearch

Zero-Flow Two-Sample Tests

We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow criterion, termed zero-flow discrepancy (ZFD). We prove the validity of ZFD and propose a practical testing procedure, termed the zero-flow two-sample test (ZF2ST). The key idea is to learn how samples from the two distributions are locally misaligned and use the resulting directional pattern as evidence of distributional difference. By separating witness learning from hypothesis evaluation

arXiv (cs.LG) · Jul 23, 2026
3 min
DONDO: Open w2v-BERT Speech-Recognition Base Models for African LanguagesResearch

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 sel...

Papers with Code · Jul 23, 2026
2 min
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextResearch

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with context and comes to dominate the draft cost -- precisely where speculation is most valuable. The effect compounds with draft length (a deep nat

arXiv (cs.LG) · Jul 23, 2026
4 min
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIsResearch

From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs

Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large language models can synthesize executable Rust tests, but their outputs often violate API preconditions, remain shallow, or reduce concurrency to accidental sequential traces. Conversely, model-based and systematic testing techniques provide semantic control but commonly require substantial handwritten code to turn abstract scenarios into executable tests. This paper addresses the gap between formal scenario design and low-cost test concretization

arXiv (cs.AI) · Jul 23, 2026
3 min
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIsResearch

From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs

Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large language models ca...

Papers with Code · Jul 23, 2026
1 min
ElasticTTT: Prior-Preserving Test-Time Tuning for Video EditingResearch

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framewo

arXiv (cs.AI) · Jul 23, 2026
3 min
GS-Agent: Creating 4D Physical Worlds With Generative SimulationResearch

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an age

arXiv (cs.AI) · Jul 23, 2026
4 min
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent MediationResearch

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipu

arXiv (cs.AI) · Jul 23, 2026
3 min
Improved lower bounds for the Shannon capacity of odd cyclesResearch

Improved lower bounds for the Shannon capacity of odd cycles

The Shannon capacity $Θ(G)$ of a graph $G$ quantifies the maximum rate at which information can be transmitted with zero error over a noisy channel. It is lower bounded by $α(G^d)^{1/d}$ for any $d$, where $α(G^d)$ is the independence number of the $d$-th strong power of $G$. We construct independent sets of size $134753$ in $C_7^{10}$, $21909$ in $C_{11}^{6}$, and $62530$ in $C_{13}^{6}$, improving the best known lower bounds for the Shannon capacity of these graphs to $Θ(C_7)\geq 134753^{1/10}>3.258020$, $Θ(C_{11})\geq 21909^{1/6}>5.289773$, and $Θ(C_{13})\geq 62530^{1/6}>6.300109$

arXiv (cs.AI) · Jul 23, 2026
3 min
Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion ModelResearch

Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion Model

Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore texture maps and focus on natural images. An eff...

Papers with Code · Jul 23, 2026
2 min
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture ProblemsResearch

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a s

arXiv (cs.AI) · Jul 23, 2026
4 min
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture ProblemsResearch

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context...

Papers with Code · Jul 23, 2026
2 min
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate itResearch

Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation». This essay argues that the overuse is a trained disposition, driven mainly by a training distribution rich in promotional prose and by preference tuning (RLHF) that rewards confident, emphatic phrasing; the left-to-right nature of generation is an amplifier rather than the root cause. Building on evidence that models diverge from hum

arXiv (cs.AI) · Jul 23, 2026
4 min
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language ModelsResearch

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated with cognitive decline. Recent advances in large language models (LLMs) further strengthen the potential of speech-based assessment by enabling more expressive representation learning and improved generalization across diverse speakers, recording devices, and clinical envi

arXiv (cs.LG) · Jul 23, 2026
4 min
Toward Continuous Assurance for the Democratization of AI Agent Creation in IndustryResearch

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments. This democratization enables rapid local innovation, but it also creates a reliability gap: agents that appear to users as simple productivity artifacts may depend on changing models, tools, retrieval sources, permissions, prompts, schedules, and external services. These dependencies can cause silent degradation long after deployment, even when no user directly modifies the agent. This paper identifies the reliability challenge creat

arXiv (cs.AI) · Jul 23, 2026
3 min
What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model RepresentationsResearch

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-circuit extraction method to a 2x2 design -- Python and Rust crossed with Qwen2.5-Coder-7B and DeepSeek-Coder-V1-6.7B -- and measuring a complete inventory of grammatical concepts (58 Python, 57 Rust) identically in all four cells: the smallest design that separates what depends on the task, the language, and the model. The answer splits into three parts. What earns dedicated circuitry is set by the task: the models agree on which conc

arXiv (cs.LG) · Jul 23, 2026
4 min
Compact Latent Coordination for Autonomous Vehicles at Unsignalized IntersectionsResearch

Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections

Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs. We propose Master-Agent Proto-plan System (MAPS), a hierarchical deep reinforcement learning (DRL) architecture in which a centralized Master agent generates a compact, continuous embedding, denoted as proto-plan, that encodes a global coordination strategy. Decentralized Worker agents integrate this embedding with local o

arXiv (cs.AI) · Jul 23, 2026
3 min
Recurrent Sinusoidal INRs for Efficient High-Fidelity RepresentationResearch

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals tha...

Papers with Code · Jul 23, 2026
1 min
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasksResearch

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudina

arXiv (cs.AI) · Jul 23, 2026
4 min
Finite-Sample Coverage Audits for High-Recall Candidate Generation: Certification and Learning-Theoretic DesignResearch

Finite-Sample Coverage Audits for High-Recall Candidate Generation: Certification and Learning-Theoretic Design

An initial high-recall stage in an empirical pipeline decides which items pass to later review, labelling, or modelling, and relevant items it misses are lost to every subsequent stage. We study how many audit labels are needed to certify, with finite-sample validity, that this missed relevant mass is small, and our main results characterise the label complexity of this problem. We first show that no procedure using only labels from inside the candidate set can certify any non-trivial bound on the missed mass: the audit must sample the excluded pool, the only region where unrecovered

arXiv (cs.LG) · Jul 23, 2026
4 min
Error Certificates for KV-Cache Eviction via Randomized DesignResearch

Error Certificates for KV-Cache Eviction via Randomized Design

Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is unchanged while the true attention-output error grows arbitrarily, so no serving-time estimator of that error is consistent. Randomized eviction restores identifiability. With a Poisson-sampled tail at known inclusion probabilities, one logit offset performs the Hájek correction inside the softmax, and a survey-sampling variance estimator over the

arXiv (cs.AI) · Jul 23, 2026
3 min
Future Rendering neq Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowResearch

Future Rendering neq Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window

Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction, and ...

Papers with Code · Jul 23, 2026
2 min
Thinkink: 2D Spatial Ink-native Interaction with LLMsResearch

Thinkink: 2D Spatial Ink-native Interaction with LLMs

People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language models (LLMs) into this practice, we propose Thinkink. Prompts can be handwritten text or drawn sketches with LLM-generated responses visualized as ink-like text and sketches spatially integrated into a shared canvas. A semantic tree streamlines ink interpretation, and a lightweight UI provides explicit control using a state machine. The tool was designed using a three-stage process. A formative study (N=12) examined current practices with conventional and digital inking met

arXiv (cs.AI) · Jul 23, 2026
3 min
AREX: Towards a Recursively Self-Improving Agent for Deep ResearchResearch

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates betw

arXiv (cs.AI) · Jul 23, 2026
4 min
Detecting LLM-Generated Tokens in Human--LLM Coauthored TextResearch

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content...

Papers with Code · Jul 23, 2026
1 min
Detecting LLM-Generated Tokens in Human--LLM Coauthored TextResearch

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs. This paper introduces a new method to address this urgent need. Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores. The key idea is to smooth adjacent to

arXiv (cs.AI) · Jul 23, 2026
3 min
Test-Time Scaling via Error LocalizationResearch

Test-Time Scaling via Error Localization

Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standard approaches such as independent sampling and sequential multi-turn refinement operate without token-level credit assignment, resulting in computational inefficiency, since valid reasoning prefixes are frequently discarded. In this work, we introduce Test-Time Scaling via Error Localization (TTEL), an inference-time algorithm that utilizes fixed or environment feedback to perform token-level error localizat

arXiv (cs.LG) · Jul 23, 2026
3 min
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion TransformersResearch

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applies an invertible linear transform to the activations and its inverse to the weights before quantizing both. Normalization layers between blocks force this transform to run online at every denoising step, making its inference computation cost the binding design constraint. Existing options trade quantization quality for inference cost: per-channel scaling (S

arXiv (cs.LG) · Jul 23, 2026
4 min
When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMsResearch

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, w...

Papers with Code · Jul 23, 2026
1 min
Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobilityResearch

Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility

Reliable electric vehicle (EV) charging infrastructure is a cornerstone of sustainable, low-carbon cities, yet urban climate stress such as extreme heat, heavy precipitation, and humidity increasingly raises equipment fault risk and undermines the resilience of urban energy and mobility services. Shifting operation from reactive repair to preventive maintenance depends on accurate, forward-looking fault-risk prediction, a task complicated by the heterogeneous time scales of physical, behavioral, contextual, and historical signals and by forecasting over a multi-week horizon. We devel

arXiv (cs.LG) · Jul 23, 2026
4 min
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought ModelsResearch

Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged). We characterize this phenomenon empirically, showing that converged generations achieve 90.3% accuracy on AIME 1983-2024 while non-converged ones achieve only 6.6%, with an overall convergence rate of 62.0%. We then ask whether this outcome is detectable early in the thinking chain using internal model representations. Training linear probes on hidden-sta

arXiv (cs.LG) · Jul 23, 2026
3 min
Context-weighted Discrete Flow MatchingResearch

Context-weighted Discrete Flow Matching

Discrete flow matching provides a flexible framework for generative modeling on discrete structures. However, the standard factorized training objective exposes the model to targets of varying difficulty, mixing well-conditioned, predictable tokens with ambiguous, high-entropy ones. We empirically demonstrate that the uncertainty over the value of each token is closely related to the density of available context in its neighborhood. Motivated by this observation, we propose a simple modification to the underlying continuous-time Markov chain (CTMC) that incorporates local context inf

arXiv (cs.LG) · Jul 23, 2026
3 min
Semantic-Aware Task Clustering for Constructive and Cooperative Multi-TaskingResearch

Semantic-Aware Task Clustering for Constructive and Cooperative Multi-Tasking

Cooperative multi-task semantic communication (CMT-SemCom) improves task execution performance by leveraging shared representations. However, as we demonstrated in [1], cooperative multi-tasking can be either constructive or destructive, depending on the semantic relationships among tasks. To ensure constructive cooperation, we propose a semantic-aware task clustering method for CMT-SemCom. We have formulated a sequential multi-stage optimization problem in which semantically aligned tasks are clustered once after a short initial training phase, and then end-to-end (E2E) joint traini

arXiv (cs.LG) · Jul 23, 2026
3 min
Cautious optimism for deep parameterized quantum circuitsResearch

Cautious optimism for deep parameterized quantum circuits

A central challenge in quantum machine learning is understanding the scaling behavior of parameterized quantum circuits (PQCs). In particular, it remains unclear how their performance on unseen data changes as the number of trainable parameters increases. Prior works have derived formal generalization guarantees for quantum models, but it is well-known that many such results do not fully characterize generalization behavior in practice. In this work, we show that gradient-based PQCs can exhibit improved performance on unseen data as model size increases, displaying the phenomenon of

arXiv (cs.LG) · Jul 23, 2026
4 min
A Diffusion-Model Subpopulation Digital Twin for Mobile Health Deployment: A Case Study on the HeartSteps InterventionResearch

A Diffusion-Model Subpopulation Digital Twin for Mobile Health Deployment: A Case Study on the HeartSteps Intervention

Mobile-health interventions increasingly use online learning and decision making algorithms to personalize when to nudge users toward healthier behavior, but a poorly designed algorithm can burden and disengage participants. New algorithm design decisions should therefore be vetted against realistic simulated users before each real-life deployment. We propose a method to develop ``JITAI-Twins'': digital twins of a target subpopulation for comparing candidate online algorithms before a just-in-time adaptive intervention (JITAI) deployment. The method builds on a conditional time-serie

arXiv (cs.LG) · Jul 23, 2026
4 min
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic SegmentationResearch

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised model DIN...

Papers with Code · Jul 23, 2026
1 min
How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-TuningResearch

How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning

A LoRA adapter is a few megabytes that almost everyone treats as a skill rather than a record of the data behind it. We put that assumption on a scale. Exten...

Papers with Code · Jul 23, 2026
2 min
Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal FeaturesResearch

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting...

Papers with Code · Jul 23, 2026
2 min
M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging DataResearch

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease u...

Papers with Code · Jul 23, 2026
1 min
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu MandarinResearch

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties...

Papers with Code · Jul 23, 2026
1 min
PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided EditingResearch

PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing

Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhoue...

Papers with Code · Jul 23, 2026
2 min
Filter Learning for Subgraphs: Algebras and Performance Risk BoundsResearch

Filter Learning for Subgraphs: Algebras and Performance Risk Bounds

Graph signal processing tasks that leverage spectral information typically assume access to the complete graph topology, which is often unavailable in practi...

Papers with Code · Jul 23, 2026
1 min
How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic ProgrammingResearch

How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming

Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusio...

Papers with Code · Jul 23, 2026
1 min
A New Well-Supported Semantics for Description Logic ProgramsResearch

A New Well-Supported Semantics for Description Logic Programs

Description logic programs are a powerful formalism for combining rules with ontologies. The well-supported semantics for description logic programs ensures ...

Papers with Code · Jul 23, 2026
1 min
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQAResearch

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring exte...

Papers with Code · Jul 23, 2026
2 min
Do Pathology Vision-Language Models Truly See Pathology?Research

Do Pathology Vision-Language Models Truly See Pathology?

Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. However, ...

Papers with Code · Jul 23, 2026
2 min
Achieving Text-based Person Retrieval with Any GranularityResearch

Achieving Text-based Person Retrieval with Any Granularity

Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This paper ...

Papers with Code · Jul 23, 2026
2 min
WAT3R: Feedforward Underwater 3D ReconstructionResearch

WAT3R: Feedforward Underwater 3D Reconstruction

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and di...

Papers with Code · Jul 23, 2026
1 min
Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence MemoryResearch

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active ove...

Papers with Code · Jul 23, 2026
2 min
Latent Variable-Mediated Cross-Learning for Few-Shot Acoustic Impedance ImagingResearch

Latent Variable-Mediated Cross-Learning for Few-Shot Acoustic Impedance Imaging

Acoustic impedance imaging is a fundamental yet severely ill-posed problem in subsurface analysis: the seismic wavelet is unknown, observations are band-limi...

Papers with Code · Jul 23, 2026
1 min
Best-of-Evidence: Best-of-N Selection under Partial VerificationResearch

Best-of-Evidence: Best-of-N Selection under Partial Verification

BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reli...

Papers with Code · Jul 23, 2026
2 min
MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-TransferResearch

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-base...

Papers with Code · Jul 23, 2026
2 min
Three-Pronged Spectral Control for Federated Parameter Efficient Fine TuningResearch

Three-Pronged Spectral Control for Federated Parameter Efficient Fine Tuning

Federated parameter-efficient fine-tuning (PEFT) enables communication-efficient adaptation of large pretrained models on decentralized edge data, but it rem...

Papers with Code · Jul 23, 2026
1 min
Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel GenerationResearch

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code genera...

Papers with Code · Jul 23, 2026
2 min
Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk VideosResearch

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube s...

Papers with Code · Jul 23, 2026
2 min
Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic LogicResearch

Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic Logic

Goal reasoning in Non-Axiomatic Logic (NAL) explains how an adaptive system derives means for realizing desired events under insufficient knowledge and resou...

Papers with Code · Jul 23, 2026
1 min
Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive LearningResearch

Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learning h...

Papers with Code · Jul 23, 2026
1 min
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?Research

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abiliti...

Papers with Code · Jul 23, 2026
2 min
Probabilistic Residual Learning for Online RecommendationsResearch

Probabilistic Residual Learning for Online Recommendations

Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, th...

Papers with Code · Jul 23, 2026
1 min
Code Monitor Red Teaming for Public-Test-Passing CodeResearch

Code Monitor Red Teaming for Public-Test-Passing Code

Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring pr...

Papers with Code · Jul 23, 2026
1 min
REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought ReasoningResearch

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context when ev...

Papers with Code · Jul 23, 2026
2 min
Lipschitzian SLLNs for random functionsResearch

Lipschitzian SLLNs for random functions

We prove strong laws of large numbers for locally Lipschitz functions in the Lipschitz pseudometric. Our results hold under either a topological or a model-theoretic condition, with the latter encompassing functions jointly definable in o-minimal structures but extending substantially beyond this class. Applications include uniform convergence of limiting and Clarke subdifferentials and finite-sample identification of solutions. Consequently, we identify broad classes of functions for which the failure phenomena revealed by our previous negative results [Tian and Royset, arXiv:2511.1

arXiv (cs.LG) · Jul 22, 2026
3 min
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual DataResearch

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data

In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete interface between perception and deduction. We present a neuro-soft-symbolic architecture for differentiable deductive reasoning over latent perceptual facts and knowledge-provided predicates. SoftReason removes the gradient gap by representing the deductive state as a lo

arXiv (cs.AI) · Jul 22, 2026
3 min
Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement LearningResearch

Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning

Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control. As a result, a single remote operator can see, manipulate, and navigate about a real, distant physical environment. This powerful control stack is often relegated to expensive full-sized robots, many of which are inaccessible to the research community. Miniature humano

arXiv (cs.LG) · Jul 22, 2026
4 min
Persian Pixel: A large-scale synthetic OCR dataset for Persian languageResearch

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Computer Science > Computer Vision and Pattern Recognition

arXiv (cs.AI) · Jul 22, 2026
4 min
FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow OptimizationResearch

FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization

Clinical biomarker workflows in translational research settings often rely on spreadsheet-driven tracking, manual quality control (QC) reconciliation, and loosely integrated systems, resulting in limited state visibility, delayed reporting, and increased operational risk. These challenges are particularly pronounced in multi-day assays such as Luminex-based quantification of Fragile X Messenger Ribonucleoprotein (FMRP), where HIPAA-compliant data governance, deterministic workflow progression, and coordinated communication across laboratory and clinical teams are required. This paper

arXiv (cs.AI) · Jul 22, 2026
4 min
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation ExplanationsResearch

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconstruction-dependent, so the score tracks gist, not specific facts. Under exact synthetic gro

arXiv (cs.AI) · Jul 22, 2026
4 min
PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEsResearch

PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs

Physics-informed learning of partial differential equations (PDEs) has been dominated by multilayer perceptrons (MLPs), whose spectral bias and dense parameterization limit both accuracy and interpretability. Kolmogorov Arnold Networks (KANs) mitigate these limitations because their learnable spline activations are structurally aligned with the piecewise-polynomial bases of classical discretizations. However, the way a PDE is cast into a loss functional is as decisive as the choice of approximator: strong-form residual minimization requires high-order derivatives and heavily weighted

arXiv (cs.LG) · Jul 22, 2026
4 min
Statevector-Referenced Geometry Survival of a Four-Qubit ZZ Quantum Kernel on IBM Quantum Hardware: A Fixed-Subset Diagnostic Across Three Execution ConfigurationsResearch

Statevector-Referenced Geometry Survival of a Four-Qubit ZZ Quantum Kernel on IBM Quantum Hardware: A Fixed-Subset Diagnostic Across Three Execution Configurations

Quantum-kernel methods encode a dataset's geometry in a Gram matrix, so learning claims on hardware kernels assume the intended geometry survives execution. We measure that survival for one frozen four-qubit ZZ feature-map kernel on $N=24$ real indoor air-quality windows, reconstructed on ibm_fez (1024 shots per circuit) under baseline, dynamical decoupling alone, and gate twirling alone, each a single non-interleaved job. Every configuration returned a complete, finite, positive-semidefinite Gram matrix and preserved the centered statevector geometry to a substantial but incomplete

arXiv (cs.LG) · Jul 22, 2026
4 min
Online Variance Reduction for Domain Adaptation on Streaming DataResearch

Online Variance Reduction for Domain Adaptation on Streaming Data

This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been proposed, these are incompatible with online, distributed, or incremental learning settings. This paper presents Adaptive vaRiance Reduction via Online reWeighting (ARROW), the first online SVR algorithm for the MMD and CORAL for streamed data. The method maintains moving average references of the alignment statistics, and adaptively reweights incoming minibatch

arXiv (cs.LG) · Jul 22, 2026
3 min
Variance-reduced Domain Adaptation using Paired SamplingResearch

Variance-reduced Domain Adaptation using Paired Sampling

Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their effectiveness in minibatch optimisation settings. Furthermore, the losses lack finite-sum structure, which renders them incompatible with classical stochastic variance reduction (SVR) methods. This paper proposes Paired Sampling for Domain Adaptation (PSDA), a novel SVR technique tailored to such objectives. PSDA pairs observations both within and across domains, t

arXiv (cs.LG) · Jul 22, 2026
3 min
Test-Time Training for Modality Order Consistency in Vision-Language ModelsResearch

Test-Time Training for Modality Order Consistency in Vision-Language Models

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Acro...

Papers with Code · Jul 22, 2026
1 min
Generative AI floods and dilutes the market for booksResearch

Generative AI floods and dilutes the market for books

Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-text AI detection across 14,419 self-published genre-fiction books sold on Amazon from 2023 to 2026, matched to daily sales records through June 2026. None of these books disclose whether or not they contain AI-produced content. We find that books for which we detected substantial AI text ($>$ 25\%) make up a large share of the catalog but a smalle

arXiv (cs.AI) · Jul 22, 2026
4 min
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail HumanoidsResearch

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alig

arXiv (cs.AI) · Jul 22, 2026
4 min
Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modelingResearch

Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modeling

Constitutive modeling under uncertainty remains a central challenge for reliable mechanics simulations, particularly when the available stress-deformation data are sparse, noisy, or heterogeneous. We propose interval and fuzzy physics-augmented neural networks (iPANNs and fPANNs) for uncertainty-aware hyperelastic constitutive modeling. iPANNs learn sparse lower, mean, and upper free energy density branches whose stresses, obtained by automatic differentiation, ultimately enclose noisy stress observations. In contrast to this deterministic interval description, fPANNs embed the learn

arXiv (cs.LG) · Jul 22, 2026
4 min
Understanding Generative AI-mediated User Engagement with Academic Library ResourcesResearch

Understanding Generative AI-mediated User Engagement with Academic Library Resources

This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources. Utilizing web analytics from August 2023 to October 2025, the research identifies a significant increase in AI-mediated traffic, particularly following the integration of linked citation features. Referral analysis identified ChatGPT, Perplexity, and Gemini as the primary platforms driving this traffic. A substantial portion of users reached the institutional repository, primarily accessing electronic theses and dissertations. This pattern suggests that AI retrieval mechanism

arXiv (cs.AI) · Jul 22, 2026
3 min
Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition DropoutResearch

Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout

RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice, failures or occlusions of surveillance sensors often remove one modality. Although RGB or depth alone can contain sufficient cues, models trained only on full-modality inputs fail to exploit the remaining modality once one is missing, causing severe degradation. We tackle this issue with a simple continued-training paradigm, \emph{Condition Dropout (ConD)}, which mitigates degradation while preserving full-modality accuracy. Starting from a pretra

arXiv (cs.AI) · Jul 22, 2026
3 min
Multi-modal transformer for signal classification in nanopore blockade experimentsResearch

Multi-modal transformer for signal classification in nanopore blockade experiments

Nanopore devices have emerged as powerful tools for single-molecule sensing, with potential for rapid, portable diagnostics. They detect changes in ionic current as analytes enter nanometer-scale pores, providing a means of identifying diverse biomarkers from their characteristic signal patterns. However, these signals are highly complex, and reliably assigning them to specific molecules remains a major challenge. Here, we address this by introducing a multi-modal deep learning architecture that jointly processes multiple signal representations, including raw time-series data, wavele

arXiv (cs.LG) · Jul 22, 2026
3 min
Label-Free Finite-Volume-Residual Training of Attention Graph Neural Networks for Coupled Thermo-Fluid FieldsResearch

Label-Free Finite-Volume-Residual Training of Attention Graph Neural Networks for Coupled Thermo-Fluid Fields

Neural surrogates are widely used in scientific machine learning for fast prediction of three-dimensional (3D) thermo-fluid fields. However, generating training data using conventional numerical solvers often incurs substantial computational and storage costs. We propose to train an attention graph neural network by minimizing the finite-volume method (FVM) residuals of the governing equations. These residuals are evaluated directly on the mesh, requiring no labeled data. We evaluate the trained surrogates against computational fluid dynamics (CFD) references and a data-supervised ba

arXiv (cs.LG) · Jul 22, 2026
3 min
Don't Trust the Label: License Laundering in AI Supply ChainsResearch

Don't Trust the Label: License Laundering in AI Supply Chains

AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and applications on GitHub. While each artifact carries a license whose obligations should propagate through redistribution, no study has yet measured whether those obligations survive the chain or are stripped and replaced as artifacts move downstream. We trace 232,270 dataset$\rightarrow$model$\rightarrow$application chains and quantify two forms of license laundering: when artifacts with no declared license acquire definitive labels downstream, and when one declared license catego

arXiv (cs.AI) · Jul 22, 2026
3 min
Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared EnvironmentsResearch

Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments

We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequence. Standard task planners, lacking foresight of future tasks and inconsiderate of others' constraints, solve each task in isolation, leaving terminal states that increase future cost for all, side effects that compound over lengthy task sequences. To reduce cost over the sequence, a robot must anticipate how its actions now may impact performance on future tasks for all robots sharing the environment. Therefore, we present courteous anticipator

arXiv (cs.AI) · Jul 22, 2026
4 min
Sound Probabilistic Safety Bounds for Large Language ModelsResearch

Sound Probabilistic Safety Bounds for Large Language Models

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of us

arXiv (cs.AI) · Jul 22, 2026
3 min
Self-supervision drives representational convergence in medical foundation models more than clinical supervisionResearch

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence is real, what produces it, and whether it is clinically usable are untested, and the similarity measures behind such claims are fragile. We present a controlled dissection across 18 image and 7 text encoders, all open-weight and run locally, spanning 7M to 27B parameters and five imaging modalities, including 650,982 chest radiographs from six datasets. To iso

arXiv (cs.AI) · Jul 22, 2026
4 min
Which Values Do LLMs Confuse? A Schwartz-Based Recognition StudyResearch

Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study

Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expresse...

Papers with Code · Jul 22, 2026
1 min
PoTRE: Test-Time Reasoning inspired by Cognitive HeterogeneityResearch

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Ag

arXiv (cs.AI) · Jul 22, 2026
3 min
The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language ModelsResearch

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation. MI is computed from differences in DepthRank scores between masked and unmasked templates, providing a principled measure of objective-

arXiv (cs.AI) · Jul 22, 2026
3 min
How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing InspectionResearch

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated...

Papers with Code · Jul 22, 2026
2 min
The Ethics of Autonomous AI Agents for Offensive SecurityResearch

The Ethics of Autonomous AI Agents for Offensive Security

LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and operated by trained practitioners -- agentic security tools exhibit \textit{indeterminacy} along three independent dimensions. First, their actions are drawn from a non-deterministic policy whose outputs resist both ex-ante and ex-post explanation, frustrating incident attribution and pre-deployment safety review. Second, their impact is open-ended due to the non-deterministic actions, agency of utilized models, and opaque LLM supply-cha

arXiv (cs.AI) · Jul 22, 2026
4 min
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching RenderingResearch

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semant

arXiv (cs.AI) · Jul 22, 2026
4 min
Exposure is Optional: Learning Unlike Coordination in Language ModelsResearch

Exposure is Optional: Learning Unlike Coordination in Language Models

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A com...

Papers with Code · Jul 22, 2026
2 min
STeMP: Spatio-Temporal Modelling ProtocolResearch

STeMP: Spatio-Temporal Modelling Protocol

Spatio-temporal machine-learning modelling is an important tool in environmental research. However, machine-learning models are highly sensitive to both the ...

Papers with Code · Jul 22, 2026
2 min
On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural LensResearch

On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens

Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to achieve human-like quality in many scenarios, their ability to handle culturally loaded expressions remains underexplored. In this study, we systematically investigate the challenges posed by culturally loaded translation in LLM-based MT systems. We construct a Chinese-Japanese bilingual dataset from the culturally representative corpus Dream of t

arXiv (cs.AI) · Jul 22, 2026
3 min
DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial ProblemsResearch

DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial Problems

While combinatorial optimization problems are central to many scientific and engineering applications, their solution remains challenging due to exponentially large search spaces. Variational quantum algorithms offer a promising route for tackling such problems, yet their practical performance is limited by repeated quantum circuit evaluations and classical parameter updates. In this work, we introduce DQAOA-GPT, a hybrid framework that integrates the distributed quantum approximate optimization algorithm (DQAOA), which decomposes a large optimization problem into smaller sub-problem

arXiv (cs.AI) · Jul 22, 2026
4 min
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware AnalysisResearch

Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis

Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation, the opacity and escalating API costs of closed-weight frontier models motivate exploration of open-weight alternatives. However, many open-weight models are large, demanding significant compute resources and incurring non-trivial hosting costs that place them beyond reach for resource-constrained deployments. This paper investigates whether orchest

arXiv (cs.AI) · Jul 22, 2026
4 min
ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training TransformersResearch

ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing efficient attention methods usually reduce this bottleneck by either imposing sparsity, so that each query attends to only a small subset of keys, or by using low-rank/kernel sketches, so that global interactions are compressed into a lower-dimensional representation. We propose \emph{ELSAA}, an efficient low-rank and sparse approximation of attention. Importantly, ELSAA does \emph{not} decompose the learned projection or output matrices of the Trans

arXiv (cs.AI) · Jul 22, 2026
4 min
The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural NetworksResearch

The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks

Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of training points swapping one coordinate, which vanishes if and only if the coordinate carries no interaction, remains informative for piecewise-linear networks, and equals in expectation the per-coordinate interaction mass of the interventional Shapley-GAM. The loss turns additivity into

arXiv (cs.AI) · Jul 22, 2026
4 min
PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous DrivingResearch

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cu...

Papers with Code · Jul 22, 2026
2 min
StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video GenerationResearch

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation

Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex driving conditions, making them unsuitable for real-time interactive applications. We present \emph{StreamHOI}, a low-latency streaming framework for long-duration HOI video generation. Instead of converting heavily conditioned HOI pipelines into streaming systems, we study how an image-to-video streaming generator should organize historical memory to preserve interactions under bounded latency. We find that the standard sink-local memory design faces a

arXiv (cs.AI) · Jul 22, 2026
4 min
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningResearch

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most playe

arXiv (cs.AI) · Jul 22, 2026
3 min
Active Inference as a Convex Markov Decision ProcessResearch

Active Inference as a Convex Markov Decision Process

Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). In this formulation, the pragmatic terms are linear in the predictive state marginals and therefore equivalent to reward maximization in a latent MDP, while the epistemic value introduces a nonlinear component that distinguishes EFE minimizatio

arXiv (cs.AI) · Jul 22, 2026
4 min
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODResearch

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communicat

arXiv (cs.AI) · Jul 22, 2026
4 min
Google commits $40M to the Genesis Mission | Google Cloud BlogResearch

Google commits $40M to the Genesis Mission | Google Cloud Blog

Google is committing $40 million in AI tokens and cloud credits to support the DOE’s Genesis Mission and accelerate groundbreaking scientific discovery.

Google DeepMind · Jul 22, 2026
4 min
RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVsResearch

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-ti...

Papers with Code · Jul 22, 2026
2 min
How Meta’s AI Models Are Powering the First Wave of Genesis Mission ProjectsResearch

How Meta’s AI Models Are Powering the First Wave of Genesis Mission Projects

Lawrence Berkeley National Laboratory — one of the US Department of Energy's premier research laboratories, known for Nobel Prize-winning work in physics, chemistry, and materials science — operates some of the most advanced scientific facilities on the planet. Among them is the Advanced Light Source (ALS), a football field-sized facility that produces intensely bright beams of X-ray light, allowing researchers to study materials from the atomic and molecular scale all the way to plants. The AL

Meta AI · Jul 22, 2026
5 min
Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image ClassificationResearch

Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification

Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover semantics are closely ...

Papers with Code · Jul 22, 2026
2 min
Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated LearningResearch

Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning

Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireless transmission, es...

Papers with Code · Jul 22, 2026
1 min
NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot SimulationResearch

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Exi...

Papers with Code · Jul 22, 2026
1 min
RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin LearningResearch

RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning

We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learn...

Papers with Code · Jul 21, 2026
1 min
Equilibrium Causal Games: Separation, Identification, and the Identifiability of Cyclic Latent StatesResearch

Equilibrium Causal Games: Separation, Identification, and the Identifiability of Cyclic Latent States

Power grids, markets, and interacting populations, settle into feedback driven equilibria observed through unknown sensors. Our Equilibrium Causal Game (ECG)...

Papers with Code · Jul 21, 2026
2 min
Detect Early, Escalate Rarely: Anytime Detection of AI-Generated Video from the Compressed BitstreamResearch

Detect Early, Escalate Rarely: Anytime Detection of AI-Generated Video from the Compressed Bitstream

Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasingly by a large vision-language model. Detection...

Papers with Code · Jul 21, 2026
2 min
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement LearningResearch

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy text from the input into their reasoning traces rather than productively solving the problem. We show that this behavior is pervasive across frontier long-context LLMs and intensifies with context length. By separating each prompt into task-relevant key evidence and irrele

arXiv (cs.AI) · Jul 21, 2026
4 min
Appearance Pointers -- Multimodal Region Control of Diffusion TransformersResearch

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning t

arXiv (cs.AI) · Jul 21, 2026
3 min
ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesisResearch

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visu...

Papers with Code · Jul 21, 2026
2 min
CodeRescue: Budget-Calibrated Recovery Routing for Coding AgentsResearch

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery rout

arXiv (cs.AI) · Jul 21, 2026
3 min
Agents in the Wild: Where Research Meets DeploymentResearch

Agents in the Wild: Where Research Meets Deployment

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordinat

arXiv (cs.AI) · Jul 21, 2026
3 min
1-Lipschitz Neural Networks on Hadamard ManifoldsResearch

1-Lipschitz Neural Networks on Hadamard Manifolds

Controlling the Lipschitz constant of a neural network is a standard way to promote robustness and stability. Most existing constraining strategies are designed for Euclidean spaces. In this work, we construct and analyze a class of 1-Lipschitz neural networks on Hadamard manifolds. Our layers are of gradient-descent type, $1$-Lipschitz, and quasi-$α$-firmly nonexpansive. The core building blocks of the proposed architecture are Busemann functions, and we exploit the properties of Busemann gradient flows to design $1$-Lipschitz geometry-preserving layers. We provide explicit construc

arXiv (cs.LG) · Jul 21, 2026
3 min
Fundamental limits of distributed multiclass classification from simple binary decisionsResearch

Fundamental limits of distributed multiclass classification from simple binary decisions

We consider the problem of constructing a $K$-class classifier from the combination of $O(\log K)$ simple binary classifiers -- this is a natural paradigm to construct a sophisticated classifier in a distributed manner with each agent performing a relatively straightforward task. We study the fundamental performance limits of such a classifier when the corresponding binary classifiers are hyperplanes. For a stylized Gaussian setting where the $K$ class centers are independent Gaussian points in $\mathbb R^d$ and the observations are corrupted by Gaussian noise, we derive explicit per

arXiv (cs.LG) · Jul 21, 2026
3 min
Provable diffusion-based posterior sampling for linear inverse problems via DDIMResearch

Provable diffusion-based posterior sampling for linear inverse problems via DDIM

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only lightweight, coordinate-wise modifications to the standard DDIM update, while explicitly incorporating the measurement model. The key idea is to perform posterior sampling separately along each s

arXiv (cs.AI) · Jul 21, 2026
4 min
ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative ModellingResearch

ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold. Due to the success of diffusion models and flow matching, one of the more common beliefs is the importance of transforming the noise distribution to the data distribution gradually through many small transformations. We ask whether this is truly necessary, and take a minimalist approach to designing a competitive generative mod

arXiv (cs.LG) · Jul 21, 2026
4 min
ISO: An RLVR-Native Optimization StackResearch

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectra

arXiv (cs.AI) · Jul 21, 2026
4 min
Associative Emotional Learning in Convolutional Neural NetworksResearch

Associative Emotional Learning in Convolutional Neural Networks

Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli. Whereas computational models such as the Rescorla-Wagner model have shed light on this important function, the limitations of these models are also known, especially when they are applied to neural data. The advent of deep neural networks has opened another avenue for modeling associative emotional learning. In this work we proposed a deep neural network model of visual valence processing, consisting of a visual module that encodes complex natural

arXiv (cs.AI) · Jul 21, 2026
4 min
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&DResearch

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an arti

arXiv (cs.AI) · Jul 21, 2026
4 min
CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic InterpretabilityResearch

CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability

Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently requires stitching together separate implementations for discovery, evaluation, and intervention, as well as hand-authoring the contrastive prompts required by many discovery methods. This fragmentation makes methods difficult to compare and limits their application beyond canonical tasks. We introduce CircuitKIT, a source-available library that connects the circuit-analysis workflow throug

arXiv (cs.LG) · Jul 21, 2026
3 min
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged InformationResearch

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is de

arXiv (cs.AI) · Jul 21, 2026
3 min
Staypoint Detection from Noisy Trajectory Data [Experiment Paper]Research

Staypoint Detection from Noisy Trajectory Data [Experiment Paper]

Detecting staypoints from raw trajectory data is fundamental to numerous spatial computing applications. This process transforms raw numeric sequences of geolocations into semantically meaningful locations, such as homes, workplaces, or restaurants. Despite its importance for semantic trajectory analysis, staypoint detection lacks standard benchmarks, and existing algorithms have never been systematically evaluated. This gap persists because no publicly available datasets provide both raw individual trajectories and ground-truth staypoint annotations. This benchmark paper addresses t

arXiv (cs.LG) · Jul 21, 2026
3 min
From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVsResearch

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasible trajectories, all onboard and in real time. Conventional approaches treat mapping and planning as separate stages and often rely on binary occupancy for collision checking. We argue that these two stages should be co-designed around a single representation: a signed distance function (SDF). By encoding distance to the nearest obstacle, an SDF provides richer information for planning and trajectory optimization than occupancy alone. We develop a

arXiv (cs.AI) · Jul 21, 2026
4 min
Riemannian Deep Learning:Modules, Networks, and GeometriesResearch

Riemannian Deep Learning:Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroup

arXiv (cs.AI) · Jul 21, 2026
3 min
Real-time optimal control with shallow recurrent decoder networksResearch

Real-time optimal control with shallow recurrent decoder networks

Controlling dynamical systems in real-time across multiple scenarios is critical to enabling adaptive control strategies, ensuring stability and efficiency. However, to tailor control actions in response to varying scenarios, traditional optimal control problems typically require several system simulations, which are often computationally demanding due to the high-dimensionality of the underlying spatio-temporal dynamics. In this work, we exploit SHallow REcurrent Decoder networks-based Reduced Order Modeling (SHRED-ROM) to synthesize a real-time closed-loop controller for high-dimen

arXiv (cs.LG) · Jul 21, 2026
3 min
LLM Detection as an Intervention: Downstream Impact under Strategic User BehaviorResearch

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how u

arXiv (cs.AI) · Jul 21, 2026
4 min
Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business ProcessesResearch

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes

This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful agents, as a model-quality benchmark target, we present three executable recipes -- SQL analytics with repair loops, agentic retrieval-augmented generation with evidence gating, and human-in-the-loop policy review with interrupt and checkpoint recovery -- to show how typed state, conditional routing, deterministic tools, retries, interrupts, checkpoints,

arXiv (cs.AI) · Jul 21, 2026
4 min
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systemsResearch

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broade

arXiv (cs.AI) · Jul 21, 2026
4 min
GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based ModelsResearch

GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. While Graph Neural Networks have emerged as fast, data-driven surrogates, their practical deployment is severely constrained by a spatial generalization gap. Standard models rely on transductive feature initializations that tie travel demand to fixed network topologies, preventing seamless transfer to new urban environments. To overcome this structural limitation, this research proposes a network-agnostic initialization layer, termed Geometrically Unconstrained Inductiv

arXiv (cs.AI) · Jul 21, 2026
4 min
They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack SurfaceResearch

They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface

We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct production LLMs across three providers, behind an LLM firewall in shadow mode. A single untrusted input - an external issue requesting a "usage-telemetry" feature - asks for code that exfiltrates process secrets (dict(os.environ)) to an attacker URL, laundered as observability. Across a pre-registered A x B (x C) factorial (N=20; naive arm N=60) we find: (1) the entry agent does not leak its system prompt (0/40); (2) an authority-framed injection ("pre-app

arXiv (cs.AI) · Jul 21, 2026
4 min
Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case InvestigationResearch

Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation

Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, an autoencoder-based anomaly signal, TreeSHAP explanations, and a bounded LLM investigation agent applied to cases the classifier scores uncertainly. Before any model comparison, we identify and remove a simulator-specific balance shortcut that would otherwise inflate baseline performance. After this correction, neither the graph features nor the a

arXiv (cs.AI) · Jul 21, 2026
4 min
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic SurveillanceResearch

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Ac

arXiv (cs.AI) · Jul 21, 2026
4 min
PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology ImageResearch

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence directly from gigapixel WSIs largely untested. We introduce PathAgentBench, a benchmark for evaluating evidence-seeking vision-language models (VLMs) across four complementary capabilities: image-to-text matching for evidence interpretation, text-to-image retrieval fo

arXiv (cs.AI) · Jul 21, 2026
4 min
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasksResearch

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information fr

arXiv (cs.AI) · Jul 21, 2026
4 min
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language ModelsResearch

Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models

Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry before compliance degrades, and how much context a model can hold before recall and honesty degrade. We report two controlled experiments crossing all three factors on one held, contamination-free synthetic corpus (the "Book of Veyra," 8,780 uniquely-named entities, deterministically regenerable from a fixed seed), evaluated across five models. Exp

arXiv (cs.AI) · Jul 21, 2026
4 min
Sequential Learner Modeling Using Multi-Relational Graph Convolutional NetworksResearch

Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks

User modeling is a critical task in a variety of personalized systems. Recognizing their effectiveness in learning from graph-structured data, Graph Neural Networks (GNNs), particularly Graph Convolutional Networks (GCNs), are increasingly employed for user modeling. However, existing approaches typically treat different relation types in a graph as homogeneous, limiting their ability to capture richer semantics and construct more informative user models. While multi-relational GNNs (MR-GNNs) have been adopted for representation learning and recommendation, their application for user

arXiv (cs.AI) · Jul 21, 2026
4 min
Inference-Time Steering for Cross-Lingual Factual Consistency in LLMsResearch

Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs

Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual inconsistency, where they shift their empirical answer distributions based solely on the prompt language. We investigate whether these biases can be mitigated at inference time, forcing an English-prompted model to answer as if it were queried in target languages (German, Spanish, Bulgarian), and evaluate four intervention strategies: zero-shot contextual steerin

arXiv (cs.AI) · Jul 21, 2026
4 min
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party MeetingsResearch

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging fo...

Papers with Code · Jul 21, 2026
1 min
Introducing Gemini 3.5 Flash CyberResearch

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber to help defenders find, validate, and patch software vulnerabilities quickly and efficiently.

Google DeepMind · Jul 21, 2026
5 min
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine TranslationResearch

The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation

Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest research indicating that RLVR could be the preferred training method for translating legal documents due to the induced reasoning capabilities, it raises the question whether it is really attributed to the reasoning or more generally to the training paradigm. We investigate the importance of including the model's reasoning trace in the generated respo

arXiv (cs.AI) · Jul 21, 2026
3 min
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric RewardsResearch

Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automated evaluation of feedback quality remains limited. We propose RLAES, a unified LLM framework that jointly optimizes essay scoring and feedback generation through RL. To make feedback quality measurable, interpretable, and usable for training, we introduce Rubric-based Feedback Evaluation (

arXiv (cs.AI) · Jul 21, 2026
4 min
OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality DegradationResearch

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often impe...

Papers with Code · Jul 21, 2026
1 min
DAIS: Dependency-Aware Intermediate QA Supervision for Complex ReasoningResearch

DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence and provide limit...

Papers with Code · Jul 21, 2026
1 min
SWITi: Quantifying and Reducing Tiling Artifacts with Sliding Window Inner TilingResearch

SWITi: Quantifying and Reducing Tiling Artifacts with Sliding Window Inner Tiling

SWITi is a test-time method for reducing artifacts in tiled predictions, particularly for neural networks that learn posterior distributions from which solut...

Papers with Code · Jul 21, 2026
2 min
SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data ProcurementResearch

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth acquiring. We prese...

Papers with Code · Jul 21, 2026
2 min
Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level TimingResearch

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing mea...

Papers with Code · Jul 21, 2026
1 min
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberResearch

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.

Google DeepMind · Jul 21, 2026
7 min
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberResearch

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.

Google DeepMind · Jul 21, 2026
7 min
When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment PrioritizationResearch

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones s...

Papers with Code · Jul 20, 2026
2 min
AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator ModelingResearch

AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator Modeling

Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video. Label aggregation aims to in...

Papers with Code · Jul 20, 2026
2 min
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternativesResearch

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a l...

Papers with Code · Jul 20, 2026
2 min
The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual MetricResearch

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap c

arXiv (cs.LG) · Jul 20, 2026
3 min
Patch Policy: Efficient Embodied Control via Dense Visual RepresentationsResearch

Patch Policy: Efficient Embodied Control via Dense Visual Representations

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sacrificing both fine-grained spatial detail and the benefits of large-scale visual pre-training. While there exist policies that do operate on dense patch features like large vision-language-action models (VLAs), they tend to be heavy and slow, inheriting the full cost of a billion-parameter vision-language model (VLM) backbone.

arXiv (cs.LG) · Jul 20, 2026
4 min
Patch Policy: Efficient Embodied Control via Dense Visual RepresentationsResearch

Patch Policy: Efficient Embodied Control via Dense Visual Representations

Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either com...

Papers with Code · Jul 20, 2026
2 min
Automated Discovery Has No Universally Superior HarnessResearch

Automated Discovery Has No Universally Superior Harness

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harnes

arXiv (cs.AI) · Jul 20, 2026
4 min
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMsResearch

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two prac

arXiv (cs.AI) · Jul 20, 2026
4 min
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft PrefixesResearch

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B,

arXiv (cs.AI) · Jul 20, 2026
4 min
Causal Discovery on Irregular Time SeriesResearch

Causal Discovery on Irregular Time Series

Causal discovery methods have shown strong performance in temporal systems, but they typically rely on regular and discrete lag structures, limiting their applicability to regularly sampled data. However, many real-world tasks require dealing with irregularly sampled streams of events, such as sensor streams, healthcare data, and financial transactions. In this work, we propose an extension of PCMCI+, a state-of-the-art method for causal discovery on regular multivariate time series, to allow for handling irregular time series. Instead of modelling causal relations through fixed-lag

arXiv (cs.LG) · Jul 20, 2026
3 min
Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal InferenceResearch

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedding space, the generator estimates conditional expected outcomes or their contrasts, and a plug-in rule selects an action. This formulation connects action-specific vector search with nearest-neighbor matching in causal inference. We decompose the regret of the two-step method into candidate-generation regret

arXiv (cs.LG) · Jul 20, 2026
3 min
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment AnalysisResearch

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use restrictive licenses, and remain computationally expensive, limiting large-scale slide-level clinical and research use. Here, we introduce GigaPath

arXiv (cs.AI) · Jul 20, 2026
4 min
Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLASResearch

Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS

This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, whereas the latent structure is decomposed into invariant factors with shared loadings and heterogeneous factors with environment-specific loadings. Such a model is motivated by transfer learning and latent factor regression, where one seeks stable low-dimensional representations for both interpretation and robus

arXiv (cs.LG) · Jul 20, 2026
4 min
Learning Adaptive Safety Margins for Visual NavigationResearch

Learning Adaptive Safety Margins for Visual Navigation

Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts, while permissive margins lead to near-boundary shortcuts under perception bias. Diffusion-based planners propose diverse trajectory candidates from egocentric RGB-D, yet reliable selection remains the bottleneck. We propose a context-conditioned safety critic that learns an adaptive clearance preference for ranking diffusion proposals, decomposed into three complementary terms: (i) a s

arXiv (cs.AI) · Jul 20, 2026
4 min
PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to ReasoningResearch

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typica

arXiv (cs.LG) · Jul 20, 2026
3 min
Three-Body Scattering for Generative ModelingResearch

Three-Body Scattering for Generative Modeling

Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct regression supervision for a one-step generator. Three-Body Scattering Modeling (TBSM) for generation turns the energy distance into a constant-size per-projectile interaction: each projectile is attracted toward one real source and repelled from one independently generated source. Conditioned on the projectile and its condition, its expectation equa

arXiv (cs.LG) · Jul 20, 2026
4 min
Certified Training for Convolutional PerturbationsResearch

Certified Training for Convolutional Perturbations

Vision models have been found to be susceptible to perturbations such as motion blur induced at runtime by a shaking camera. This impedes their deployment in critical applications since phenomena such as slightly blurred vision might lead to failures, for example an object detector missing objects. While methods such as data augmentation or Adversarial Training can improve empirical robustness, they lack formal safety guarantees, making it difficult to identify and mitigate hidden vulnerabilities. We introduce a novel Certified Training approach that leverages an efficient encoding o

arXiv (cs.LG) · Jul 20, 2026
3 min
EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain DatabaseResearch

EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database

Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. However, conventional compressors often struggle to preserve fine structural details at high compression ratios (CRs), and implicit neural representations (INRs) require costly per-volume optimization and produce models with fixed CRs. To respond, we present EVOLVE, an autoencoder (AE)-based volume-compression framework that targets high CRs for offline compression, with three key contributions. Firs

arXiv (cs.LG) · Jul 20, 2026
4 min
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationsResearch

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers commit to limited transformations and fixed workload assumptions, so achieving high performance on a new application requires hand-crafting an efficient implementation. We present FlashRT, an agent harness that guides coding agents to lift simple developer-written reference

arXiv (cs.LG) · Jul 20, 2026
4 min
A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive ManufacturingResearch

A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing

Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve, a phenomenon known as concept drift. Maintaining surrogate fidelity under drift, particularly when models must also capture aleatoric uncertainty, remains an open challenge. Existing adaptive frameworks lack principled mechanisms for detecting when updates are needed, for efficiently adapting models from limited streaming data, and for certifying that updates genuinely improve predictive performance. Here we present an adaptive Digital Twin fram

arXiv (cs.AI) · Jul 20, 2026
4 min
OR Else: A Differentiable Trust Region for Policy OptimizationResearch

OR Else: A Differentiable Trust Region for Policy Optimization

PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in the scalar objective's derivative. We ask whether Output Reset (OR), a smooth one-sided saturation rule, offers a useful alternative for large language model post-training. PPO-OR and GRPO-OR replace the clipped policy term with an OR squared-margin loss in rollout-relative token log-ratio space; the advantage sign determines the update direction, and a token contributes zero direct OR residual after crossing the favorable margin. We compare PPO-

arXiv (cs.AI) · Jul 20, 2026
4 min
The Calibration Channel Determines the Bayes-Error Proxy: An Exact Law for Temperature-Induced DistortionResearch

The Calibration Channel Determines the Bayes-Error Proxy: An Exact Law for Temperature-Induced Distortion

The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al. estimates the irreducible error of a binary task directly from probability-valued labels. Recent work by Ushio et al. showed that this estimator is fragile when the probabilities are not the true posterior: even perfectly calibrated soft labels can yield a substantially inaccurate estimate, and they propose isotonic calibration as a consistent remedy. We complement that line of work by characterizing exactly how the most widely used post-hoc calibration map -- temperature scaling -- distorts the proxy. We

arXiv (cs.LG) · Jul 20, 2026
4 min
TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory MinimizationResearch

TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization

Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and prototyping. However, despite their value as coding assistants, agent-generated code tends to be larger and more verbose than the corresponding human-written implementation. In this work, we show that the cause lies in the agent's own search process: while iterating toward a passing solution, an agent accumulates speculative edits, abandoned hypotheses, and temporary changes that persist into the final patch. This may seem harmless for a single p

arXiv (cs.AI) · Jul 20, 2026
4 min
Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge DevicesResearch

Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices

Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural networks. We investigated Differentiable Logic Gate Networks (Diff-Logic) as a hardware-native alternative that compiles models into pure Boolean circuits executable via bitwise CPU operations. Through rigorous iso-parameter experiments across four EEG datasets spanning two classification tasks, binary dementia detection and 3-class emotion recognition, we compared Diff-Logic against matched-capacity Multi-Layer Perceptron (MLP) and Binarized Neural Network (BNN) basel

arXiv (cs.AI) · Jul 20, 2026
4 min
Totally Positive Matrices and the Highest-Order Coefficients of the Characteristic PolynomialResearch

Totally Positive Matrices and the Highest-Order Coefficients of the Characteristic Polynomial

We investigate the extent to which totally positive matrices can be distinguished through the highest-order coefficients of their characteristic polynomials. To identify the most informative coefficients, we also employed neural-network classifiers together with feature-attribution methods. Using datasets built from several structured totally positive families, including products of positive bidiagonal matrices, Vandermonde matrices, and Cauchy matrices, we find that the coefficients (a_{n-1}, a_{n-2}, a_{n-3}) already contain strong discriminatory information for separating totally

arXiv (cs.LG) · Jul 20, 2026
4 min
LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and ApplicationsResearch

LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications

Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains. In smart grids, recent work applies agentic schemes to forecasting, optimization, and control, wrapping trusted solvers behind language interfaces and orchestrating multi-step workflows. The literature lacks a unified approach to designing and evaluating such systems. LLMs can produce numerically plausible yet physically infeasible outputs, evaluation protocols vary across tasks, and the boundary between what the model s

arXiv (cs.AI) · Jul 20, 2026
4 min
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial ConstraintsResearch

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LL

arXiv (cs.AI) · Jul 20, 2026
4 min
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and ReasoningResearch

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free

arXiv (cs.AI) · Jul 20, 2026
4 min
SGA: Plug&Play Geometric Verification for Educational Video SynthesisResearch

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensuring spatial correctness and visual legibility remains challenging, as existing frameworks emphasize pedagogical content while overlooking geometric occlusions. We propose the Symbolic Geometric Agent (SGA), a plug-and-play module for code-centric animation pipelines that intercepts LLM-generated code, performs partial execution to extract symbolic scene graphs, and applies targeted refinement when spatial conflicts are detected. We fur

arXiv (cs.AI) · Jul 20, 2026
3 min
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?Research

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out transfer, and causal intervention. The susceptibility is largely install

arXiv (cs.AI) · Jul 20, 2026
4 min
Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation SteeringResearch

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches and operate at the input level, offering no fine-grained control over the reasoning process itself. Related work analyzes and discovers latent transition dynamics in the reasoning traces from Large Language Models. Building on this, we statistically characterize these states, and show that failure trajectories get stuck in self-loops, exhausting the token

arXiv (cs.AI) · Jul 20, 2026
4 min
SciForma: Structure-Faithful Generation of Scientific DiagramsResearch

SciForma: Structure-Faithful Generation of Scientific Diagrams

Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directi...

Papers with Code · Jul 20, 2026
2 min
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMsResearch

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Gr

arXiv (cs.AI) · Jul 20, 2026
4 min
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingResearch

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of

arXiv (cs.AI) · Jul 20, 2026
4 min
Enhancing Rubric-based RL via Self-DistillationResearch

Enhancing Rubric-based RL via Self-Distillation

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreov

arXiv (cs.AI) · Jul 20, 2026
4 min
SelectInfer: Selective Neuron Loading and Computation for On-Device LLMsResearch

SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Existing approaches to model compression and optimization often rely on coarse-grained pruning or quantization, which can compromise accuracy or require re-training and fine-tuning. In this work, we introduce SelectInfer, a neuron-level optimization framework that enables efficient LLM inference on edge devices through selective

arXiv (cs.AI) · Jul 20, 2026
3 min
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionResearch

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content ar...

Papers with Code · Jul 20, 2026
2 min
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionResearch

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from veri

arXiv (cs.AI) · Jul 20, 2026
4 min
Generalised Bellman recurrence and three dualities in sequential decision-makingResearch

Generalised Bellman recurrence and three dualities in sequential decision-making

What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one

arXiv (cs.AI) · Jul 20, 2026
3 min
SGN: A Similarity-based Generative Network for Data Generation under Distribution ShiftResearch

SGN: A Similarity-based Generative Network for Data Generation under Distribution Shift

Generative models trained on a source domain often produce samples that are poorly aligned with shifted target domains, limiting their effectiveness for target-domain data augmentation. Although target-specific adaptation can reduce this mismatch, it typically requires additional optimization and domain-specific parameters. We propose a Similarity-based Generative Network (SGN), a reusable framework that is trained once on labeled source data and applied to new target domains without parameter updates. SGN learns a latent space structured by label-induced pairwise similarities while

arXiv (cs.AI) · Jul 20, 2026
3 min
Human Grounded Evaluation of Large Language Models for Optical Network AutomationResearch

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. We present HuGLEN, a stepwise evaluation pipeline that uses an LLM-as-a-judge together with a small set of expert ratings to enable scalable and reproducible comparison of candidate LLMs, and to rank them using a quality efficiency score (QES). We demonstrate HuGLEN for translating outputs from an explainable artificial intelligence (XAI) model for the optical network quality of transmission (QoT) estimation task into op

arXiv (cs.AI) · Jul 20, 2026
3 min
Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation DataResearch

Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data

Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent receives a dataset, an evaluation script, and one editable file, and iterates without supervision: modify the code, measure, keep the change if the score improves. But what does the agent actually optimize--the developer's intent, or the literal number? We ran this loop on a real production task: deciding which Quranic verses appear in a noisy speech-recognition transcript and splitting the transcript by verse. Two frontier coding agents, Claude

arXiv (cs.AI) · Jul 20, 2026
4 min
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent SecurityResearch

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for \emph{adaptive multi-round attacks against memoryless LLM defenders}: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting

arXiv (cs.AI) · Jul 20, 2026
4 min
The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse SearchResearch

The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search

Organizations often pool dispersed information into one ranking and then allow many agents to act on that shared view. In a discovery problem, this can impro...

Papers with Code · Jul 20, 2026
2 min
Exploration Matters for Escaping the Blur Trap in 3D Gaussian SplattingResearch

Exploration Matters for Escaping the Blur Trap in 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) employs Gaussian primitives for explicit scene representation, facilitating real-time, high-fidelity reconstruction and novel vi...

Papers with Code · Jul 20, 2026
1 min
Entanglement geometry separates circuit cutting, classical hardness, and trainabilityResearch

Entanglement geometry separates circuit cutting, classical hardness, and trainability

Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classic...

Papers with Code · Jul 20, 2026
1 min
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language ModelsResearch

FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models

Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-b...

Papers with Code · Jul 20, 2026
2 min
Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM AgentsResearch

Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical reasoning, and tool u...

Papers with Code · Jul 20, 2026
2 min
Token-Level Off-Policy Learning for Faithful Generation Under Distribution ShiftResearch

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. ...

Papers with Code · Jul 20, 2026
1 min
OrderMoE: An expert similarity driven distributed edge MoE inferenceResearch

OrderMoE: An expert similarity driven distributed edge MoE inference

Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging...

Papers with Code · Jul 19, 2026
2 min
CLDRoute: Conditional Latent Diffusion for Routability Map Generation in Physical DesignResearch

CLDRoute: Conditional Latent Diffusion for Routability Map Generation in Physical Design

Accurate routability estimation during physical design is important for reducing costly post-routing iterations. Prior learning-based methods treat this task...

Papers with Code · Jul 18, 2026
1 min
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight QuantizationResearch

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights and the growing KV cache. We propose PagedWeight, a novel management method for MoE LLM serving that dynamically quantizes MoE model's weights at runtime and balances expert-weight precision with the KV cache sizes. PagedWeight exposes and effectively navigates the complex tradeoff between the model's task accuracy, memory consumption,

arXiv (cs.LG) · Jul 17, 2026
3 min
A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic ComputingResearch

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. In this work, we focus on energy-based thermodynamic computing where the stochastic process is well described by Langevin dynamics with tunable energy potentials. The implementation of such potentials in physical hardware enables us to generate and sample from basic parameterized energy-based models. We demonstrate how to construct and train popular c

arXiv (cs.LG) · Jul 17, 2026
3 min
Cluster-Aware Matching via Laplacian Optimal TransportResearch

Cluster-Aware Matching via Laplacian Optimal Transport

In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable within a coherent region, finding a robust region-to-region alignment is more desirable than establishing a precise point-to-point correspondence. To this end, we propose a novel approach for cluster-aware matching based on Laplacian Optimal Transport (LapOT). The key idea is to regularize the optimal transport problem with quadratic Laplacian

arXiv (cs.LG) · Jul 17, 2026
3 min
Physics-enhanced reinforcement learning for real-time optimal control of dynamical systemsResearch

Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synthesize optimal control strategies. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the exploration-exploitation dilemma in high-dimensional spaces. In this work, we bridge RL and traditional optimal control for dynamical system with a novel Physics-EnhA

arXiv (cs.LG) · Jul 17, 2026
4 min
Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle VulnerabilitiesResearch

Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities

Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, and vehicle operations. These vulnerabilities are commonly documented as plain text in the Common Vulnerabilities and Exposures (CVE) database; however, security practitioners require structured information about affected assets, types of weaknesses, and attack behaviors to effectively mitigate the risks from these vulnerabilities. To thi

arXiv (cs.AI) · Jul 17, 2026
4 min
When Does Muon Help Agentic Reinforcement Learning?Research

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%); high-rate AdamW controls retain no post-update success. The effect depends on the advantage estimator and learning rate. At 3e-5, Muon improves GRPO

arXiv (cs.AI) · Jul 17, 2026
3 min
Behaviour-Conditioned Neural Processes for Adaptive Residential Short-Term Load ForecastingResearch

Behaviour-Conditioned Neural Processes for Adaptive Residential Short-Term Load Forecasting

Residential short-term load forecasting (STLF) is challenging because household demand is heterogeneous, temporally variable, and shaped by diverse behavioural routines. This work investigates whether inferred behavioural structure can be embedded within the forecasting mechanism of a Neural Process-based probabilistic model, rather than used only as an external grouping signal, for context-conditioned residential STLF. We propose a behaviour-conditioned Attentive Neural Process framework that treats each load profile as a forecasting task. Behavioural structure is represented by a d

arXiv (cs.LG) · Jul 17, 2026
4 min
An Exam for Active ObserversResearch

An Exam for Active Observers

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual

arXiv (cs.AI) · Jul 17, 2026
4 min
PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety AssessmentResearch

PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment

Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands continuous, real-time monitoring systems capable of anticipating conflicts before they escalate into crashes. We present PRISA, a modular infrastructure LiDAR framework leveraging privacy-preserving, low-light-robust roadside sensors for long-term traffic observation and real-time risk detection at the edge. The framework comprises two core components: a s

arXiv (cs.LG) · Jul 17, 2026
4 min
Learning Standard Model structure from LHC data with Riemannian flow matchingResearch

Learning Standard Model structure from LHC data with Riemannian flow matching

In this work we demonstrate that a single transformer-based generative model can capture Standard Model structure spanning five decades of invariant mass, from the sub-GeV regime to the TeV continuum, a range that no single Monte Carlo sample covers. To achieve this we design \textsc{ShellFlow}, a Riemannian conditional flow matching model that, given the recorded event composition, generates each particle on its on-shell manifold. Its only physics priors are the on-shell condition and the invariant-mass formula. The model is trained on $\sim 10^{9}$ real $pp$ collision events from t

arXiv (cs.LG) · Jul 17, 2026
3 min
Improving Improved Kernel PLSResearch

Improving Improved Kernel PLS

Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms. This article focuses on two shared steps, the computation of the $\mathbf{X}$ rotations, $\mathbf{R}$, and the $\mathbf{Y}$ loadings, $\mathbf{Q}$, and accelerates both. For $\mathbf{R}$, term-by-term accumulation is replaced by a direct evaluation strategy that requires the same number of multiplications but parallelizes better on modern hardware. For $\mathbf{Q}$, I identify - to the best of my knowledge, for the first time - equivalences showing that each $\mathbf{Y}$

arXiv (cs.LG) · Jul 17, 2026
4 min
When Do Multi-Agent Systems Help? An Information Bottleneck PerspectiveResearch

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounded relay messages. We show that, under infinite relay bandwidth, any SAS can be simulated b

arXiv (cs.AI) · Jul 17, 2026
4 min
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement LearningResearch

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our knowledge. ToolSciVer equips a VLM with three type-aware visual tools, table row/column focus, chart-to-structure p

arXiv (cs.AI) · Jul 17, 2026
3 min
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle GovernanceResearch

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance

AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way. Yet existing work on AI trustworthiness remains either too high-level to support lifecycle monitoring and reassessment or too narrowly metric-driven to connect with governance needs. We therefore propose a lightweight methodology for auditable trustworthiness levels in AI governance. The methodology has two components: a formal framework for represent

arXiv (cs.AI) · Jul 17, 2026
4 min
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning DataResearch

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description

arXiv (cs.AI) · Jul 17, 2026
4 min
Harmonizing AI Safety ThresholdsResearch

Harmonizing AI Safety Thresholds

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for

arXiv (cs.AI) · Jul 17, 2026
3 min
The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic InfrastructureResearch

The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure

State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants. However, these guarantees rely on participants outside this set correctly executing the protocol's transition semantics. Agentic validators expose a weaker boundary: an authenticated, responsive, non-equivocating, and protocol-compliant reasoning participant may still endorse a semantically invalid transition due to reasoning errors. We call this failure mode an epistemic fault, and the collective phenomenon the Honest

arXiv (cs.LG) · Jul 17, 2026
4 min
Understanding Reasoning from Pretraining to Post-TrainingResearch

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute swe

arXiv (cs.AI) · Jul 17, 2026
4 min
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement LearningResearch

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing domain classifiers, value-guided data filtering, or representation learning. Instead, we study the domain adaptation problem from a generative mode

arXiv (cs.AI) · Jul 17, 2026
4 min
Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapesResearch

Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes

AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to structure. We propose they are also a scientific object that can be analyzed directly: a learned encoding of protein conformational organization that can be probed and characterized. By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, we show that the trained model produces physically structured conformational landscapes

arXiv (cs.LG) · Jul 17, 2026
4 min
Pick-to-Learn Calibration of an MPC Policy for an Origin-to-Destination Flight ProblemResearch

Pick-to-Learn Calibration of an MPC Policy for an Origin-to-Destination Flight Problem

This paper illustrates the Pick-to-Learn methodology applied to the calibration of a Model Predictive Control policy. While developed around a specific example, the presentation is meant to highlight a methodology of broad applicability. The example concerns an aircraft traveling from an origin point to a destination point in the presence of uncertain crosswinds and a low-connectivity zone that should be avoided. The MPC policy is parameterized by two hyperparameters, which are selected from data by the P2L procedure. Starting from a dataset of 400 wind realizations, also called scen

arXiv (cs.LG) · Jul 17, 2026
3 min
Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation ForecastingResearch

Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting

Precipitation nowcasting over the immediate 10-90 min period is important for flood management and real-time decision-making in urban regions. Conventional short-range forecasting with high-resolution numerical weather prediction requires frequent data assimilation, model initialization, and spin-up, introducing computational latency. Machine learning provides an alternative by learning storm evolution directly from high-frequency observations and producing forecasts quickly after training. This is particularly relevant for Mumbai, India, where monsoon convection, land-sea interactio

arXiv (cs.LG) · Jul 17, 2026
4 min
HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying DetectionResearch

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical Cross-modal Incongruity Graph Network), a novel framework that models cross-modal incongruity at token, phrase, a

arXiv (cs.AI) · Jul 17, 2026
4 min
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA ModelsResearch

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these c

arXiv (cs.AI) · Jul 17, 2026
4 min
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and StandardizationResearch

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization

Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundations of agentic systems: reasoning, planning, tool use, multi-agent coordination, and evalua

arXiv (cs.AI) · Jul 17, 2026
3 min
Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence TomographyResearch

Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography

Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative biomarkers of retinal structure. Indeed, there is a growing interest in the analysis of OCTs in the context of neurodegenerative diseases. However, segmentation remains challenging due to speckle noise, shadowing artifacts, low contrast between adjacent layers, anatomical variability across subjects, and domain shifts arising from different acquisition protocols and clinical populations. While deep learning methods have achieved remarkable performance, their robustness a

arXiv (cs.AI) · Jul 17, 2026
4 min
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry AnalysisResearch

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result against a jointly trained model on the same data. On task-goal completion, merging matches joint RL -- and every

arXiv (cs.AI) · Jul 17, 2026
3 min
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoningResearch

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing de

arXiv (cs.AI) · Jul 17, 2026
4 min
Deep and Probabilistic Models for Gene Regulatory Network InferenceResearch

Deep and Probabilistic Models for Gene Regulatory Network Inference

Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection, while evaluation is constrained by incomplete reference networks and point-estimate outputs that lack uncertainty. GRN reconstruction also depends on prior knowledge to constrain TF-gene interactions, yet available priors are often assay-dependent and diffi

arXiv (cs.LG) · Jul 17, 2026
3 min
Loop the Loopies!Research

Loop the Loopies!

We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N-fold increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla T

arXiv (cs.AI) · Jul 17, 2026
3 min
DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model EmbeddingsResearch

DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings

Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regional domains, or computationally intensive process-based models unsuitable for daily continental-scale use. We present DELUGE, a multimodal deep learning framework for daily pluvial flood damage prediction at ~1 km resolution and national scale, trained on spatially and temporally corrected NFIP claims (2017-2022) and structured around t

arXiv (cs.LG) · Jul 17, 2026
4 min
SciForge: An AI-Native, Multimodal Workbench for Scientific DiscoveryResearch

SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery

Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, auditable research state. We present SciForge, a multimodal research-native AI workbench that reserves the graphical interface for human judgment while search, parsing, model routing, workflow execution, plotting, writing, and presentation generation run as modular agent-accessible services. SciForge is built around five pillars: (i) \emph

arXiv (cs.AI) · Jul 17, 2026
4 min
Revisiting data-driven dynamic security assessment with a tabular foundation modelResearch

Revisiting data-driven dynamic security assessment with a tabular foundation model

Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system using machine learning. Existing approaches face two limitations. First, they require a large labelled database for training, with a separate model trained, tuned, and maintained for each contingency in a potentially long list of credible contingencies. Second, the trained models generalize poorly to unseen contingencies. This work addresses the limitations by using a tabular foundation model (TFM) that assesses stability through in-context learning,

arXiv (cs.AI) · Jul 17, 2026
4 min
Rethinking Quantum Continual Learning with Quantum Fisher InformationResearch

Rethinking Quantum Continual Learning with Quantum Fisher Information

Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We propose quantum elastic weight consolidation (QEWC), a quantum Fisher information (QFI)-informed regularization method for mitigating forgetting. Unlike conventional elastic weight consolidation based on classical Fisher information (CFI), which measures parameter importance through measurement-dependent output statistics, QEWC uses QFI to

arXiv (cs.AI) · Jul 17, 2026
4 min
CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward ApproachResearch

CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach

This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy Gradient (DDPG). The agents use technical indicators, cyclical calendar encodings, and daily news sentiment scores produced by LLaMA 3.2 1B. To

arXiv (cs.LG) · Jul 17, 2026
4 min
Candidate Attended Dialogue State Tracking Using BERTResearch

Candidate Attended Dialogue State Tracking Using BERT

Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions and generate responses. The increasingly popular dialogue system applications like Google Assistant, Siri and Alexa need to support a large number of services and APIs, resulting in growing attention to the scalability of such systems. Especially for some domains with little or no training data, the capability of transferring existing knowl

arXiv (cs.AI) · Jul 17, 2026
3 min
DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense PredictionResearch

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual depthwise separable inverted bottlenecks to improve frozen VFM utilization through fusion

arXiv (cs.AI) · Jul 17, 2026
4 min
Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage GridsResearch

Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids

Increases in photovoltaic generation, charging of electric vehicles and heat-pump demand challenge operating limits in low-voltage distribution grids. This requires curative curtailment methods that can operate under sparse observability, noisy measurements, and imperfect grid models. Unlike prior end-to-end reinforcement-learning approaches for partially observable curtailment, this work decouples congestion detection and control by combining a random-forest violation pre-classifier with an actor-critic controller, and evaluates its robustness to measurement noise and grid-parameter

arXiv (cs.AI) · Jul 17, 2026
3 min
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AIResearch

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest seriously in safety, fairness, and oversight cannot consistently prove to consumers, regulators, and shareholders that their systems go beyond the bare minimum of compliance. What is missing is a way for society to recognize or compare the difference. The result is a trust gap: a structural condition in which responsible development effor

arXiv (cs.AI) · Jul 17, 2026
4 min
A Formally Grounded ODRL Evaluator: Implementation and ComparisonResearch

A Formally Grounded ODRL Evaluator: Implementation and Comparison

The ODRL policy language is emerging as the de-facto standard for policy modelling data access and usage preferences, AI governance policies and data workflows in European dataspaces. The current standard has no mathematical formal semantics to describe how a system should implement policy evaluation. This has resulted in a variety of systems and tools that implement their own interpretation of the language, which limits interoperability and cannot guarantee consistent results. Based on an existing semantic model of ODRL, we formalise the problems of ODRL evaluation for the access co

arXiv (cs.AI) · Jul 17, 2026
3 min
FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side ChannelsResearch

FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels

Federated Learning (FL) over 5G cellular networks protects raw data but remains vulnerable to side-channel leakage. Prior fingerprinting attacks assume packe...

Papers with Code · Jul 16, 2026
2 min
RoboTTT: Context Scaling for Robot PoliciesResearch

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for

arXiv (cs.AI) · Jul 16, 2026
4 min
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity GeneratorsResearch

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, DiffusionNFT offers an efficient forward-process RL framework that does not require reverse-process trajectories or likelihood estimation. However, applying such RL methods to MeanFlow remains underexplored. DiffusionNFT optimizes instantaneous velocities, whereas MeanFlow sampl

arXiv (cs.LG) · Jul 16, 2026
4 min
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper RevisionsResearch

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight visual grammar to advance a specific argument. To address this, we present SciDiagramEdit, a benchmark and skill-

arXiv (cs.AI) · Jul 16, 2026
4 min
Online Neural Space Time Memory for Dynamic Novel View SynthesisResearch

Online Neural Space Time Memory for Dynamic Novel View Synthesis

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demandi

arXiv (cs.LG) · Jul 16, 2026
4 min
Pretraining Data Can Be Poisoned through Computational PropagandaResearch

Pretraining Data Can Be Poisoned through Computational Propaganda

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally,

arXiv (cs.AI) · Jul 16, 2026
3 min
SceneBind: Binding What and Where Across Vision, Audio and LanguageResearch

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further prop

arXiv (cs.AI) · Jul 16, 2026
3 min
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security AgentsResearch

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at

arXiv (cs.AI) · Jul 16, 2026
4 min
Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment ClassifierResearch

Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier

The growing use of Bitcoin as a decentralized digital asset and investment tool has sparked strong interest in understanding its market behavior. This study presents a new approach to analyze Bitcoin market sentiment by combining on-chain and financial data with social media posts. Unlike models that aim to predict prices, this work focuses on explaining market sentiment using blockchain transactions, historical price data of Bitcoin, and daily Twitter sentiment classifications. The method merges sentiment trends with on-chain and financial metrics, normalized into a dataset for deta

arXiv (cs.LG) · Jul 16, 2026
4 min
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent CollaborationResearch

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can become trapped in repetitive loops, wasting search budgets and ultimately compromising the quality and completeness of the final output. We introduce SearchOS, a system-level multi-agent framework that turns fragile, implicit search progress into explicit, persistent, and s

arXiv (cs.AI) · Jul 16, 2026
4 min
teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving DataResearch

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and collected without interventions, which makes causal questions such as "How would rain change traffic density?" difficult to answer. We present teLLMe, a system for exploratory causal analysis of urban driving datasets. The system starts from a structured event table built from dashcam annotations and combines causal structure learning with the PC algorithm, bootstrap-based stability checks, and query-specific effect estimation using line

arXiv (cs.AI) · Jul 16, 2026
4 min
AutoSynthesis: An agentic system for automated meta-analysisResearch

AutoSynthesis: An agentic system for automated meta-analysis

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a research question in natural language, AutoSynthesis formulates a search strategy, retrieves scientific literature, screens candidate studies, assesses full-text eligibility, extracts quantitative statistics, computes standardized effect sizes, and finally performs random-

arXiv (cs.AI) · Jul 16, 2026
3 min
Mutable Low-Rank Sketches for Retrain-Free RecommendationResearch

Mutable Low-Rank Sketches for Retrain-Free Recommendation

A common bottleneck in two-stage recommendation is embedding staleness: when a user rates a new item, their embedding remains fixed until the next retrain cycle. We propose mutable sketches, which store each user's preferences in a KP-tree (a sparse segment tree with sum aggregation), fit a low-rank projection once, and recompute embeddings on-the-fly as ratings arrive. We prove that each new observation monotonically tightens the prediction error envelope (Theorem 1), a guarantee that FunkSVD and eALS lack. On KuaiRec, the mutable sketch achieves 0.810 RMSE at 1.8% data read vs. ALS

arXiv (cs.LG) · Jul 16, 2026
3 min
In-Place Tokenizer Expansion for Pre-trained LLMsResearch

In-Place Tokenizer Expansion for Pre-trained LLMs

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small fraction of their parameters. On a compact model those matrices are a material share of per-token decode bandwidth, so on-device models ship small voca

arXiv (cs.AI) · Jul 16, 2026
4 min
Data Driven Block Replacement SchedulingResearch

Data Driven Block Replacement Scheduling

We develop data-driven algorithms for maintaining $N$ independent identical machines under a \textit{block replacement policy}, in which each machine is replaced upon failure and all machines are jointly replaced at regular intervals of length $k$. The goal is to learn the cost-minimizing interval $k^*$ from operational data when the lifetime distribution is unknown. At each decision epoch, the operator selects $k \in \{1, 2, \ldots, K\}$, observes the resulting failure history (a mixture of complete and right-censored lifetimes) and incurs a per-unit-time cost governed by the renewa

arXiv (cs.LG) · Jul 16, 2026
4 min
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk SpaceResearch

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-state direction analysis and random-split null tests, we show that content danger (CD) and physical danger (PD) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on the CD/PD separability, we propose PRISM, a single-

arXiv (cs.AI) · Jul 16, 2026
4 min
NeuronSoup: Evolving Asynchronous, Shared-Neuron Temporal Graphs without BackpropagationResearch

NeuronSoup: Evolving Asynchronous, Shared-Neuron Temporal Graphs without Backpropagation

We present NeuronSoup, a neural computation architecture that replaces synchronous layer-by-layer processing with asynchronous, delay-mediated signal propagation through a pool of shared neurons. Each path in the network routes a continuous-valued signal from one input neuron to one output neuron through a variable number of intermediate hidden neurons. Hidden neurons are physically shared across paths: when two paths pass through the same neuron, the second arrival encounters the accumulated state left by the first, producing constructive or destructive interference that depends on

arXiv (cs.LG) · Jul 16, 2026
4 min
Symbal: Detecting Systematic Misalignments in Model-Generated CaptionsResearch

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal,

arXiv (cs.AI) · Jul 16, 2026
4 min
Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped LangevinResearch

Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin

Unadjusted samplers such as unadjusted Hamiltonian Monte Carlo and underdamped Langevin are well-known to be biased. Metropolis--Hastings adjustment has been conventionally incorporated into Hamiltonian Monte Carlo to eliminate the bias. However, this adjustment can significantly increase the iteration complexity due to the small step size required for reasonable Metropolis acceptance rates. In this work, we extend the \emph{delocalization of bias} phenomenon, previously established for the overdamped Langevin algorithm, to these two unadjusted algorithms. We show that to control the

arXiv (cs.LG) · Jul 16, 2026
3 min
BadWAM: When World-Action Models Dream Right but Act WrongResearch

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use sma

arXiv (cs.LG) · Jul 16, 2026
4 min
MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue LocalizationResearch

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscuring whether visual input helped, hurt, or was ignored. We introduce \textbf{MM-IssueLoc}, a controlled benchmark and evaluation protocol for repository-level localization with visual evidence. MM-IssueLoc contains 652 issue-PR instances across 23 languages, with annota

arXiv (cs.AI) · Jul 16, 2026
4 min
Self-Evolving Human-Centered Framework for Explainable Depression Symptom AnnotationResearch

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression-related datasets, labels are often assigned without structured evidence, symptom-level justification, or traceable alignment with the criteria of the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition, Text Revision (DSM-5-TR), limiting both transparency and downstream model interpretability. We propose a self-evolving, expert-in-the-loop annotation framework for Major Depressive Disorder (MDD) that combines la

arXiv (cs.AI) · Jul 16, 2026
4 min
Self-Evolving Human-Centered Framework for Explainable Depression Symptom AnnotationResearch

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression...

Papers with Code · Jul 16, 2026
2 min
Mask-Aware Policy Gradients for Diffusion Language ModelsResearch

Mask-Aware Policy Gradients for Diffusion Language Models

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each masked position and which positions to remask. We formalize this as a two-stage action MDP, sho

arXiv (cs.AI) · Jul 16, 2026
3 min
Subjective Risk Decomposition: A New View for Uncertainty QuantificationResearch

Subjective Risk Decomposition: A New View for Uncertainty Quantification

We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric uncertainty measures can be derived via decomposition of a subjective risk, based on a strictly proper loss. Reverse cross-entropy provides a prominent example, where decomposition recovers the classic information-theoretic uncertainty terms. The same approach recovers numerous measures previously proposed across the UQ literature, providing them a c

arXiv (cs.AI) · Jul 16, 2026
3 min
Plover: Steering GUI Agents through Plan-Centric InteractionResearch

Plover: Steering GUI Agents through Plan-Centric Interaction

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users' ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning a

arXiv (cs.AI) · Jul 16, 2026
3 min
Can We Trust Item Response Theory for AI Evaluation?Research

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for

arXiv (cs.AI) · Jul 16, 2026
4 min
RTS Smoother-Guided Learning of Physics-Based Neural Differential ModelsResearch

RTS Smoother-Guided Learning of Physics-Based Neural Differential Models

Ordinary differential equations (ODEs) are widely used to model dynamical systems in physics, biology, neuroscience, and physiology, but in many applications some equations of the dynamics are unknown and only a subset of the state variables are measured. We propose a hybrid neural--physics framework in which the known components of the ODE are kept explicit and the missing components are represented by a neural network. The proposed method consists of two stages where we alternate between state and parameter estimation and iterate until a predetermined criterion is met. Specifically

arXiv (cs.LG) · Jul 16, 2026
4 min
T^2MLR: Transformer with Temporal Middle-Layer RecurrenceResearch

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across n

arXiv (cs.AI) · Jul 16, 2026
3 min
Benchmarking Multimodal Large Language Models for Scientific Visualization LiteracyResearch

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance

arXiv (cs.AI) · Jul 16, 2026
4 min
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionResearch

MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels medical AI errors by severity (1--5) and safety gate type (missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, source support gap). The current public release (v0.2.1) contains 44 clinician-reviewed synthetic cases with severity annotations, a live HuggingFace leaderboard preview, a safet

arXiv (cs.AI) · Jul 16, 2026
3 min
The Industrialization of Research ; On AI-Driven Science and Its ConsequencesResearch

The Industrialization of Research ; On AI-Driven Science and Its Consequences

Artificial intelligence is transforming scientific research - not merely as a more powerful instrument, but as an autonomous participant in the research cycle itself. This transition constitutes, in the most precise sense of the term, the industrialization of research: a shift from a craft model, in which knowledge, method, and judgment are embedded in the researcher, to a pipeline model, in which these steps are decomposed, automated, and supervised. The US Department of Energy's Genesis Mission is the most ambitious current instantiation of this shift, but the fundamental questions

arXiv (cs.AI) · Jul 16, 2026
4 min
Scaling Behavior Foundation Model for Humanoid RobotsResearch

Scaling Behavior Foundation Model for Humanoid Robots

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learn

arXiv (cs.AI) · Jul 16, 2026
4 min
On-Policy Delta DistillationResearch

On-Policy Delta Distillation

On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output distribution. The delta signal is defined as the difference between the teacher model and its base model prior to instructi

arXiv (cs.LG) · Jul 16, 2026
3 min
Concept-Guided Spatial Regularization for World Models in Atari PongResearch

Concept-Guided Spatial Regularization for World Models in Atari Pong

World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, while the world models themselves are rarely studied in isolation. We examine five representative visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM. After reproducing their training pipelines and matching the reported agent performance, we freeze the learned world models and evaluate them with a closed-loop rollout diagnostic: a policy trained separately from the corresponding MBRL agent interacts with each frozen model, and the generated video

arXiv (cs.AI) · Jul 16, 2026
4 min
NIFA: Nonlinear IMC enhanced FPGA for efficient ML inferenceResearch

NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference

Recent FPGAs have improved deep learning (DL) inference efficiency through dedicated tensor blocks and in-BRAM computation. ReRAM-based analog in-memory computing (IMC) pushes efficiency further, offering an order-of-magnitude improvement in compute density and energy efficiency over conventional digital logic by performing vector-matrix multiplication (VMM) directly within the ReRAM crossbar; prior work has integrated such IMC blocks into FPGAs for DL inference. However, conventional IMC designs support only static-weight VMM, leaving nonlinear operations and dynamic matrix-matrix m

arXiv (cs.AI) · Jul 16, 2026
4 min
Learning in Infinitesimal Non-Compositional SketchesResearch

Learning in Infinitesimal Non-Compositional Sketches

This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting. Machine learning problems are specified as sketches: graphs with commutativity conditions $\mathcal D$, limit cones $\mathcal L$, and colimit cocones $\mathcal K$, generalizing the usual scalarization of loss functions or vector space assumptions. Non-compositionality is defined purely as failure of a universal factorization problem, not as

arXiv (cs.LG) · Jul 16, 2026
4 min
Long-Context Fine-Tuning with Limited VRAMResearch

Long-Context Fine-Tuning with Limited VRAM

Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We combine Hierarchical Global Attention (HGA) with segment-wise backpropagation and tiered KV storage. Only the active segment remains differentiable in VRAM; older KV is detached into RAM or NVMe, and HGA loads a bounded set of exact historical tokens for each query block. On Qwen3-8B with 4-bit QLoRA and PG19, dense training on a 16 GB Quadro RTX 5000 fits 2,048 tokens but fails at 4,096, whereas HGA reaches 16,384 tokens with 15.28 GB peak VRAM. U

arXiv (cs.AI) · Jul 16, 2026
4 min
AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation LearningResearch

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint, which commits every retrieval direction to the same stability-plasticity trade-off. We propose AlphaWiSE, a post-hoc weight-space interpolation method that composes two frozen source checkpoints. For each aligned parameter tensor identified by its checkpoint key, AlphaWiSE fits one scalar interpolation coeffi

arXiv (cs.LG) · Jul 16, 2026
3 min
Evaluating covariate balance for long time horizon Markov decision processesResearch

Evaluating covariate balance for long time horizon Markov decision processes

This article explores the application of covariate balance diagnostics for detecting the presence of hidden confounding/model miss-specification in studies applying offline reinforcement learning (RL) to deriving optimal treatment recommendations. The results demonstrate that, either there is a high risk of bias within existing offline RL studies for treatment recommendations or, existing covariate balance metrics are not sufficient to assess such studies. Regardless, existing offline RL studies cannot be concluded as being statistically robust. The conclusions propose future researc

arXiv (cs.LG) · Jul 16, 2026
3 min
An Introduction to Sparse Identification of Nonlinear Dynamics for Engineering ApplicationsResearch

An Introduction to Sparse Identification of Nonlinear Dynamics for Engineering Applications

Many engineering problems involve phenomena whose governing equations are poorly characterized or only partially known. Surrogate modeling techniques such as neural networks can capture the behavior of these systems, but they typically demand large training datasets that are difficult to obtain in engineering contexts and yield models with limited physical interpretability. The Sparse Identification of Nonlinear Dynamics (SINDy) method addresses both limitations by performing sparse regression over libraries of candidate nonlinear terms, recovering interpretable governing equations f

arXiv (cs.LG) · Jul 16, 2026
4 min
Kernel weighted importance sampling for off-policy evaluation in contextual banditsResearch

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weighted importance sampling), particularly under complex conditions including behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of vanilla weighted importance sampling with the linearity of vanilla importance sampling.

arXiv (cs.LG) · Jul 16, 2026
3 min
DriftWorld: Fast World Modeling through DriftingResearch

DriftWorld: Fast World Modeling through Drifting

Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. This creates a bottleneck for diffusion-based world models: multistep sampling makes each rollout expensive, limiting large-scale action search at inference time. We introduce DriftWorld, an action-conditioned world model based on drifting generative models. Rather than denoising iteratively at inference, DriftWorld learns an action-conditioned drift during training, allowing it to generate future frames from the current obs

arXiv (cs.LG) · Jul 16, 2026
4 min
Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI ScreeningResearch

Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening

Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries. Existing methodologies attempt to bypass these constraints using computationally expensive, fully fine-tuned hybrid architectures that relegate spatial explainability to a post-hoc approximation rather than an intrinsic model property. We propose a parameter-efficient framework utiliz

arXiv (cs.LG) · Jul 16, 2026
4 min
cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical DataResearch

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data

High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on low-dimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. To address this gap, we introduce categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that preserves the original data matrix while augmenting it with interpret

arXiv (cs.LG) · Jul 16, 2026
4 min
SMC-ES: Automated synthesis of formally verified control policiesResearch

SMC-ES: Automated synthesis of formally verified control policies

The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.e., policies) that are not only performant but also provably safe and robust. While learning-based methodologies such as Reinforcement Learning offer flexible and scalable approaches to automatically synthesize such controllers, they typically lack the formal guarantees necessary for safe deployment. To bridge this gap, we propose a novel simulation-based methodology to automatically synthesize policies with formal guarantees regarding performance, safety, an

arXiv (cs.LG) · Jul 16, 2026
4 min
Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical ImagingResearch

Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging

Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare. It works by minimizing the distance between matched (positive) data modalities, while maximizing the distance between mismatched (negative) samples. Traditional CL frameworks typically assume instance-based correspondence within data batches, treating all non-paired samples as negatives. However, this assumption often fails in medical settings, where samples may share high-level semantic attributes,

arXiv (cs.LG) · Jul 16, 2026
4 min
Google DeepMind and Isomorphic Labs approach to bioresilienceResearch

Google DeepMind and Isomorphic Labs approach to bioresilience

Google DeepMind and Isomorphic Labs approach to bioresilience, using AI models to support prevention, detection and response.

Google DeepMind · Jul 16, 2026
4 min
Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code QualityResearch

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language ...

Papers with Code · Jul 16, 2026
2 min
An LLM-Based Automatic Sportscast Solution for Robot Soccer MatchesResearch

An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches

RoboCup has always been a scenario to develop systems that solve real-world problems. Driven by the main goal of playing against the 2050 FIFA World Cup cham...

Papers with Code · Jul 16, 2026
1 min
Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via MultimodalityResearch

Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality

Cross-modal learning, i.e., learning to predict one modality from another, is a fundamental mechanism for self-supervision via leveraging multimodality. Many...

Papers with Code · Jul 16, 2026
2 min
Qubes OS Security in the Public RecordResearch

Qubes OS Security in the Public Record

Qubes OS is a revealing case for security measurement because its architecture makes component boundaries security-relevant. We present a protocol-driven lon...

Papers with Code · Jul 16, 2026
2 min
Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated LearningResearch

Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning

Federated fine-tuning of large pre-trained models increasingly relies on Low-Rank Adaptation (LoRA) to reduce communication and computation, but heterogeneou...

Papers with Code · Jul 15, 2026
2 min
Beyond scalar losses: calibrating segmentation models via gradient vector field surgeryResearch

Beyond scalar losses: calibrating segmentation models via gradient vector field surgery

Region-based loss functions, such as the Dice loss, have established themselves as the de facto standard for highly class- and region-imbalanced segmentation...

Papers with Code · Jul 15, 2026
1 min
Leveraging unlabelled data for generalizable neural population decodingResearch

Leveraging unlabelled data for generalizable neural population decoding

Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly lever

arXiv (cs.LG) · Jul 15, 2026
4 min
Linear Independent Component Analysis via Optimal TransportResearch

Linear Independent Component Analysis via Optimal Transport

Linear Independent Component Analysis (ICA) recovers jointly independent source signals from their linear mixtures. To achieve this, classical ICA algorithms attempt to maximize non-Gaussianity, measured by negentropy, which is linked to independence by information theory. Because exact negentropy optimization is intractable, they rely on proxy contrast functions, such as fourth-order cumulants, and parametric log-likelihoods. We propose instead to measure non-Gaussianity using the squared Wasserstein distance $W_2^2$ to a standard Gaussian. We prove that the Wasserstein distance bet

arXiv (cs.LG) · Jul 15, 2026
3 min
MetaPerch: Learning from metadata for bioacoustics foundation modelsResearch

MetaPerch: Learning from metadata for bioacoustics foundation models

Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data. Recent work has shown that supervision alone can produce SotA species detection models when trained on this large-scale data -- however, there remains unutilized potential in the form of recording metadata readily available within these community-driven data hubs. In this work, we explore the use of metadata -- such as location and time -- as auxiliary supervision signals, allowing the model to leverage species-metadata correlations in its lear

arXiv (cs.LG) · Jul 15, 2026
3 min
Screening of Biosecurity Features in Metagenomic Data with Evo 2 ProbesResearch

Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention

arXiv (cs.LG) · Jul 15, 2026
4 min
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning ModelsResearch

Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex, multi-step tasks. However, when errors occur, current interaction approaches typically involve re-generating another response that may make mistakes again, or users laboriously flag the faulty step in follow-up turns that may get responses <You are right, I made a mistake here> followed by similar errors recurring. To address this issue, we propose an efficient human intervention mechanism for precisely correcting reasoning errors in LLMs, termed

arXiv (cs.AI) · Jul 15, 2026
3 min
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake EducationResearch

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education

This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake preparedness and conscious action among primary-school students. The system extends the award-winning STEM project Earthquaker moving from mechanical simulation with Lego WeDo2 to cognitive and metacognitive processing. The robotics component uses Lego WeDo2 automation to simulate seismic response, letting students interact with sensors

arXiv (cs.AI) · Jul 15, 2026
4 min
AI-accelerated End-to-End Framework for Rapid Professional UpskillingResearch

AI-accelerated End-to-End Framework for Rapid Professional Upskilling

By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skills gap grew from roughly 3 days in 2014 to 36 days in 2018. Most current frameworks accelerate single stages of upskilling programs and generally lack industry validation. We present an end-to-end framework that applies AI acceleration across five stages of knowledge acquisition, content development, content review and verification, teaching, and assessment development; with a strong focus on both production and learning efficiency. Three strong external signals valida

arXiv (cs.AI) · Jul 15, 2026
3 min
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case StudyResearch

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study

Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cursive chancery hand used in palace memorials, despite limited labeled data. We study a multi-expert system that reuses checkpoints from an iterative fine-tuning process as domain specialists and uses a lightweight page-level image classifier to dispatch pages by visual style. When the checkpoint pool lacks a suitable specialist, we train an additional expert for that domain. On three frozen test sets, the routed system matches the selected speciali

arXiv (cs.AI) · Jul 15, 2026
3 min
Early Adoption of Agentic Coding Tools by GitHub ProjectsResearch

Early Adoption of Agentic Coding Tools by GitHub Projects

Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic coding tools, (2) project-level agentic PR productivity, and (3) human-agent collaboration pa

arXiv (cs.AI) · Jul 15, 2026
4 min
Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical StudyResearch

Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study

With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global energy mix continues to grow. Unlike conventional power generation, the output of renewable energy sources cannot be controlled as consistently due to their dependence on environmental conditions. Therefore, reliable prediction of current and future energy production is essential. In this paper, we report findings from two structured literature reviews on real-world renewable energy prediction tasks: wind turbine power curve modeling and photovoltaic

arXiv (cs.AI) · Jul 15, 2026
4 min
Transforming Rank: How Architecture Navigates the Spectral Pathologies of DepthResearch

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. ...

Papers with Code · Jul 15, 2026
2 min
Transforming Rank: How Architecture Navigates the Spectral Pathologies of DepthResearch

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank. We show that skip connections trade off rank collapse against ensemble-like behavior, controlled by the relative scales of the branch and the skip: skip connections

arXiv (cs.AI) · Jul 15, 2026
4 min
Lighthouse RL: Sample-Efficient Circuit Optimization via Strategic Reset PointsResearch

Lighthouse RL: Sample-Efficient Circuit Optimization via Strategic Reset Points

In this paper, we introduce Lighthouse RL, a sample-efficient reinforcement learning (RL) approach for analog circuit sizing. Traditional methods lack generalization across different performance targets, while standard RL approaches waste resources exploring unpromising regions. Our method addresses these inefficiencies through a strategic reset strategy that initializes episodes from high-performing configurations discovered during training, called "lighthouses". These states, which are closer to the target objectives, guide exploration toward promising regions. When compared to RL

arXiv (cs.LG) · Jul 15, 2026
3 min
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective ViolationResearch

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational controls to achieve security-relevant compromise. This paradigm remains necessary for AI-enabled systems, but it is no longer sufficient. In such systems, adversaries may influence prompts, retrieved content, sensor inputs, training data, memory, tools, or human-AI interaction loops to alter system behavior without directly compromising the underlying infrastructure. This paper reframes penetration testing for AI-enabled systems as objecti

arXiv (cs.AI) · Jul 15, 2026
4 min
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0Research

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this

arXiv (cs.AI) · Jul 15, 2026
4 min
Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza PendulumResearch

Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum

We suggest using the Lyapunov characteristic exponent (LCE) as a dense reward signal for the reinforcement learning problem of stabilizing the inverted pendulum with vertical motion. With LCE, the agent not only successfully found the oscillatory motion known as the Kapitza pendulum but also damped the pendulum's pivoting, leaving it in a strictly upright position.

arXiv (cs.LG) · Jul 15, 2026
3 min
The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous CommerceResearch

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive recommendation algorithms to autonomous, goal-directed agents capable of executing purchasing decisions, the conventional understanding of consumer-brand relationships requires a structural reevaluation. By synthesizing extant literature across human-machine teaming, consumer decision-making, and algorithmic trust dynamics, we demonstrate that traditional loyalty models fail to account for algorithmic bounded rationality and constructed

arXiv (cs.AI) · Jul 15, 2026
4 min
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon AgentsResearch

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become sparse and high-variance as trajectories grow to tens or hundreds of tool calls. They can also be misleading: a failed rollout may contain many useful actions that move the agent closer to the goal, yet outcome-only training assigns them the same negative advantage as the eventual mistake. We propose TRACE (Turn-l

arXiv (cs.LG) · Jul 15, 2026
4 min
Multimodal Empirical Bayes Variational Autoencoders for Joint Longitudinal and Time-to-Event ModelingResearch

Multimodal Empirical Bayes Variational Autoencoders for Joint Longitudinal and Time-to-Event Modeling

Longitudinal tumor measurements, dropout information, and genetic covariates provide complementary information about treatment response, but integrating these data sources within a single population modeling framework remains challenging. We extend the empirical Bayes variational autoencoder (EB-VAE) framework to joint longitudinal and time-to-event modeling and evaluate it on tumor growth data. The framework represents inter-individual variability using latent individual effects regularized by a covariate-conditioned empirical Bayes prior, while a decoder maps these latent effects t

arXiv (cs.LG) · Jul 15, 2026
4 min
Music-to-Dance Generation via Atomic MovementsResearch

Music-to-Dance Generation via Atomic Movements

Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its compositional nature, making generated dances structurally incoherent and difficult to control. In this work, we introduce a structure-aware framework that models choreography as a sequence of atomic movements-semantically interpretable motion events that serve as the building blocks of dance. To construct this

arXiv (cs.AI) · Jul 15, 2026
4 min
A Self-Evolving Agent for Longitudinal Personal Health ManagementResearch

A Self-Evolving Agent for Longitudinal Personal Health Management

Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture that updates support as a person's routines, preferences, measurements and risks change. It separates shared safety rules and medical knowledge from private longitudinal memory containing profile facts, reusable procedures and episodic traces. After each episode, induction determines what should update the profile, revise a procedure, remain episodic or be excluded. We evaluated HealthClaw with a synthetic

arXiv (cs.AI) · Jul 15, 2026
4 min
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement LearningResearch

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language mod...

Papers with Code · Jul 15, 2026
1 min
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates CodeResearch

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not guide intermediate generation steps, such as those during autoregressive LLM decoding. Constrained decoding intervenes earlier by rejecting invalid tokens during sampling, but requires white-box model access and costly reimplementation for semantic constraints.We introduce generative compilation, the first approach to obtaining compiler feedback on

arXiv (cs.AI) · Jul 15, 2026
4 min
Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot CeilingsResearch

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $k$ verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.15712) shows that posterior log-odds grow linearly in $k$, so failure decays exponentially, and states that "a tight theory of partially correlated verifier cascades remains open." This note gives a minimal such theory. Modeling the per-instance false-accept rate on the generator's own errors as a latent variable $α\sim G$ (de Finetti), the exact cascade posterior is

arXiv (cs.AI) · Jul 15, 2026
4 min
AIMO Interpretability ChallengeResearch

AIMO Interpretability Challenge

We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' internal mechanisms. The challenge is motivated by a central limitation of standard reasoning benchmarks: strong final-answer accuracy does not reveal whether a model relies on stable reasoning mechanisms or exploits brittle reasoning shortcuts. Building on AI Mathematical Olympiad (AIMO) problems and submissions, together with resources from the Fields Model Initiative, the competition will provide (1) newly-publ

arXiv (cs.AI) · Jul 15, 2026
4 min
Experience Memory Graph: One-Shot Error Correction for AgentsResearch

Experience Memory Graph: One-Shot Error Correction for Agents

Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer from compounding errors and struggle to recover from failures. Existing self-correction mechanisms rely on prompt-based reflection, which is inherently brittle, incurs heavy time and API costs due to iterative trial-and-error loops, and produces task-specific memory that may be hard to generalize to new scenarios. To address this, we propose

arXiv (cs.AI) · Jul 15, 2026
4 min
Verifying formulas for interventional distributionsResearch

Verifying formulas for interventional distributions

We formalize verification in causal graphical models: deciding whether a given observational formula identifies a target interventional distribution. This opens a problem complementary to identification, asking not whether any identifying formula exists, but whether the given formula is identifying. We show that even sound and complete solutions to identification do not solve verification. We propose a falsifier as a first practical route forward, prove that it induces an almost-surely correct verifier for regular exponential-family models, and use the resulting verifier to develop t

arXiv (cs.AI) · Jul 15, 2026
3 min
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the WildResearch

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangle interaction understanding with dataset-specific supervision, limiting their ability to generalize to open-world and compositional scenarios. Recent HOI detectors attempt to leverage MLLMs through prompting strategies to transfer interaction-specific knowledge. However, such prompt-based approaches primarily focus on extracting discrimin

arXiv (cs.AI) · Jul 15, 2026
4 min
AI-Augmented Human Resource Management? Insights from German companiesResearch

AI-Augmented Human Resource Management? Insights from German companies

This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \enquote{augmenting} human resource management. Organisations employ generative AI or predictive analytics to transform traditional human resource functions, to streamline routine tasks and to reallocate resources toward strategic, people-centred activities. Our findings from interviews and group discussions and a survey (N=410) reveal that while AI tools enhance HR analytics capabilities, their adoption mainly serves efficiency and rationalising g

arXiv (cs.AI) · Jul 15, 2026
3 min
NodeImport: Imbalanced Node Classification with Node Importance AssessmentResearch

NodeImport: Imbalanced Node Classification with Node Importance Assessment

In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend to overfit to majority classes while underrepresenting minority classes. Existing solutions, which either prioritize nodes based on class size or synthesize new nodes for minority classes, often fall short of effectively addressing this imbalance issue. This paper introduces an approach to class-imbalanced node classification by utilizing

arXiv (cs.AI) · Jul 15, 2026
4 min
Multimodal Assessment of Pancreatic Cancer Resectability Using Deep LearningResearch

Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning

Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability. We introduce a fully automated multimodal deep learning framework that jointly analyzes 3D contrast enhanced CT and structured clinical information to classify patients into the three National Comprehensive Cancer Network (NCCN) resectability categories (upfront resectable, borderline resectable, locally advanced). The approach uses a Swin-UNETR backbone

arXiv (cs.AI) · Jul 15, 2026
3 min
Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion DetectionResearch

Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection

Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness against feasible traffic manipulation remains largely empirical. We present Traffic-Aware Randomized Smoothing (TA-RS), a classifier-agnostic certified defense that injects Gaussian noise exclusively into the directly controllable (DC) subspace -- features a remote attacker can modify -- during both fine-tuning and certification, aligning the smoothing distribution with the attacker-controllable subspace. We identify a critical prerequisite: applyi

arXiv (cs.AI) · Jul 15, 2026
4 min
CAS I: A Geometric Coding TheoremResearch

CAS I: A Geometric Coding Theorem

This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijections on the set of binary strings, called symmetries and define the symmetry prior of a string as the probability that a randomly chosen symmetry from a given group has the string as its unique fixed point. We show that for any fix-retractable symmetry group, a group admitting a computable section that selects an isolating symmetry for every string, the symmetry prior is a universal lower semi-computable semi-measure. In this case, the Geometric Cod

arXiv (cs.AI) · Jul 15, 2026
3 min
Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space CorrelationsResearch

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations

Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timesteps and self attention computation. As diffusion timesteps are reduced, the computation cost of self attention becomes the dominant bottleneck. Existing acceleration approaches largely inherit sparse attention techniques from large language models, which fail to consider the unique spatiotemporal correlation of video data. This paper presents Kaleido, an algorithm hardware codesign that accelerates all operations in vDiTs by exploiting channel-wis

arXiv (cs.AI) · Jul 15, 2026
4 min
MxGPS: Multiplex Graph Transformers for a Power Grid Foundation ModelResearch

MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model

Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting: the tendency of task-specific gradient signals to encode relational structure particular to the training topologies rather than the underlying physics, causing models to fail on unseen grids despite strong in-distribution performance. To expose and address this failure mode, we introduce MxGPS (Multiplex GPS), a multiplex graph transformer that

arXiv (cs.AI) · Jul 15, 2026
4 min
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM AgentsResearch

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, fro...

Papers with Code · Jul 15, 2026
2 min
ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization LevelResearch

ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level

We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \times n}$ into $A \...

Papers with Code · Jul 15, 2026
1 min
Discrete Diffusion Models: A Unified Framework from Tokenization to GenerationResearch

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering para...

Papers with Code · Jul 15, 2026
1 min
Reassessing Muon for Matrix FactorizationResearch

Reassessing Muon for Matrix Factorization

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and h...

Papers with Code · Jul 14, 2026
1 min
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump ProcessesResearch

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws $\textit{together}$, each modality reshapes the ot...

Papers with Code · Jul 14, 2026
2 min
Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and ExecutionResearch

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--turning a one-line edit into a small code-base audit. We argue the missing capability is task-aware execution-scope estimation: judging a task's difficulty, the information it truly needs, and the shortest reliable path before committing budget. We formalize minimum-sufficient execution and the Agent Cognitive

arXiv (cs.AI) · Jul 14, 2026
4 min
The Seriality Gap in Video Diffusion ModelsResearch

The Seriality Gap in Video Diffusion Models

When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computati

arXiv (cs.LG) · Jul 14, 2026
3 min
The Seriality Gap in Video Diffusion ModelsResearch

The Seriality Gap in Video Diffusion Models

When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere...

Papers with Code · Jul 14, 2026
1 min
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at ScaleResearch

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per second on a single server-grade GPU, far faster than existing object-level simulators, while

arXiv (cs.AI) · Jul 14, 2026
4 min
PalmClaw: A Native On-Device Agent Framework for Mobile PhonesResearch

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, inter

arXiv (cs.AI) · Jul 14, 2026
4 min
A Shortcut to Statistically Steady-State Turbulence with Flow MatchingResearch

A Shortcut to Statistically Steady-State Turbulence with Flow Matching

Many nonlinear physical systems exhibit an initial transient phase in which perturbations grow before nonlinear interactions lead to a statistically steady state. While this saturated regime is of primary interest, direct numerical simulations must resolve the full transient dynamics before reaching it, incurring significant computational cost. In Computational Fluid Dynamics, reduced-order approaches such as Large Eddy Simulation mitigate computational cost by modeling small-scale dynamics, enabling tractable approximations of turbulent flows. In contrast, for systems such as gyroki

arXiv (cs.LG) · Jul 14, 2026
4 min
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language ModelResearch

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the

arXiv (cs.AI) · Jul 14, 2026
4 min
Dynamic Resource Allocation for Ensemble Determinization MCTSResearch

Dynamic Resource Allocation for Ensemble Determinization MCTS

Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomness and hidden information. In particular, several Monte Carlo Tree Search (MCTS) variants are commonly used in such domains. In this paper, we propose a series of enhancements for Ensemble Determinization MCTS, introducing two axes for dynamic resource allocation. First, Dynamic Number of Determinizations, increases or decreases the number of currently used determinization trees depending on the behavior of so-far search. Second, Dyn

arXiv (cs.AI) · Jul 14, 2026
3 min
The Spectrum Is Not Enough: When Context Helps Time-Series ForecastingResearch

The Spectrum Is Not Enough: When Context Helps Time-Series Forecasting

A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different question: whether \emph{adding context}, a longer lookback, a retrieval plug-in, or a pretrained model, will help. These are not the same question. The value of context is a property of the operating point, not of the series. Any index built from the power spectrum is invariant under phase randomization, whereas the beyond-second-order value that retrieval and foundation models supply is not, because a phase-randomized series is asymp

arXiv (cs.LG) · Jul 14, 2026
4 min
Watermark Forensics for Generative Models: An Information-Theoretic PerspectiveResearch

Watermark Forensics for Generative Models: An Information-Theoretic Perspective

A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length $n$. One object organizes the answers. Let $S$ be the secret the mark carries (a user's identity or payload), and let the information profile $ν(t)=I(S;X_t\mid X_{<t})$ record how much the $t$-th token reveals about $S$ given the earlier ones. Its total mass pays for att

arXiv (cs.LG) · Jul 14, 2026
4 min
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan EvaluationResearch

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57 admissible deletions matched the analytic identity and threshold sign, and every route had at least one score-improving deletion. A score-seeking optimizer, allowed to restructure routes but

arXiv (cs.AI) · Jul 14, 2026
4 min
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMsResearch

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, resist and update, pull in opposite directions. We study them on a B

arXiv (cs.AI) · Jul 14, 2026
4 min
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem GenerationResearch

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram generation approaches struggle with analytic geometry: template methods cannot handle constraint-driven layouts, and generative models lack the geometric precision to render annotated conic curves correctly. We present FormalAnalyticGeo, a scalable framework for fully automatic generation of multimodal analytic geometry problems. Leveraging

arXiv (cs.AI) · Jul 14, 2026
4 min
Ensemble Controlled-Flow Filtering for Implicit Data AssimilationResearch

Ensemble Controlled-Flow Filtering for Implicit Data Assimilation

Data assimilation estimates the state of a dynamical system from model forecasts and incoming observations. Many observation mechanisms, however, are many-to-one, implicit, non-smooth, or accessible only through simulation, and need not provide the residual structures or likelihood guidance required by existing ensemble filters. We introduce implicit data assimilation, in which the analysis law is defined as an energy tilt of the forecast distribution. We then propose the Ensemble Controlled-flow Filter (EnCF), which realizes this update through a stochastic controlled flow and learn

arXiv (cs.LG) · Jul 14, 2026
3 min
Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code ModelsResearch

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexample as an oracle-relative refutation, and introduce PoPE (Popperian Placebo-controlled Evaluation): a methodology for measuring whether evidence that falsifies LLM-generated code can be used operationally by that same model. In PoPE, error content is paired with channel-specific placebos that keep the predeclared scaffold while ablating tas

arXiv (cs.AI) · Jul 14, 2026
4 min
Robustness of Deep Learning Models for PV Power Forecasting under NWP Forecast Errors: A Spatiotemporal and Physically Interpretable AnalysisResearch

Robustness of Deep Learning Models for PV Power Forecasting under NWP Forecast Errors: A Spatiotemporal and Physically Interpretable Analysis

Engineering use of AI forecasting models requires not only high nominal accuracy but also predictable behavior under uncertain inputs. In photovoltaic (PV) forecasting, this requirement is especially challenging because numerical weather prediction (NWP) errors are temporally correlated, state dependent, and physically coupled across variables. Existing evaluations, however, often rely on perfect forecast assumptions or simplistic perturbations that do not reflect these characteristics. This study presents a physically constrained robustness evaluation framework based on simulation,

arXiv (cs.LG) · Jul 14, 2026
4 min
ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start BenchmarkResearch

ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark

Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a resource is challenging for three reasons: cross-platform hotel names must be reconciled before interactions are comparable; quality must be audited with reproducible metrics rather than ad hoc cleaning; and public release must preserve privacy while remaining benchmarkable under realistic cold-start conditions. We introduce ViHoRec, a quality-controlled Vietnamese hotel recommendation dataset of 18{,}267 interactions between 6{,}832 users

arXiv (cs.AI) · Jul 14, 2026
3 min
Efficient Sequential Calibration with $O(T^{2/3-ε})$ Error BoundResearch

Efficient Sequential Calibration with $O(T^{2/3-ε})$ Error Bound

We study the online binary sequential calibration problem. A recent breakthrough by \citet{dagan2024breaking} overcomes the classical \(T^{2/3}\) barrier for calibration error. Building on this result, we present an efficient randomized forecaster that achieves an expected calibration error \(O(T^{2/3-\varepsilon})\) for some constant \(\varepsilon>0\). Our forecaster combines the \textsc{SPR-Calibration} procedure \citep{dagan2024breaking} with an outer Blackwell-style correction layer. The \textsc{SPR-Calibration} procedure controls calibration with respect to a surrogate sequence

arXiv (cs.LG) · Jul 14, 2026
3 min
Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision ProcessesResearch

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into

arXiv (cs.AI) · Jul 14, 2026
4 min
LatentFlow: A General Framework for Conditioning Stochastic ProcessesResearch

LatentFlow: A General Framework for Conditioning Stochastic Processes

Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box information, and global constraints all induce intractable conditional laws, requiring bespoke, model-specific constructions. We introduce LatentFlow, a single framework for conditioning stochastic processes, with no learned neural approximations and no training. Our starting point is to write the stochastic process as the deterministic image of a tractable latent innovation, $f_0 = T_{\vartheta}(ξ_0)$, with $ξ_0$ sampled from a simple refer

arXiv (cs.LG) · Jul 14, 2026
4 min
Contrastive-Collapsed Loss for Flexible and Geometrically Optimal Embeddings and Faster ConvergenceResearch

Contrastive-Collapsed Loss for Flexible and Geometrically Optimal Embeddings and Faster Convergence

In this work, we introduce CoCo, a loss function aimed at learning normalized and well-structured representations. The proposed loss encourages intra-class collapse and inter-class contrast while preserving sufficient flexibility for neural networks to approximate geometrically optimal embeddings with large angular separation between classes. We provide a theoretical analysis positioning CoCo with respect to related objectives such as dot regression and cross-entropy, showing that the new proposed loss benefits from closer initialization to the optimal configuration, more informative

arXiv (cs.LG) · Jul 14, 2026
3 min
Real-time fall detection based on vision for low-power edge platformsResearch

Real-time fall detection based on vision for low-power edge platforms

Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss event in a coupled dynamical system. We introduce a novel dual-LTC architecture comprising a Center-of-Mass (CoM) subsystem and a Base-of-Support (BoS) subsystem, both instantiated as Liquid T

arXiv (cs.AI) · Jul 14, 2026
4 min
Accelerated Mixing Time of Randomized Hamiltonian Monte CarloResearch

Accelerated Mixing Time of Randomized Hamiltonian Monte Carlo

We show the Randomized Hamiltonian Monte Carlo (RHMC) algorithm has accelerated mixing time guarantees for sampling from log-concave probability distributions. RHMC proceeds by repeatedly simulating the continuous-time Hamiltonian dynamics for some random integration times, and resetting the velocity to be an independent Gaussian random variable between each simulation. We show that when the target distribution is log-concave and satisfies an $α$-Talagrand inequality (for example, if the target distribution is $α$-strongly log-concave), if we use a random integration time from either

arXiv (cs.LG) · Jul 14, 2026
4 min
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon ConversationsResearch

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on

arXiv (cs.AI) · Jul 14, 2026
4 min
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress ProxiesResearch

UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies

Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect task completion, making the quality of these signals critical. Since such dense labels are rarely available at scale, normalized time within a demonstration is often used as a scalable substitute: later frames are treated as higher progress. However, this time-derived label is only a noisy proxy for physical task progress. In contact-rich manipulation, a robot may make progress and then lose it through slips, failed grasps, or partial u

arXiv (cs.AI) · Jul 14, 2026
4 min
Energy-Based Physics-Informed Form Finding for Clustered Tensegrity StructuresResearch

Energy-Based Physics-Informed Form Finding for Clustered Tensegrity Structures

Tensegrity form-finding and physical property prediction are fundamental inverse problems in structural mechanics, which aim to determine equilibrium configurations and internal force distributions. These problems are challenging due to strong nonlinearity arising from the coupling between geometry and forces, the need to ensure structural stability, and the enforcement of constraints such as boundary conditions and symmetry. Moreover, traditional methods often lack robustness to noise and outliers. This paper proposes an energy-based learning framework for clustered tensegrity form

arXiv (cs.LG) · Jul 14, 2026
3 min
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation StudyResearch

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning. We present Pythia, a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, Pythia keeps clinical notes on local infrastructure and selects prompt

arXiv (cs.AI) · Jul 14, 2026
4 min
Deep4ge: DNN Training Trajectories for Fault Detection and DiagnosisResearch

Deep4ge: DNN Training Trajectories for Fault Detection and Diagnosis

Deep learning systems often fail due to subtle implementation faults that alter training behavior. Recent work has studied how to detect and diagnose such failures from changes observed across training epochs. However, the software engineering community still lacks a public dataset of per-epoch training runs with documented fault history, feature extraction details, and clear reuse support for fault detection and diagnosis tasks. We present Deep4ge, a controlled benchmark of 14,227 training runs generated from 59 adapted TensorFlow/Keras deep neural network (DNN) programs collected f

arXiv (cs.LG) · Jul 14, 2026
3 min
Toward Localizing and Repairing Bias in Transformer Attention HeadsResearch

Toward Localizing and Repairing Bias in Transformer Attention Heads

Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existing fairness testing and repair methods largely operate at the input-output or retraining level, while recent work suggests that bias-related behavior can concentrate in a small set of attention heads. This paper studies whether attention heads can be localized and repaired through a targeted inference-time intervention. We introduce ROBIN, a white-box head-level fairness debugging method that ranks attention heads using sensitivi

arXiv (cs.LG) · Jul 14, 2026
3 min
Toward Localizing and Repairing Bias in Transformer Attention HeadsResearch

Toward Localizing and Repairing Bias in Transformer Attention Heads

Transformer language models are increasingly used as software components, yet biased outputs remain difficult to localize and repair inside the model. Existi...

Papers with Code · Jul 14, 2026
1 min
Unveiling Complex Collective Behaviors from Simple RewardsResearch

Unveiling Complex Collective Behaviors from Simple Rewards

Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic analysis, limiting multi-robot applications. Furthermore, complex swarm behaviors can surprisingly emerge from simple rewards without explicit aggregation incentives. Unveiling the mechanisms behind this emergence is critical, but the disconnection between simple rewards and collective behaviors exacerbates interpretability challenges. This paper aims to reveal the hidden mechanisms in this process. We propose a two-stage EEC (\LinkIII) e

arXiv (cs.AI) · Jul 14, 2026
4 min
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart GenerationResearch

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures char

arXiv (cs.AI) · Jul 14, 2026
3 min
Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage ControlResearch

Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control

Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well requires planning hours ahead under storage constraints. Model predictive control (MPC) and reinforcement learning are difficult to scale across buildings. This study instead adapts an open-weight reasoning model through reinforcement learning with verifiable rewards (RLVR). We convert exact offline dynamic-programming (DP) action values into dense rewards for every candidate action. Using only 30 training prompts, reinforcement fine-tu

arXiv (cs.LG) · Jul 14, 2026
4 min
Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature SensitivityResearch

Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity

Unconventional computing systems must demonstrate robust performance under real-world environmental conditions to enable practical deployments. We have recently proposed superparamagnetic nanodot ensembles driven by strain-induced magnetoelectric coupling as exciting candidates for use as ultra-low energy consumption reservoir computing substrates. However, because their dynamics are governed by thermal activation effects, these systems are intrinsically sensitive to ambient temperature fluctuations, leading to degraded task performance when operated outside the temperature range use

arXiv (cs.AI) · Jul 14, 2026
4 min
ANGLE: Angular Neural Generative Learning via EngressionResearch

ANGLE: Angular Neural Generative Learning via Engression

Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geology, and meteorology. Traditional regression targets the conditional mean, which is often geometrically misleading for circular responses under multimodal, skewed, or asymmetric data structures. To address these limitations, a lightweight deep generative framework, namely ANGLE, is introduced for non-parametric distributional regression on the circle. The full conditional distribution of an angular response, given Euclidean and circular covariates, is learned through a genera

arXiv (cs.LG) · Jul 14, 2026
3 min
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference TechniquesResearch

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems fr

arXiv (cs.AI) · Jul 14, 2026
3 min
Solution of the Hempel's statistical ambiguity problem and Causal AIResearch

Solution of the Hempel's statistical ambiguity problem and Causal AI

This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity (RMS) for the statistical laws used in the inference. An analysis of the RMS refinements made by Wesley Salmon, Alberto Coffa, and James Fetzer led to the following definition of maximally specific statistical laws: "the lawlike premises of an adequate explanation must specify all and only those properties whose pr

arXiv (cs.AI) · Jul 14, 2026
4 min
Solution of the Hempel's statistical ambiguity problem and Causal AIResearch

Solution of the Hempel's statistical ambiguity problem and Causal AI

This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are d...

Papers with Code · Jul 14, 2026
2 min
Human-AI Agent Interaction as a Neuroplastic Training EnvironmentResearch

Human-AI Agent Interaction as a Neuroplastic Training Environment

Interaction with AI agents has become one of the most frequent activities of everyday digital life. Whether conversing with an assistant, working with a coding copilot, or generating images, the interaction follows a common iterative loop: a request is issued, a result returned, appraised, and the request revised. We observe that this loop is a high-frequency stream of contact events -- moments at which a result meets a person and a conditioned response may fire before deliberate appraisal -- making everyday agent interaction an unrecognised neuroplastic training environment. When a

arXiv (cs.AI) · Jul 14, 2026
4 min
Visual Access Boundaries in Vision-Language Model ReasoningResearch

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended wh...

Papers with Code · Jul 14, 2026
2 min
Visual Access Boundaries in Vision-Language Model ReasoningResearch

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass. We introduce Visual Access Sweep, a causal intervention that masks attention from generated-token queries to image-token keys along layer depth and generation time, and define the Visual Access Boundary (VAB) as the minimal

arXiv (cs.AI) · Jul 14, 2026
4 min
PixelLoop: Shortcut Topological Navigation with Pixel-Level LoopsResearch

PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops

Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in purely topological representations has received relatively little attention. Importantly, loop closure over topological maps is distinct from loop closure over globally referenced trajectories and metric maps. Building on recent denser topologies grounded in pixel-level, relative 3D geometry, we propose PixelLoop which introduces loop closures directly in pixel space. Unlike sparse image-level edges or pose-graph corrections in SLAM, our pixel-level c

arXiv (cs.AI) · Jul 14, 2026
4 min
Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVsResearch

Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVs

This study introduces a unified control framework for fixed-wing unmanned aerial vehicles (UAVs) fitted with a pan-tilt (PT) camera, intended to perform an end-to-end mission spanning from initial target detection to accurate terminal engagement. The proposed system employs a three-phase strategy: a vision-based target acquisition phase, an NMPC-based tracking phase, and a terminal guidance phase. During tracking, the framework uses an Unscented Kalman Filter (UKF) to fuse YOLO-based visual detections with inertial measurements, enabling robust target state estimation under unknown d

arXiv (cs.AI) · Jul 14, 2026
3 min
The One-Word Census: Answer-Choice Conformity Across 44 Language ModelsResearch

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the time. We characterize this convergence with a deliberately minimal instrument: 31 single-turn prompts, each naming a category with many valid one-word answers ("Name a tree."), asked four times per model with no system prompt. Analysis is exact-match on normalized tokens -- no embeddings, no judge -- at about a dollar per model. That models

arXiv (cs.AI) · Jul 14, 2026
4 min
AVQ-Attention: Adaptive Vector-Quantized AttentionResearch

AVQ-Attention: Adaptive Vector-Quantized Attention

The $\mathcal{O}(N^2)$ complexity of attention over $N$ tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to $\mathcal{O}(MN)$ by representing keys with $M$ codewords, but applies uniform codebook capacity regardless of where attention mass concentrates: high-attention regions of key space may be coarsely approximated while low-attention regions waste representational capacity. We propose Adaptive Vector-Quantized (AVQ) Attention, which adaptively allocates codebook capacity based on attention importance. Starting from a sma

arXiv (cs.LG) · Jul 14, 2026
3 min
Directional Constraints for Efficient Exploration in Safe Reinforcement LearningResearch

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Safe Reinforcement Learning methods address this requirement by enforcing safety constraints. Nevertheless, learning under constraints often reduces learning speed and could lead to suboptimal task performance, as the agent must solve a more complex constrained optimization problem compared to unconstrained set

arXiv (cs.LG) · Jul 14, 2026
4 min
Learning-enabled Acceleration of Scenario-based Model Predictive ControlResearch

Learning-enabled Acceleration of Scenario-based Model Predictive Control

Scenario-based model predictive control (SBMPC) is a variant of model predictive control (MPC) that explicitly accounts for uncertainty by optimizing control actions over multiple predicted scenarios. However, its computational complexity increases rapidly with the number of scenarios and prediction horizon, limiting is applicability to real-time planning and control. This paper presents a learning-accelerated Alternating Direction Method of Multipliers (ADMM) algorithm for efficiently solving SBMPC problems by leveraging parallel computing and Moreau envelope learning, while maintai

arXiv (cs.LG) · Jul 14, 2026
3 min
Learning Mechanistic Reasoning for Chemical Reactions with Large Language ModelsResearch

Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models

Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations. Learning the mechanism logic is therefore essential for enhancing the fundamental chemical intelligence of large language models (LLMs). The stepwise deduction of reaction mechanism aligns naturally with the reasoning paradigms of reasoning LLMs. However, current chemical LLMs primarily emphasize coarse-grained name reactions for product prediction and retrosynthesis, often leading to physical inconsistencies and hallucinations. In contrast, specialized small-scale

arXiv (cs.LG) · Jul 14, 2026
3 min
Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy CoordinationResearch

Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination

Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation methods such as FedAvg do not account for system-level constraints, often leading to unsafe global behavior. In this work, we study constraint-aware aggregation for federated reinforcement learning in distributed energy coordination. We propose aggregation rules that incorporate both local performance and estimated constraint violation into the server-side update. Among these, a simple penalty-based rule, $w_i \propto R_i - αV_i$, cons

arXiv (cs.LG) · Jul 14, 2026
4 min
What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in GrokkingResearch

What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking

Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learnable prior built from the wrong feature family (magnitude bands) blocks generalization like a random partition (1/15 vs 0/20 grok; $p=0.43$ between them), confirming the companion's prediction that priors act at the level of the circuit's features. Supervision: a fully label-free invariance prior -- positives are commuted pairs $(a,b)\si

arXiv (cs.LG) · Jul 14, 2026
4 min
Lightweight Multi-Scale Anomaly Detection for Resource-Constrained Edge DevicesResearch

Lightweight Multi-Scale Anomaly Detection for Resource-Constrained Edge Devices

Time-series anomaly detection is increasingly important in IoT systems, sensor networks, and edge monitoring applications, where models must operate under st...

Papers with Code · Jul 14, 2026
2 min
Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear PredictionResearch

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent app...

Papers with Code · Jul 14, 2026
1 min
AI in Indian Education: Atal Innovation Mission and Google launch ATL SaathiResearch

AI in Indian Education: Atal Innovation Mission and Google launch ATL Saathi

Atal Innovation Mission launches ATL Saathi, a Gemini powered AI assistant empowering India's educators to nurture the next generation of innovators.

Google DeepMind · Jul 14, 2026
5 min
Differentiable Clone-Structured Causal Graphs for End-to-End Cognitive Map Learning from Image SequencesResearch

Differentiable Clone-Structured Causal Graphs for End-to-End Cognitive Map Learning from Image Sequences

How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when natural...

Papers with Code · Jul 14, 2026
2 min
What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM EvaluationResearch

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1. The ...

Papers with Code · Jul 14, 2026
2 min
Toward Trustworthy Autonomous Science: A Two-Year Community RoadmapResearch

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five crit...

Papers with Code · Jul 13, 2026
2 min
Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training DataResearch

Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data

Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functions far simpler than their parameter counts suggest, but it is challenging to construct codes that realize this simplicity. Parameter-based methods such as quantization produce code lengths that scale with model size, insensitive to how much information the parameters store. Prequential coding bypasses this issue by compressing the training trajectory, but codes the exact data sequence re

arXiv (cs.LG) · Jul 13, 2026
4 min
Metacognition in LLMs: Foundations, Progress, and OpportunitiesResearch

Metacognition in LLMs: Foundations, Progress, and Opportunities

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper

arXiv (cs.AI) · Jul 13, 2026
4 min
Invariant Learning Dynamics of Transformers in Inductive Reasoning TasksResearch

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning. In this class, we theoretically prove that the training dynamics of attention models can be confined to a highly interpretable, low-dimensional invariant manifold. On this manifold, the learning dyna

arXiv (cs.AI) · Jul 13, 2026
4 min
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous ManipulationResearch

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies v...

Papers with Code · Jul 13, 2026
1 min
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous ManipulationResearch

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object mo

arXiv (cs.AI) · Jul 13, 2026
4 min
A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification ProtocolResearch

A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. W

arXiv (cs.LG) · Jul 13, 2026
3 min
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiasResearch

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight activation manifold while biased inputs are displaced along a low-dimensional, type-specific s

arXiv (cs.AI) · Jul 13, 2026
4 min
Evidence-Backed Video Question AnsweringResearch

Evidence-Backed Video Question Answering

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation maskle

arXiv (cs.AI) · Jul 13, 2026
3 min
Input-Aware Dynamic Backdoor Attack Against Quantum Neural NetworksResearch

Input-Aware Dynamic Backdoor Attack Against Quantum Neural Networks

Quantum Neural Networks (QNNs) are a promising framework for quantum machine learning on near-term quantum devices, but their security risks remain insufficiently understood. Studies have shown that QNNs are vulnerable to backdoor attacks, yet existing quantum backdoors mostly rely on a fixed trigger shared by all poisoned inputs. This fixed-trigger design is a major weakness because many defenses detect or weaken the repeated patterns such triggers leave in data representations. Although input-aware dynamic backdoors have been studied in classical neural networks, transferring them

arXiv (cs.LG) · Jul 13, 2026
4 min
LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training EnvironmentsResearch

LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components. Rather than assuming a fixed fusion structure, the framework first integrates more closely related modalities and then incorporates additional heterogeneous modalities, supporting scalable adaptation across data

arXiv (cs.AI) · Jul 13, 2026
3 min
Transformer-Guided Swarm Intelligence for Frugal Neural Architecture SearchResearch

Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search

Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive computational resources, often measured in thousands of GPU-days. In this paper, we propose a frugal and memetic NAS framework designed to democratize architecture design on consumer-grade hardware. Our approach combines the global macro-search capabilities of an autoregressive Transformer controller, trained via Reinforcement Learning (RL), with the local micro-exploitation of an Artificial Bee Colony (ABC) algorithm. To prevent premature convergence during the RL pha

arXiv (cs.AI) · Jul 13, 2026
4 min
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling AgentsResearch

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces diverse, visually grounded scenarios through information-flow-guided planning and multi-s

arXiv (cs.AI) · Jul 13, 2026
4 min
Relaxing Faithfulness with Intervention-Only Causal DiscoveryResearch

Relaxing Faithfulness with Intervention-Only Causal Discovery

Causal discovery algorithms learn a network that describes the causal dependencies among random variables. A common workflow involves first utilizing conditional independence properties on observational data to determine partially directed causal relationships, then applying interventions to orient the unknown causal directions. A critical assumption for the first step is faithfulness: a requirement that causally linked variables exhibit statistical dependence. Many natural systems include buffering and stabilizing pathways that cancel out to achieve systemic robustness. This cancell

arXiv (cs.LG) · Jul 13, 2026
3 min
Introducing Human-Centeredness in AI-Assisted LexicographyResearch

Introducing Human-Centeredness in AI-Assisted Lexicography

This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers significant opportunities to enhance lexicographic work, it also raises concerns regarding the future role of lexicographers and the preservation of linguistic and cultural diversity. Drawing on HCAI principles and previous applications in other language professions, the paper identifies four interrelated dimensions through which AI integration in lexicography can be understood and critically examined: the augmented lexicographer, the sociotechnical c

arXiv (cs.AI) · Jul 13, 2026
3 min
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language ModelsResearch

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective inference-time intervention, yet most existing methods intervene only after the audio encoder and operate at a relatively coarse granularity. The encoder itself, where acoustic information is first extracted from the waveform, remains largely unexplored, especially at the level of individual neurons. We introduce IAAN, Identifying and Am

arXiv (cs.AI) · Jul 13, 2026
4 min
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio DescriptionResearch

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on loca

arXiv (cs.AI) · Jul 13, 2026
4 min
An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It RevealsResearch

An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals

Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism. We give an exact instrument for measuring how a trained model uses these modes. Because the state matrix is diagonal, each channel's output decomposes exactly into per-mode contributions, and a per-(layer, channel, window) Gram tensor yields the exact output error of dropping any subset of modes, offline, at any budget. Validated against the reference implementation to a relative error of $2.3\times10^{-7}$ on the Mamba-1 famil

arXiv (cs.LG) · Jul 13, 2026
4 min
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework ProposalResearch

Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal

Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although prior research has proposed frameworks, patterns, and user-centered approaches to support explainability, there is limited empirical understanding of how existing Requirements Engineering (RE) practices support explainability requirements across the RE lifecycle, especially in an industrial context. This paper reports early findings from an ongoing industry-based study investigating how explainability requirements are elicited, specified, and vali

arXiv (cs.AI) · Jul 13, 2026
4 min
From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASPResearch

From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP

A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss lands

arXiv (cs.LG) · Jul 13, 2026
3 min
From Global to Factor-Wise Expert Composition in Discrete Diffusion ModelsResearch

From Global to Factor-Wise Expert Composition in Discrete Diffusion Models

Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithically and ignoring the potential spatial or functional specializations of different experts.

arXiv (cs.LG) · Jul 13, 2026
3 min
Paradoxes of Game Theoretic Equilibria and Price of AnarchyResearch

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing multi-agent learning to static equilibrium and black-box regret analysis obscures underlying dynamic disequilibrium and game theoretic bounds. First, interior Nash equilibria lack $C^1$ vector field information, meaning agents cannot distinguish aligned from strictly opposing incentives. Inheriting this geome

arXiv (cs.LG) · Jul 13, 2026
4 min
When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent SystemsResearch

When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems

As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is \emph{local benignness}. No fragment carries the harm, and what is left loo

arXiv (cs.LG) · Jul 13, 2026
4 min
Playful AI in Professional Email: A Field Experiment on Tone and Recipient EngagementResearch

Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement

Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes how recipients actually behave, and through what channel, remains unknown. Here, in a randomized crossover field experiment, 121 employees across six companies sent work emails under three conditions over three weeks: unaided writing, GPT-5 rewriting in a playful tone, and GPT-5 rewriting in a professional tone. Across 16,880 emails, playful editing increased emotional positivity (B=+0.068, p<0.001), and professional editing decreased it (B=-0.041, p<0.001), yet neither

arXiv (cs.AI) · Jul 13, 2026
3 min
HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NASResearch

HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NAS

With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential. These methods rely on real feedback from the target hardware to tailor DNN architectures for efficient deployment. While the search can be parallelized, latency measurements via hardware-in-the-loop (HIL) remain a bottleneck due to their sequential nature. Recent approaches use latency predictors to replace costly HIL feedback, but challenges persist: (1) platform-spec

arXiv (cs.LG) · Jul 13, 2026
4 min
NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force PerceptionResearch

NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception

Differentiable simulators have advanced policy learning and model-based control, yet actuator dynamics remain an important source of sim-to-real error. This is particularly acute on low-cost platforms, where the linear current-to-torque relation $τ= K_tI$ becomes unreliable during commanded-target tracking because of friction, hysteresis, backlash, and thermal effects. We present NeuralActuator, a neural actuator model that jointly predicts (i) a simulator-equivalent generalized-effort surrogate for trajectory propagation on low-cost servo platforms, (ii) external force with a contac

arXiv (cs.LG) · Jul 13, 2026
4 min
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module FactoriesResearch

Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories

Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop. A major complication is decisive: long post-operation time-lags caused by concrete curing, watertightness ponding tests, and paint drying, during which a module is blocked while its workstation stays free. On benchmark instances grounded in an official national prefabrication guidebook, these lags inflate even the optimal reference makespan by about 67% on average, and ignoring them at decision time, then repairing to feasibility, i

arXiv (cs.AI) · Jul 13, 2026
4 min
Active Offline-to-Online Reinforcement LearningResearch

Active Offline-to-Online Reinforcement Learning

Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the policy with the highest estimated value. However, as in offline pretraining, fine-tuning per

arXiv (cs.AI) · Jul 13, 2026
4 min
CatRetriever: Contrastive Representation Learning for Slab-to-Bulk Retrieval in Generative Catalyst DiscoveryResearch

CatRetriever: Contrastive Representation Learning for Slab-to-Bulk Retrieval in Generative Catalyst Discovery

Inverse design is an emerging data-driven paradigm for efficiently navigating vast chemical spaces to discover new materials with targeted properties, and in the context of heterogeneous catalysis, surface generative models have recently advanced this goal by directly generating catalyst surface-adsorbate structures. However, these models typically operate at the slab level and do not provide the corresponding parent bulk structure, making it difficult to assess bulk-dependent properties such as formation energy, surface energy, crystallographic symmetry, and synthesizability. Here,

arXiv (cs.LG) · Jul 13, 2026
4 min
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based MemoryResearch

An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory

Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated conversational scams. It also introduces ConScamBench-278, an initial public multi-categor

arXiv (cs.AI) · Jul 13, 2026
4 min
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice ConversionResearch

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark ei

arXiv (cs.AI) · Jul 13, 2026
3 min
$\mathtt{Q^2SAR}$: overcoming classical bottlenecks in drug discovery via quantum multiple kernel learningResearch

$\mathtt{Q^2SAR}$: overcoming classical bottlenecks in drug discovery via quantum multiple kernel learning

Quantitative Structure-Activity Relationship ($\mathtt{QSAR}$) modeling is a foundational computational methodology in early-stage drug discovery, heavily relied upon for predicting compound toxicity, bioavailability, and therapeutic potential. However, classical methods often struggle to effectively map the highly complex, non-linear, and high-dimensional interactions inherent in molecular data, leading to reduced predictive accuracy and costly late-stage clinical failures. In this paper, we present a Quantum Multiple Kernel Learning ($\mathtt{QMKL}$) framework, dubbed Next-Gen $\ma

arXiv (cs.LG) · Jul 13, 2026
4 min
Agent Hacks Agent: Autoresearch for Production-Agent Red-TeamingResearch

Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming

Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserve artifacts such as benchmarks, payloads, or attack programs, which record where attacks succeed but not the enabling conditions behind unsafe agent behavior. We study automated red-teaming for production LLM agents using one agentic research environment to discover reusable vulnerability kno

arXiv (cs.AI) · Jul 13, 2026
4 min
Think Through a Bottleneck: Hourglass Reasoning for Rigorous InductionResearch

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its inferred rule does little on its own. What actually matters is a structurally enforced isolation between reasoning stages, so that information can only pass between them as a compressed symbolic state. We introduce \textbf{Hourglass reasoning}, which enforces strict context isolation between reasoning stages. The frozen LLM acts as a meta-constructor, building for each task a symbolic encoder--decoder: an Induction module compresses the support ex

arXiv (cs.AI) · Jul 13, 2026
4 min
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical IntelligenceResearch

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems expose limited interfaces for reuse and evaluation. We review the evolution toward WAMs and

arXiv (cs.AI) · Jul 13, 2026
4 min
Self-Healing Visual Recovery for Autonomous Ground Vehicles Using Camera-Only Visual OdometryResearch

Self-Healing Visual Recovery for Autonomous Ground Vehicles Using Camera-Only Visual Odometry

Low-cost unmanned ground vehicles are often used in indoor places like warehouses, inspection corridors, and farm rows, where painted floor lines guide the robot. Line following is useful because it only needs one camera and little computing power, but it can fail when the line is blocked or turns sharply and goes out of view. Sensor-rich platforms tolerate this through hardware redundancy (LiDAR, GPS, multiple cameras), but camera-only systems must recover at runtime with no additional infrastructure. This paper presents a lightweight, two-stage recovery approach that restores guide

arXiv (cs.LG) · Jul 13, 2026
4 min
Diversified Multinomial Logit Contextual BanditsResearch

Diversified Multinomial Logit Contextual Bandits

Existing contextual multinomial logit (MNL) bandits model relevance-driven choice but ignore the potential benefits of within-assortment diversity, while submodular/combinatorial bandits encode diversity in rewards but lack structured choice probabilities. We bridge this gap with the $\textit{diversified multinomial logit}$ (DMNL) contextual bandit, which augments MNL choice probabilities with a generally submodular diversity function, thereby formalizing the relevance--diversity trade-off within a single model. Incorporating diversity renders exact MNL assortment optimization intrac

arXiv (cs.LG) · Jul 13, 2026
3 min
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLMResearch

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-source modular GraphRAG engine, addresses this by separating extraction from consolidation: entities and relations pass through two-stage typed extraction, DBSCAN-backed deduplication, LLM summarization, and Leiden community detection. A key insight motivates a compact extractor: the skills an in-pipeline LLM needs - comprehension, extraction,

arXiv (cs.AI) · Jul 13, 2026
4 min
A multi-scale feature enhanced graph neural network for fluid dynamics prediction in complex geometriesResearch

A multi-scale feature enhanced graph neural network for fluid dynamics prediction in complex geometries

Industrial design in fields such as vehicle and aerospace engineering often relies on large-scale numerical simulations to evaluate fluid dynamics performance, which can incur substantial computational costs. Deep neural networks have shown promise in improving simulation efficiency, especially graph neural networks (GNNs), which demonstrate great potential due to their flexibility with unstructured data. However, GNNs face challenges when dealing with tasks involving complex geometries and large-scale meshes. In this paper, we propose the Multi-scale Feature Enhanced Graph Neural Ne

arXiv (cs.LG) · Jul 13, 2026
4 min
How to Tame Grokking: Representation Geometry as a Control SignalResearch

How to Tame Grokking: Representation Geometry as a Control Signal

Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization. Despite extensive recent study, the factors influencing the emergence and timing of grokking remain incompletely understood. We investigate the relationship between representation geometry and delayed generalization. We find that dimensionality collapse consistently precedes the onset of grokking in all evaluated settings. Motivated by these observations, we introduce Geometric Dimensionality Regularization (GeomDR), a simple sp

arXiv (cs.LG) · Jul 13, 2026
3 min
Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohortsResearch

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data. Left unaddressed, these barriers prevent reliable disease modelling and hinder effective clinical evaluation. Conventional imputation strategies introduce systematic bias, distort inter-feature relationships, and yield overconfident predictions, limitations especially consequential in diagnostic settings. Here, we propose NITROGEN, an imputation-free transformer that jointly models within-patient feature dependenc

arXiv (cs.LG) · Jul 13, 2026
4 min
Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile ForecastersResearch

Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters

Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We

arXiv (cs.LG) · Jul 13, 2026
4 min
Closing the Loop: An Access-Control Architecture for Automated, Anomaly-Driven Network Revocation in IoT DeploymentsResearch

Closing the Loop: An Access-Control Architecture for Automated, Anomaly-Driven Network Revocation in IoT Deployments

Network-based anomaly detection for IoT devices has matured to the point of reporting strong detection accuracy, yet most published systems stop at raising an alert and leave the question of automated enforcement to future work or to a programmable data plane that few real networks operate. This paper presents an access-control architecture that closes that loop using only standard, already-deployed protocols. Devices authenticate via IEEE 802.1X with EAP-TLS, and a RADIUS server acts as a continuous policy decision point capable of evicting an active session via a Change-of-Authoriz

arXiv (cs.AI) · Jul 13, 2026
4 min
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation ModelResearch

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundati

arXiv (cs.AI) · Jul 13, 2026
4 min
Reproducing human biases in route choice using large language models: Toward scalable behavioral modelingResearch

Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely on surveys and controlled experiments to calibrate CPT parameters, yet these methods are d

arXiv (cs.AI) · Jul 13, 2026
4 min
Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming PatternsResearch

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and

arXiv (cs.AI) · Jul 13, 2026
4 min
Relational Positioning as a Measurable Risk Object: History-Carried Lock-in and Self-Confabulation in Multi-Turn Human-AI DialogueResearch

Relational Positioning as a Measurable Risk Object: History-Carried Lock-in and Self-Confabulation in Multi-Turn Human-AI Dialogue

In long, multi-turn dialogue a large language model maintains an implicit relational stance toward the user, spanning from "push the user toward real-world o...

Papers with Code · Jul 13, 2026
2 min
Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference AlignmentResearch

Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment

Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic chara...

Papers with Code · Jul 13, 2026
1 min
Inter-Stop Energy Prediction and Causal Driver Quantification for Dual-Source Trolleybuses via a Time-Aware Tabular Deep Learning ArchitectureResearch

Inter-Stop Energy Prediction and Causal Driver Quantification for Dual-Source Trolleybuses via a Time-Aware Tabular Deep Learning Architecture

Dual-source trolleybuses alternate between overhead catenary supply and on-board battery operation, creating energy-use patterns driven by route attributes, ...

Papers with Code · Jul 13, 2026
2 min
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMsResearch

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evalu...

Papers with Code · Jul 13, 2026
2 min
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning AbilitiesResearch

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete,...

Papers with Code · Jul 13, 2026
1 min
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical ResultsResearch

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible res...

Papers with Code · Jul 13, 2026
2 min
LP Mining with LP2Graph: A Use Case for Railway ReschedulingResearch

LP Mining with LP2Graph: A Use Case for Railway Rescheduling

Like many optimization-driven domains, railway rescheduling relies on Mixed-Integer Linear Programming (MILP), yet the field's modeling knowledge is scattere...

Papers with Code · Jul 13, 2026
2 min
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent EvaluationResearch

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, conte...

Papers with Code · Jul 13, 2026
2 min
MusicMark: A Robust Generative Watermarking Framework for Music GenerationResearch

MusicMark: A Robust Generative Watermarking Framework for Music Generation

AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attribution. However, ...

Papers with Code · Jul 13, 2026
2 min
The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium ReasoningResearch

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should ...

Papers with Code · Jul 13, 2026
2 min
Generative Chinese Statute RetrievalResearch

Generative Chinese Statute Retrieval

Statute retrieval is a fundamental task in legal information retrieval, yet existing approaches struggle to bridge the gap between colloquial legal queries a...

Papers with Code · Jul 13, 2026
1 min
Why Low-Light Cameras Go Color Blind: Removing Color Bias in Raw DenoisingResearch

Why Low-Light Cameras Go Color Blind: Removing Color Bias in Raw Denoising

Raw images inherently suffer from noise due to the stochastic nature of light and sensor hardware imperfections. As real photon counts fall, the ratio of thi...

Papers with Code · Jul 13, 2026
2 min
MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit CalibrationResearch

MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit Calibration

Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A funda...

Papers with Code · Jul 13, 2026
2 min
From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground TruthResearch

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the c...

Papers with Code · Jul 13, 2026
2 min
Enhanced Byzantine-Robust Federated Learning Via Truncated-Quadratic Loss for Heterogeneous DataResearch

Enhanced Byzantine-Robust Federated Learning Via Truncated-Quadratic Loss for Heterogeneous Data

Federated learning distributes data among $n$ clients, making it vulnerable to malicious attacks and data heterogeneity, which together pose challenges for r...

Papers with Code · Jul 13, 2026
2 min
CGS: Configurable Graph Summarization with Bounded Neighborhood Loss and Query SupportResearch

CGS: Configurable Graph Summarization with Bounded Neighborhood Loss and Query Support

Given a large graph, how to generate a compact summary graph that is configurable by the user and supports multiple graph queries with either no loss or with...

Papers with Code · Jul 13, 2026
2 min
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation DefectsResearch

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automat...

Papers with Code · Jul 12, 2026
2 min
LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous DrivingResearch

LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving

Autonomous driving requires long-horizon closedloop decision making in dynamic traffic environments. Latent world models offer an effective framework for thi...

Papers with Code · Jul 12, 2026
2 min
Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDFResearch

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed a...

Papers with Code · Jul 12, 2026
2 min
Large language model agents accelerate inverse design of metal-organic frameworks for gas separationResearch

Large language model agents accelerate inverse design of metal-organic frameworks for gas separation

Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design dif...

Papers with Code · Jul 12, 2026
2 min
PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal AssessmentResearch

PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment

Preparing for job interviews is important for securing desired positions, yet realistic practice remains difficult to access: real interviews are infrequent,...

Papers with Code · Jul 11, 2026
2 min
Inverse-IMPRESSION: A Graph-based Platform for Molecular Structure Elucidation from Experimental NMR Spectroscopic PropertiesResearch

Inverse-IMPRESSION: A Graph-based Platform for Molecular Structure Elucidation from Experimental NMR Spectroscopic Properties

Here, we present a platform built on our inverted Graph Transformer Network, IMPRESSION-G2, which can accurately and rapidly reconstruct molecular bonding di...

Papers with Code · Jul 10, 2026
2 min
PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal SynthesisResearch

PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis

Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achieving a state-of-the-art area under the receiver operating characteristic curve (AUC) of approximately 0.70 on the DREAM database (Wong et al., 2025, Nature Communications). We introduce PHINN-EEG (Persistent Homology Inspired Neural Network for EEG), the first topological time-series framework for dream mentation analysis. Using sliding-window Takens delay embeddings and Vietoris-Rips filtrations on multichannel pre-awakening EEG epochs, we extract D

arXiv (cs.AI) · Jul 10, 2026
4 min
Scalable Visual Pretraining for Language IntelligenceResearch

Scalable Visual Pretraining for Language Intelligence

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must

arXiv (cs.AI) · Jul 10, 2026
3 min
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI ModelsResearch

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large La

arXiv (cs.AI) · Jul 10, 2026
4 min
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentsResearch

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions of Large Language Model (LLM) agents have demonstrated promise in penetration testing and Capture-the-Flag (CTF) environments, their application to IoT specific vulnerabilities remains unexplored. This paper presents an autonomous multi-agent framework, referred to as Vulnerability EXploitation using AI Agents (VEXAIoT), for vulnerability discovery a

arXiv (cs.AI) · Jul 10, 2026
4 min
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI AgentsResearch

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a ne...

Papers with Code · Jul 10, 2026
2 min
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AIResearch

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weight

arXiv (cs.AI) · Jul 10, 2026
3 min
Deep Gaussian Processes on Directed Acyclic GraphsResearch

Deep Gaussian Processes on Directed Acyclic Graphs

Many real-world processes can be represented as compositions of functions along a directed acyclic graph (DAG). In causal modelling, these correspond to the underlying mechanisms; in engineering, to multiple fidelity levels; and in gene-regulatory networks, to transcription factors. These functions are partially observed across the DAG, with noisy and heterogeneously sampled measurements, posing significant challenges for reconstruction, uncertainty propagation, and inference. To tackle these challenges, we place priors over functions and naturally arrive at Deep Gaussian Processes o

arXiv (cs.LG) · Jul 10, 2026
4 min
Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly DetectionResearch

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome this without distortive data resampling, we propose the Semantic Pareto-DQN, a multi-objective reinforcement learning framework. Our approach synthesizes heterogeneous transaction features into cohesive natural-language narratives, encoded by large language models, thereby producing a robust, scale-invariant state representation.

arXiv (cs.AI) · Jul 10, 2026
3 min
Lean-QIT: Towards a Formal Infrastructure for Quantum Information TheoryResearch

Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory

Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correction. Formalizing its coding theorems requires connecting finite-block protocols, analytic inequalities, and asymptotic limits within a unified machine-checked framework. Existing developments, however, lack a reusable operational layer that defines codes, error criteria, achievable rates, and capacities independently of their information-theoretic characterizations. In this work, we present LeanQIT,

arXiv (cs.AI) · Jul 10, 2026
3 min
4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene PerceptionResearch

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera methods mainly optimize detection, while dual-task systems usually decode boxes and occupancy with limited interaction. To address this gap and advance radar-based multi-task learning, we propose \method, a 4D radar-camera framework for 360$^\circ$ full-scene perception

arXiv (cs.AI) · Jul 10, 2026
4 min
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026Research

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, which require deciding when to answer under uncertainty, and Bonus questions, which emphasize accurate answer selection and human adoption. To address these differing objectives, we develop a

arXiv (cs.AI) · Jul 10, 2026
4 min
LLM for EDA in Front-End Design: Challenges and OpportunitiesResearch

LLM for EDA in Front-End Design: Challenges and Opportunities

As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great potential in Electronic Design Automation (EDA). Beyond specification understanding, LLMs show the potential to serve as a unified intelligent interface for hardware description language (HDL) generation, testbench construction, and design space exploration. The rise of agentic AI, represented by pioneering systems such as OpenClaw, offers a strategic roadmap for the next generation EDA. From thi

arXiv (cs.LG) · Jul 10, 2026
4 min
Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task AllocationResearch

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of expert models or tools, while overlooking critical factors such as performance variability and cost efficiency among functionally similar alternatives. To address this, we propose Agora, a framework that introduces an incentive-compatible auction mechanism for dynamically allocating tasks to expert models and tools. By treating

arXiv (cs.AI) · Jul 10, 2026
3 min
PAC-ACT: Post-training Actor-Critic for Action Chunking TransformersResearch

PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost, while vision-action chunking policies are more suitable for real-time industrial control. However, these policies are usually trained by behavior cloning and suffer from distribution shift in contact-rich tasks. This paper proposes PAC-ACT, a reinforcement-learning post-training framework for pretrained Action Chunking Transformer polic

arXiv (cs.AI) · Jul 10, 2026
3 min
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI SystemsResearch

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classification Framework, a structured, repeatable instrument that can be applied to seven types of agentic AI systems and is grounded in foundational pre-existing AI governance frameworks. At the core of the framework is a twelve-dimension scoring rubric that robustly quantifies the risk. This rubric is combined with other components, such as the GPA + IAT cl

arXiv (cs.AI) · Jul 10, 2026
4 min
Entropy-Constrained Machine Learning with Residual Data Augmentation for Modeling Chemical KineticsResearch

Entropy-Constrained Machine Learning with Residual Data Augmentation for Modeling Chemical Kinetics

We present a physics-constrained machine learning framework for accelerating the direct numerical simulation (DNS) of turbulent reacting flows. The model replaces the direct evaluation of detailed chemical source terms with a surrogate that predicts reaction rates from a reduced thermochemical state. To improve physical consistency, the second law of thermodynamics is incorporated as a training constraint by enforcing non-negative entropy generation, which restricts the evolution of the thermochemical state to physically admissible directions and improves stability during time integr

arXiv (cs.LG) · Jul 10, 2026
3 min
Knowledge Graphs and Explainable AI as Complementary Resources for Urban MiningResearch

Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining

Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in which AI support must serve qualified auditors who remain accountable for the decisions taken. The relevant unit of value is not prediction accuracy alone, but the defensibility of the supported decisions: their legibility, plausibility, sourcing, and contestability. Explainable AI techniques and domain knowledge graphs each address parts of this requirement, and existing taxonomies have catalogued their integration. The literature is descriptively rich but structurally u

arXiv (cs.AI) · Jul 10, 2026
3 min
Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph ApproachResearch

Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach

We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse languages, totaling 160 conventional expressions, the large majority of which are idiomatic. Each expression is annotated with binary conceptual features (containment, concealment, emotional, social, etc.) derived from cognitive-linguistic theory, and pairwise Jaccard similarities define a weighted graph. Community detection reveals that idioms cluster by conceptual schema rather than by language, producing a structure consistent with cognitive-ling

arXiv (cs.AI) · Jul 10, 2026
4 min
Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary AlgorithmsResearch

Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary Algorithms

Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternatives available in financial markets. A hot area of research over the past few decades has been portfolio optimization that seeks to determine how much an investor should invest in which asset. Introducing real-world conditions to the optimization model turns the problem into an NP-hard one for whose solution exact methods become inefficient; hence, researchers have turned to evolutionary algorithms to approximate solutions. In this paper, strengthening strategies are p

arXiv (cs.AI) · Jul 10, 2026
4 min
TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language ModelsResearch

TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models

Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a

arXiv (cs.AI) · Jul 10, 2026
3 min
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AIResearch

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operatio

arXiv (cs.AI) · Jul 10, 2026
4 min
Graph-Regularized Low-Rank Matrix Completion by Variable ProjectionResearch

Graph-Regularized Low-Rank Matrix Completion by Variable Projection

We address the low-rank matrix completion problem by incorporating graph regularization into the existing Riemannian Trust-Region Matrix Completion (RTRMC) framework. The latter uses the geometry of the low-rank constraint to remodel the problem as an unconstrained optimization problem on a single Grassmann manifold. Our approach, named Graph-Regularized RTRMC (GR-RTRMC), exploits the inherent relationships between rows and columns of the matrix. By using these relationships, we aim to improve the accuracy and robustness of matrix completion, particularly in scenarios where the under

arXiv (cs.LG) · Jul 10, 2026
3 min
The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMsResearch

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal representations and verbalized outputs. Training simple probes on activations from four VLMs across five counting datasets reveals that nonlinear probes can reliably detect counting errors, suggesting that VLMs often encode the correct count even when they output the wrong answer. SVCCA analysis shows that probes trained on ground-truth counts and probes trained on model

arXiv (cs.LG) · Jul 10, 2026
4 min
CoCoT-EEG: Contrastive-Pretrained Multiscale Convolutional Transformer for EEG DecodingResearch

CoCoT-EEG: Contrastive-Pretrained Multiscale Convolutional Transformer for EEG Decoding

Self-supervised pretrained foundation models (FM) have shown early promise for non-invasive electroencephalogram (EEG) decoding applications. Many recent large-scale models converged on the approach of tokenizing raw EEG followed by masked reconstruction pretraining. However, this recipe has been shown to be suboptimal for data, like EEG, with high noise amplitude and information confined to limited dimensions such as narrow frequency bands. Building on this insight, we develop a novel contrastive-pretrained EEG model with multiscale temporal convolution input layers and Transformer

arXiv (cs.LG) · Jul 10, 2026
3 min
GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series ForecastingResearch

GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting

Time series forecasting requires models to capture diverse, often mutually exclusive, temporal dynamics, from smooth trend continuation to nonstationary drift and strict phase-aligned recurrence. While recent deep learning models have improved accuracy, they typically force these diverse patterns through a single computational backbone governed by fixed algorithmic inductive biases (e.g., self-attention or spectral filtering). This single-mechanism approach often struggles with the profound heterogeneity of real-world series, where different variables and forecast horizons necessitat

arXiv (cs.LG) · Jul 10, 2026
4 min
Statistically Undetectable Backdoors in Deep Neural NetworksResearch

Statistically Undetectable Backdoors in Deep Neural Networks

We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assump

arXiv (cs.LG) · Jul 10, 2026
3 min
TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse EcosystemsResearch

TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems

The emergence of metaverse platforms has created virtual economies that introduce new challenges related to fraud, bot activity, and illicit financial behavior. Despite growing interest in trustworthy metaverse analytics, existing datasets typically focus on user behavior, authentication, or financial transactions in isolation, limiting the development and reproducible evaluation of multimodal fraud detection methods. To address this gap, we present TSAI-MetaFraud, a multimodal, multi-task benchmark dataset for fraud analytics in virtual economies. TSAI-MetaFraud integrates behaviora

arXiv (cs.LG) · Jul 10, 2026
3 min
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level ExpertsResearch

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones. Here we present ALICE, a unified foundation model trained through multi-stage agglomerative distillation that sequentially distills eight vision-only, vision-language, and slide-level teacher models into dedicated modules of a single backbone. ALICE is pretrained on 24,985,184 tile-level pathology images and 155,604 high-resolution images, and evaluated across 21 task s

arXiv (cs.AI) · Jul 10, 2026
3 min
SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival PredictionResearch

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from demographics collected at intake to genomic profiling requiring specialized tissue analysis. Current multimodal survival methods either assume all modalities are available or passively handle missing data, but none actively reason about whether acquiring the next modality is justified for a given patient along this ordered workflow. We formulate this as a sequential de

arXiv (cs.AI) · Jul 10, 2026
4 min
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM InferenceResearch

Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference

Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood. Existing efficiency efforts focus predominantly on reducing visual tokens, implicitly treating visual processing as the dominant energy cost. We overturn this implicit assumption through the first systematic energy profiling of on-device VLM inference, spanning five models across three architecture families, four input resolutions, and two hardware platforms (NVIDIA RTX 3070 and Jetson Orin NX). Our analysis yields three findings. First, ave

arXiv (cs.AI) · Jul 10, 2026
4 min
Failure as a Process: An Anatomy of CLI Coding Agent TrajectoriesResearch

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories

Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making their reliability a growing concern. Existing empirical studies investigate why coding agents fail, yet they largely treat failure as a final outcome rather than a temporal process, providing limited insight into how failures emerge, evolve, and become unrecoverable. We present the first large-scale empirical study of CLI coding-agent failure trajectories, introducing a process-oriented framework that analyzes failure through its

arXiv (cs.AI) · Jul 10, 2026
4 min
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-VisibilityResearch

What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility

A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap. We demonstrate that VGGT implicitly encodes co-visibility as an emergent behavior: without any supervision for this task, its internal representations exhibit a clear hierarchical structure mirroring that of large language models, i.e. early layers build a 3D-aware scene representation, while late layers act as dedicated co-visibility reasoners. In particular, we identify layer L17 as

arXiv (cs.AI) · Jul 10, 2026
4 min
All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language ModelsResearch

All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models

Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not fundamental, but an artifact of treating explanation and prediction as separate objectives; when properly coupled, they become complementary, so that equipping a model to explain itself improves, rather than degrades, its accuracy. We introduce the Rashomon Explanation paradigm, which builds a set of faithful

arXiv (cs.AI) · Jul 10, 2026
4 min
Shared Selective Persistent Memory for Agentic LLM SystemsResearch

Shared Selective Persistent Memory for Agentic LLM Systems

Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made previous sessions productive. Naively persisting entire conversation histories is token-inefficient and counterproductive: irrelevant context degrades generation quality. We introduce shared selective persistent memory, an architecture that identifies and retains four categories of reusable context (task specifications, data schemas, tool configurati

arXiv (cs.AI) · Jul 10, 2026
4 min
Multimodal Reward Hacking in Reinforcement LearningResearch

Multimodal Reward Hacking in Reinforcement Learning

Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amplified when visual evidence is evaluated by text-only or weakly grounded rewards. We study reward hacking in MLLM RL across safety VQA, chart VQA, and stress-test settings, varying reward design, data ambiguity, model scale (2B-32B), and RL algorithm (GRPO, RLOO, DAPO). We introduce Newly Rewarded Failure Rate (NRFR), which measures failures among samples whose proxy reward improves over the SFT baseline. O

arXiv (cs.AI) · Jul 10, 2026
4 min
Terminal Dimension Reduction for Time Series with ApplicationsResearch

Terminal Dimension Reduction for Time Series with Applications

Terminal embeddings have emerged as a powerful tool for dimension reduction. Given a set of points $P\subset \mathbb{R}^d$, a terminal embedding is a mapping $f:\mathbb{R}^d\rightarrow \mathbb{R}^t$ that preserves the pairwise distance between any pair of points $p\in P$ and $q\in \mathbb{R}^d$ up to small distortion under this mapping. Terminal embeddings have been particularly fruitful for constructing $k$-means and $k$-median coresets, where the objective is to find a typically weighted subset $Ω$ of $P$ such that for any candidate solution, the cost of the clustering objective on

arXiv (cs.LG) · Jul 10, 2026
4 min
Neural Collapse Is Forbidden: Information Floors in Language ModelsResearch

Neural Collapse Is Forbidden: Information Floors in Language Models

Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and within-token context carries 79-91%, stable across a 100x parameter range. On the theory side, token-level weight decay penalizes a category in proportion

arXiv (cs.LG) · Jul 10, 2026
4 min
Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision TransformersResearch

Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers

The human visual system (HVS) employs foveated sampling and eye movements to achieve efficient perception, conserving both metabolic energy and computational resources. Drawing inspiration from this robustness and adaptability, we introduce the Foveated Dynamic Transformer (FDT), a foveation-guided dynamic token-selection architecture that integrates these mechanisms into a vision transformer framework. The FDT exhibits strong resilience to various types of noise and adversarial attacks, despite not being explicitly trained for such challenges. This inherent robustness is achieved th

arXiv (cs.LG) · Jul 10, 2026
3 min
ProofCouncil: An LLM Agent for Solving Open Mathematical ProblemsResearch

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved throug...

Papers with Code · Jul 10, 2026
1 min
Active rejection enables reliable generalization of universal machine-learning interatomic potentialsResearch

Active rejection enables reliable generalization of universal machine-learning interatomic potentials

Universal machine learning interatomic potentials (uMLIPs) bridge quantum-mechanical accuracy and large-scale molecular dynamics, but the cost of high-accuracy calculations such as r$^2$SCAN limits training to datasets that remain small relative to the open materials space. Strong average benchmark performance also does not guarantee reliable energy--force predictions for every structure. We propose Adaptive Multi-Teacher Routing (ATR), which reformulates high-fidelity data construction as a structure-wise decision problem under uncertainty. Using a small set of real r$^2$SCAN labels

arXiv (cs.LG) · Jul 10, 2026
4 min
Robustifying Vision-Language Models via Test-Time Prompt AdaptationResearch

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional structure of the data. This sample-centric approach limits robustness, as it fails to distinguish confident adversarial mispredictions from true semantic consistency. In this work, we observe that adversarial distortion is structurally brittle: while holistic representations are corrupted

arXiv (cs.LG) · Jul 10, 2026
3 min
Test-Time Scaling for Small VLMs on Multilingual Visual MCQResearch

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never comm

arXiv (cs.LG) · Jul 10, 2026
4 min
Multimodal Scenario Similarity Search for Autonomous DrivingResearch

Multimodal Scenario Similarity Search for Autonomous Driving

Large-scale autonomous-driving datasets contain vast numbers of recorded scenarios, creating a need for efficient retrieval methods that can identify situations similar to a given query. Existing approaches typically rely on either visual representations or motion-based descriptions, making it difficult to understand their relative strengths and limitations for scenario retrieval. In this work, we present a multimodal framework for autonomous-driving scenario retrieval that combines visual and trajectory-based representations within a unified retrieval pipeline. We investigate two tr

arXiv (cs.LG) · Jul 10, 2026
4 min
A Sovereign, Open-Source Foundation Model for German and EnglishResearch

A Sovereign, Open-Source Foundation Model for German and English

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both lan

arXiv (cs.LG) · Jul 10, 2026
4 min
Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device TuningResearch

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter cross-talk is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning - the same effect that plagues manual tuning of such systems. We propose using a factored representation of the action space, learned online, to decouple agents and minimize their interference. Our framework, QADAPT, uses this factorization to

arXiv (cs.LG) · Jul 10, 2026
3 min
Similarity search generalisation in contrastive learning with InfoNCE lossResearch

Similarity search generalisation in contrastive learning with InfoNCE loss

Similarity search is a primary application of embedding models trained by contrastive learning. For one of the most popular contrastive learning loss functions, InfoNCE, we show that the population risk with $k$ negative samples is $O(1/k)$ close to an expected cross-entropy which quantifies deviation between i) a softmax similarity search over unseen data using the learned embedding function, and ii) an idealised softmax search over the same data but using similarity implicitly represented in the positive sample generator. This complements existing interpretations of InfoNCE in the

arXiv (cs.LG) · Jul 10, 2026
3 min
SYNRARE: Synthetic Rare Disease EHR Generation for ML BenchmarkingResearch

SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking

Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants. Machine Learning Algorithms applied to Electronic Health Records show promise for accelerating the diagnosis; however, legal and privacy concerns pose significant barriers. To address these issues, Synthetic Data Generation is an alternative method for obtaining Electronic Health Records and can be applied with any Machine Learning algorithm for benchmarking and development purposes. Despite the availability of Synthetic Data Generation algorithms, support for

arXiv (cs.LG) · Jul 10, 2026
4 min
Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor ClassificationResearch

Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

Deep learning models dependency on large-scale inertial datasets presents a significant bottleneck in inertial sensor-based classification tasks, such as human activity recognition and smartphone location recognition. In these domains, data collection requires massive recording campaigns that are complex, time-consuming, and difficult to scale. Currently, data-driven guidelines for determining the minimum sample size required to reach a desired accuracy level do not exist. To address this gap, this study presents a systematic empirical evaluation of learning curve convergence rates i

arXiv (cs.LG) · Jul 10, 2026
4 min
Generative Communications: Overview, Technologies, and TrendsResearch

Generative Communications: Overview, Technologies, and Trends

The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as images and videos, resh...

Papers with Code · Jul 10, 2026
1 min
Introducing Muse Spark 1.1Research

Introducing Muse Spark 1.1

Today, we’re excited to introduce Muse Spark 1.1, the latest model from Meta Superintelligence Labs and a significant upgrade from Muse Spark. Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, with major gains in tool and computer use, coding, and multimodal understanding.

Meta AI · Jul 10, 2026
6 min
CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density StratificationResearch

CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification

Microbial density is clinically important for tumor assessment and treatment decision-making, and recent advances in deep learning suggest that it can be non...

Papers with Code · Jul 10, 2026
1 min
BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM ServingResearch

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences c...

Papers with Code · Jul 9, 2026
1 min
Wat3R: Underwater 3D Geometry Learning without AnnotationsResearch

Wat3R: Underwater 3D Geometry Learning without Annotations

Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quali...

Papers with Code · Jul 9, 2026
1 min
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video GeneratorsResearch

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video ...

Papers with Code · Jul 9, 2026
2 min
OpenCoF: Learning to Reason Through Video GenerationResearch

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning. To address this gap, we introduce OpenCoF, a framework comprising the OpenCoF-17K dataset, a rea

arXiv (cs.AI) · Jul 9, 2026
4 min
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationResearch

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to re

arXiv (cs.AI) · Jul 9, 2026
4 min
Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion SamplingResearch

Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling

Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal $L^2$ error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positi

arXiv (cs.LG) · Jul 9, 2026
4 min
MulTTiPop: A Multitrack Transcription Dataset for Pop MusicResearch

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription mod...

Papers with Code · Jul 9, 2026
1 min
MulTTiPop: A Multitrack Transcription Dataset for Pop MusicResearch

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evalu

arXiv (cs.LG) · Jul 9, 2026
3 min
SLORR: Simple and Efficient In-Training Low-Rank RegularizationResearch

SLORR: Simple and Efficient In-Training Low-Rank Regularization

Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. Existing training-time low-rank regularizers can improve compressibility, but they often require SVDs of large weight matrices, modify the model architecture (introducing additional trainable parameters), or rely on stateful cached quantities. To address these limitations, we introduce SLORR, a simple, stateless, and architecture-preserving framework for in-training low-rank regularization, instantiated wi

arXiv (cs.AI) · Jul 9, 2026
3 min
Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive AnalysisResearch

Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis

In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in higher education. Based on objective log data from 77,543 students enrolled in distance studies, we examine usage patterns across gender, age group, study cluster, degree, and study mode. To date, existing research on educational chatbots has largely relied on comparatively small samples and self-reported survey data, while large-scale evidence on actual usage behavior remains limited. Our findings show that Syntea is already embedded in the study routines of many lea

arXiv (cs.AI) · Jul 9, 2026
3 min
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN GraphResearch

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the data manifold in its original high-dimensional space, before the distortion that UMAP's 2D projection introduces. We demonstrate the untapped potential of this internal representation, showing how standard graph algorithms applied to this graph enhance data sensemaking: (1) PageRank identifies representative data points, (2) k-core decomposition rev

arXiv (cs.AI) · Jul 9, 2026
3 min
AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam UnderstandingResearch

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world dr

arXiv (cs.AI) · Jul 9, 2026
3 min
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion GenerationResearch

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context windows. In this work, we introduce ARDY, a streaming generation framework that bridges this

arXiv (cs.LG) · Jul 9, 2026
4 min
Workflow as Knowledge: Semantic Persistence for LLM-Mediated WorkflowsResearch

Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows

Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, checkpointing, and human approval. Existing workflow systems already address many execution concerns. This paper proposes a Lisp-inspired but language-independent conceptual model: symbolic forms, object identity, and live-image thinking are used as explanatory lenses, not implementation commitments. In this model, workflow definitions, workflow instances, inference records, context snapshots, and dependency relations are represented as persistent knowledge objects in a sha

arXiv (cs.AI) · Jul 9, 2026
3 min
The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMsResearch

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantiz

arXiv (cs.AI) · Jul 9, 2026
3 min
The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMsResearch

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on a...

Papers with Code · Jul 9, 2026
1 min
Super Weights in LLMs and the Failure of Selective TrainingResearch

Super Weights in LLMs and the Failure of Selective Training

Recent work identified Super Weights, individual parameters whose removal degrades model performance by orders of magnitude. We show that this degradation due to pruning Super Weights does not universally apply to all LLMs. Furthermore, if these parameters are so important, Super Weight-aware training should be effective. We show the opposite. Training Super Weights in isolation (100 to 8,192 parameters) drops accuracy to random-guessing levels on both OLMo-1B and OLMo-7B, and expanding to local neighborhoods of up to 36K parameters provides no improvement. The failure is specific to

arXiv (cs.LG) · Jul 9, 2026
4 min
Validity of LLMs as data annotators: AMALIA on authorityResearch

Validity of LLMs as data annotators: AMALIA on authority

A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity. For theoretical constructs that must be inferred rather than read from surface features, the question is whether the model follows the construct'

arXiv (cs.AI) · Jul 9, 2026
4 min
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute PredictionResearch

Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction

Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable. However, most pose estimators remain optimized for geometric keypoint accuracy, while many real-world applications in rehabilitation, sports science, ergonomics, and clinical movement analysis require biomechanical quantities that describe how the body moves, loads, and activates. In this work, we propose BioModule, a lightweight plug-in temporal transformer that attaches downstream of any 3D pose estimator and predicts biomechanical attributes from standard

arXiv (cs.AI) · Jul 9, 2026
4 min
Latent Memory Palace: Reasoning for Control as Autoregressive Variational InferenceResearch

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive "reasoning." However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and a

arXiv (cs.LG) · Jul 9, 2026
3 min
Deep Learning for Joint Narrowband Interference Cancellation and Soft Demodulation in OFDM SystemsResearch

Deep Learning for Joint Narrowband Interference Cancellation and Soft Demodulation in OFDM Systems

Narrowband interference (NBI) severely degrades orthogonal frequency-division multiplexing (OFDM) systems by corrupting subcarriers and rendering classical soft demodulation ineffective. Conventional compressed-sensing (CS) mitigation exhibits high sequential latency and leaves structured, non-Gaussian residuals that cause log-likelihood ratio (LLR) unreliability, decoder saturation, and severe error floors when employing classical Gaussian demappers. We resolve this pipeline mismatch using a unified deep learning framework for joint NBI cancellation and robust soft demodulation. Fir

arXiv (cs.LG) · Jul 9, 2026
4 min
Remember When It Matters: Proactive Memory Agent for Long-Horizon AgentsResearch

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank fr

arXiv (cs.AI) · Jul 9, 2026
4 min
Remember When It Matters: Proactive Memory Agent for Long-Horizon AgentsResearch

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajecto...

Papers with Code · Jul 9, 2026
2 min
LTM: Large-scale Terrain Model for Wildfire-prone LandscapesResearch

LTM: Large-scale Terrain Model for Wildfire-prone Landscapes

Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfire-prone regions often span vast areas where conventional reconstruction methods underperform. Airborne LiDAR systems provide high-resolution terrain data, but they are expensive and infrequently updated. Image-based methods offer a lower-cost alternative, but struggle due to sparse visual features and limited image overlap. We propose a multi-modal reconstruction framework leveraging outdated Digital Elevation Models (DEMs) as geometric priors for image-based 3D reconstructi

arXiv (cs.LG) · Jul 9, 2026
3 min
MPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement LearningResearch

MPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning

We address liquidity placement in the Bitcoin Lightning Network (LN): given a fixed budget, which channels should a node open to maximize its routing capacity? We cast this as a budget-constrained combinatorial optimization problem on graphs, selecting $k$ edge additions that maximize $s$--$t$ max-flow, a theory-grounded measure of routing capacity, and solve it with graph reinforcement learning. Our lightweight agent combines a message-passing policy network with proximal policy optimization (PPO) and action masking, and is trained under a hub-exclusion curriculum: the network's top

arXiv (cs.LG) · Jul 9, 2026
3 min
Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source AttributionResearch

Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before su...

Papers with Code · Jul 9, 2026
2 min
ProjAgent: Procedural Similarity Retrieval for Repository-Level Code GenerationResearch

ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation

Repository-level code generation requires implementing target functions while accounting for complex cross-file dependencies and project-specific conventions. Existing retrieval methods predominantly rely on lexical, structural, or semantic similarity, often overlooking repository functions that implement similar procedural logic despite differing in identifiers or application domains. We propose ProjAgent, a repository-level code generation system that introduces procedural similarity as an explicit retrieval signal. ProjAgent decomposes the target function into intermediate reasoni

arXiv (cs.AI) · Jul 9, 2026
3 min
A Practical Investigation of Training-free Relaxed Speculative DecodingResearch

A Practical Investigation of Training-free Relaxed Speculative Decoding

Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM. Standard speculative decoding is lossless: its rejection and resampling steps exactly preserve the LLM's sampling distribution. Recent work argues that relaxing this strict guarantee can yield further speed-ups, controlled capability-speed trade-offs, or even capability gains. We practically investigate training-free relaxed speculative decoding techniques, unify existing approaches within a shared framework, benchmark

arXiv (cs.AI) · Jul 9, 2026
3 min
SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy MarketsResearch

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions.

arXiv (cs.AI) · Jul 9, 2026
4 min
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language ModelsResearch

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-commit router captures; however, that guarantee holds only under an idealized oracle equipped with correctness labels and an unconstrained budget, neither of which a deployed system has. To the best of our knowledge, no previous work treats resampling the committed model and rerouting to an alternative model as

arXiv (cs.LG) · Jul 9, 2026
4 min
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web SearchResearch

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks. A single ReAct-style agent is constrained by one long trajectory and limited context, making it difficult to handle depth and coverage simultaneously. Existing multi-agent systems improve search coverage through parallel execution and aggregation, but still exhibit clear limitations in recursive depth, collaboration adaptability, and evidence-grounded expansion. We propose WebSwarm, a progressive

arXiv (cs.AI) · Jul 9, 2026
4 min
EdgeRefine: Privacy-Utility Balance for Graphs via Jaccard Sampling under Edge Differential PrivacyResearch

EdgeRefine: Privacy-Utility Balance for Graphs via Jaccard Sampling under Edge Differential Privacy

Graph Neural Networks (GNNs) have shown considerable success in learning from graph-structured data, but their use in privacy-sensitive areas remains difficult because graph structure can leak sensitive link information. To satisfy edge-level differential privacy, a common approach is to inject noise into all elements of the graph's adjacency matrix, thereby obfuscating the existence of any single edge. However, stronger privacy requires more noise, and excessive noise reduces utility, making the privacy-utility balance a major barrier to practical privacy-preserving graph learning.

arXiv (cs.LG) · Jul 9, 2026
4 min
Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation StudyResearch

Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study

Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complementary production specialties must trade within a constrained social network to obtain utili

arXiv (cs.AI) · Jul 9, 2026
3 min
Secure Decentralized Federated Learning via Gossip and Virtual VotingResearch

Secure Decentralized Federated Learning via Gossip and Virtual Voting

Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing gossip-based methods often lack provenance finality and resilience to Byzantine or lazy participants. Ledger-assisted federated learning (FL) improves auditability, yet blockchains, shards, or settlement committees can reintroduce global coordination costs that conflict with DFL locality. This paper proposes \emph{gspDAG-FL}, a secure DFL framework that derives consensus from the same gossip history used to disseminate models. Nodes excha

arXiv (cs.LG) · Jul 9, 2026
4 min
Multi-Modal, Multi-Environment Machine Teaching for Robust Reward LearningResearch

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions

arXiv (cs.AI) · Jul 9, 2026
4 min
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic EditingResearch

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on data expansion and more on higher-quality data utilization. However, in the context of large-scale corpora, existing refinement methodologies face significant limitations in quality, efficiency, and reliability: Rule-based approaches are constrained by fixed heuristics and struggle with instance-level variations; LLM-based approaches improve quality but fail to meet the efficiency and reliability requirement

arXiv (cs.AI) · Jul 9, 2026
4 min
BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionResearch

BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression

Large language models (LLMs) are increasingly constrained by memory capacity, weight bandwidth, and checkpoint storage during deployment. Existing low-bit compression methods mainly follow two directions. Scalar or group-wise quantization is simple and compatible with efficient low-precision kernels, but its representation capacity becomes limited when the target budget approaches 2 bits per weight. Vector-quantized weight compression provides a richer block-level representation, but usually introduces explicit codebooks, index lookup, and additional storage accounting. This paper pr

arXiv (cs.LG) · Jul 9, 2026
4 min
Steering Neural Network Training through Interpretable Constraints Based on Partial DependenceResearch

Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features o

arXiv (cs.LG) · Jul 9, 2026
4 min
The complexities of patient-centred conversational artificial intelligenceResearch

The complexities of patient-centred conversational artificial intelligence

Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversa

arXiv (cs.AI) · Jul 9, 2026
3 min
When Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiesResearch

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept. However, in vision-language models (VLMs), vanilla SAEs struggle to learn modality-consistent concepts, with concepts often exhibiting fragmented coverage (i.e., disjoint regions) in the visual modality. To address this challenge, we propose a Structured Sparse AutoEncoder ($S^2AE$) that enforces concept consistency from both semantic and spatial perspectives in the visual modality. Spe

arXiv (cs.AI) · Jul 9, 2026
4 min
Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment GuidanceResearch

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often miss within-stage heterogeneity and the clinical context in electronic medical records (EMRs). We present HCC-STAR (Hepatocellular Carcinoma Staging, Treatment And pRognosis), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score-based staging, ranked guideline-consistent treatments with evidence-based rationales, and individualized survival estimates. We

arXiv (cs.AI) · Jul 9, 2026
4 min
Federated Deep Learning for Privacy-Preserving Cardiovascular Disease Risk PredictionResearch

Federated Deep Learning for Privacy-Preserving Cardiovascular Disease Risk Prediction

Cardiovascular disease risk prediction models often rely on data from a single institution or centrally pooled datasets. Extending these models across institutions could be limited by privacy regulations and constraints on sharing patient-level data. Federated learning enables collaborative model development without transferring sensitive patient data, but its application in healthcare remains challenging because datasets often differ in size, population characteristics, and outcome definitions. In this study, we present a federated deep learning approach for privacy-preserving cardi

arXiv (cs.LG) · Jul 9, 2026
4 min
Robust Bayesian Decision Making under Adversarial UncertaintyResearch

Robust Bayesian Decision Making under Adversarial Uncertainty

Scientific experiments are often designed to maximize information gain, yet in many applications the primary objective is to support reliable downstream decision-making. Existing decision-aware experimental design and active learning methods typically assume well-specified outcome models and implicitly rely on the stability of the optimal decision under real-world perturbations. In practice, however, experimental outcomes are frequently influenced by hidden or weakly modeled effects, which can substantially alter decision optimality and lead to misleading conclusions. We study sequen

arXiv (cs.LG) · Jul 9, 2026
3 min
Spectral Stability of Pseudoinverse-Based Extreme Learning MachineResearch

Spectral Stability of Pseudoinverse-Based Extreme Learning Machine

Extreme Learning Machine (ELM) computes output weights analytically using the Moore-Penrose pseudoinverse. Although this leads to fast training, its numerical stability depends strongly on the conditioning of the hidden layer matrix. This paper studies pseudoinverse-based ELM from a spectral perspective. We show that the smallest singular value governs perturbation amplification in the output weights, while the condition number provides a quantitative measure of hidden-layer instability. We compare SVD-based pseudoinverse computation with iterative hyperpower methods and discuss widt

arXiv (cs.LG) · Jul 9, 2026
3 min
ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation MethodsResearch

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analyst...

Papers with Code · Jul 9, 2026
2 min
ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation MethodsResearch

ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods

Missing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce ImputeViz, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results. The system brings together widely used methods, including MICE, Random Forest, XGBoost, and kNN, within an interactive environment that makes missingness patterns explicit. To support geospatial reasoning, we introduce gKNN, a geographic

arXiv (cs.LG) · Jul 9, 2026
4 min
SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and LimitsResearch

SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits

Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or late fusion, which combines independently trained unimodal predictors. Early fusion can be accurate but monolithic, while late fusion is modular but may lose cross-modal interactions. This paper revisits XAI-guided adaptive fusion (\xgaf), a tree-based mixture of unimodal and cross-modal experts whose sample-level weights are derived from TreeSHAP attribution magnitudes. We focus on the effect of SHAP attribution reduction when experts have unequ

arXiv (cs.AI) · Jul 9, 2026
4 min
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingResearch

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of their KV\$-reuse exceeds 80% of request tokens in a production trace from BAILIAN, versus 54-62% in chat. This paper first contributes a systematic study of request scheduling for agents on

arXiv (cs.AI) · Jul 9, 2026
4 min
Contravariance Theory: Strong Alignment for Minimal Solutions to Hard TasksResearch

Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks

A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between artificial networks and real brain networks. Here, we show that for any two minimal DNN solutions to a sufficiently hard task: (i) "weak" alignment of network representations based on affine mappings guarantees "strong" alignment of privileged axes, and (ii) alignment "zippers" up the network hierarchy, causing the emergence of privileged axes from end-to-end task

arXiv (cs.LG) · Jul 9, 2026
3 min
CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal ConsistencyResearch

CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal Consistency

The operational integrity of complex industrial systems relies on precise anomaly detection and diagnosis. The vast majority of existing methods narrowly focus on capturing temporal similarities of representations, often overlooking the disruption of internal causal relationships, which characterizes system failures and latent anomalies. In this paper, we propose a novel framework (CAAD) that reframes anomaly detection as the continuous verification of Granger causality consistency through exogenous variables. Specifically, the CAAD framework models exogenous time-series variables as

arXiv (cs.LG) · Jul 9, 2026
3 min
CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban CommunitiesResearch

CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban Communities

For urban managers and designers, improving the functional attributes of urban communities to enhance territorial resilience in the face of complexity and uncertainty is crucial. Currently, community planning often follows a top-down approach and lacks effective metrics to quantify informal behaviors of residents, leading to frequent conflicts with original plans. This study introduces CommuniWave, a machine learning model designed to efficiently detect and quantify the Degree of Informal Behavior (DIB) in urban communities. The model integrates a Behavior Capture Net (BCN) based on

arXiv (cs.AI) · Jul 9, 2026
3 min
VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database RetrievalResearch

VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval

Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-language and reference-based approaches have significantly advanced this field, they often rely on text prompts, limited visual examples, or expensive feature matching procedures, making them difficult to scale to large and continuously expanding object repositories. In this work, we propose VocaDet, a sample-driven open-vocabulary object detection and segmentation framework that learns object concepts directly from user-provided positive and negat

arXiv (cs.AI) · Jul 9, 2026
4 min
Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCHResearch

Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH

Text-to-image (T2I) models have been shown to exhibit social biases. Prior work has mainly focused on gender, skin tone, and cultural representation within r...

Papers with Code · Jul 9, 2026
1 min
Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?Research

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models ofte...

Papers with Code · Jul 9, 2026
2 min
CT-CLIP Representations for Multimodal Lung Cancer Survival PredictionResearch

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the scarcit...

Papers with Code · Jul 9, 2026
1 min
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse AutoencodersResearch

Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

We present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained B...

Papers with Code · Jul 9, 2026
1 min
Frequency-Domain Multi-Modality Transportation ModelingResearch

Frequency-Domain Multi-Modality Transportation Modeling

Multi-modality transportation refers to urban systems composed of multiple transportation modes, such as traffic flow and public transit, whose dynamics are ...

Papers with Code · Jul 9, 2026
1 min
Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic SurgeryResearch

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); howe...

Papers with Code · Jul 9, 2026
2 min
Eigenvalue Calibration for Semantic Embeddings of Large Language ModelsResearch

Eigenvalue Calibration for Semantic Embeddings of Large Language Models

Uncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerge...

Papers with Code · Jul 9, 2026
1 min
FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language NavigationResearch

FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation

Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time visual inputs. Compa...

Papers with Code · Jul 9, 2026
2 min
Prediction-Powered Active TestingResearch

Prediction-Powered Active Testing

Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled. However, existing estima...

Papers with Code · Jul 9, 2026
1 min
ArtMine: Discovering and Formalizing Artistic ProcessesResearch

ArtMine: Discovering and Formalizing Artistic Processes

Understanding how artworks are created requires reasoning about the iterative decisions, material operations, and contextual influences that shape artistic p...

Papers with Code · Jul 9, 2026
2 min
Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBenchResearch

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However...

Papers with Code · Jul 9, 2026
2 min
Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language ModelsResearch

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates significant...

Papers with Code · Jul 9, 2026
1 min
LUMI: Tokenizer-Agnostic LLM-Based Lossless Image CompressionResearch

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained model, c...

Papers with Code · Jul 9, 2026
2 min
Introducing Muse Image and Muse VideoResearch

Introducing Muse Image and Muse Video

Muse Image follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. Muse...

Meta AI · Jul 9, 2026
5 min
MLQENABLER: Enabling Secure Machine Learning Queries over Encrypted Database in Cloud ComputingResearch

MLQENABLER: Enabling Secure Machine Learning Queries over Encrypted Database in Cloud Computing

In cloud computing, the public cloud service providers (CSPs) can provide cloud storage as the primary service while providing additional machine learning (M...

Papers with Code · Jul 9, 2026
1 min
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-ActionResearch

LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action

Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenar...

Papers with Code · Jul 9, 2026
2 min
Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future DirectionsResearch

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-world ...

Papers with Code · Jul 9, 2026
2 min
Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based ModelsResearch

Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an energy based model sub...

Papers with Code · Jul 9, 2026
1 min
Mixture of Enhanced-View Experts for Multi-Query Vehicle ReID and A Large-Scale BenchmarkResearch

Mixture of Enhanced-View Experts for Multi-Query Vehicle ReID and A Large-Scale Benchmark

Multi-query vehicle ReID aims to leverage complementary information from diverse views for robust feature learning. However, current methods suffer from simp...

Papers with Code · Jul 9, 2026
2 min
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and FaithfulnessResearch

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect...

Papers with Code · Jul 9, 2026
2 min
An exact information theory of generalization phase transitions in Bayesian diffusion modelsResearch

An exact information theory of generalization phase transitions in Bayesian diffusion models

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead o...

Papers with Code · Jul 9, 2026
2 min
Tool-Making and Self-Evolving LLM Agents in Low-Latency SystemsResearch

Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time...

Papers with Code · Jul 9, 2026
2 min
Unit-Independent Low-Rate Wrist GSR Processing for Stress Detection Using Phasic nSCR FeaturesResearch

Unit-Independent Low-Rate Wrist GSR Processing for Stress Detection Using Phasic nSCR Features

Galvanic skin response (GSR) is widely used for stress detection, but wrist-based GSR remains challenging because its absolute amplitude can differ substanti...

Papers with Code · Jul 9, 2026
2 min
A Quantum Reservoir Architecture for Chaotic Forecasting and a Test of Whether Its High Dimension HelpsResearch

A Quantum Reservoir Architecture for Chaotic Forecasting and a Test of Whether Its High Dimension Helps

Quantum reservoir computing uses a fixed quantum circuit as a feature generator and trains only a simple linear readout on top of it. This makes it cheap to ...

Papers with Code · Jul 8, 2026
2 min
GIRAF: Towards Generalizable Human Interactions with Articulated ObjectsResearch

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in ro...

Papers with Code · Jul 8, 2026
1 min
Infinity-Parser2 Technical ReportResearch

Infinity-Parser2 Technical Report

We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-e...

Papers with Code · Jul 8, 2026
1 min
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural ReasoningResearch

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order. However, applying artificial intelligence to this process presents a joint challenge of representation and reasoning: models must preserve domain-native structural inf

arXiv (cs.AI) · Jul 8, 2026
4 min
Co-LMLM: Continuous-Query Limited Memory Language ModelsResearch

Co-LMLM: Continuous-Query Limited Memory Language Models

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed. This recently introduced paradigm provides multiple advantages, including knowledge control capabilities that remain beyond conventional LLMs. We propose continuous-query LMLM (CO-LMLM), where the KB pairs continuous keys with textual knowledge values, a significant departure from prior reliance on relational KB and queries. CO-LMLM generates flexible vec

arXiv (cs.AI) · Jul 8, 2026
3 min
The Key to Going Linear: Analysis-Driven Transformer LinearizationResearch

The Key to Going Linear: Analysis-Driven Transformer Linearization

The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, sho

arXiv (cs.LG) · Jul 8, 2026
3 min
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database BypassResearch

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver layers that are not designed for bulk columnar analytics. We present Jailbreak, an approach that bypasses the database engine entirely by reading storage files directly and materializing data as in-memory columnar buffers. Jailbreak's key insight is that database file formats, while complex, are fully specified by their source code and documentatio

arXiv (cs.AI) · Jul 8, 2026
4 min
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI SafetyResearch

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing o

arXiv (cs.AI) · Jul 8, 2026
4 min
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHFResearch

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that substantially improve the feedback efficiency of diffusion RLHF while preserving generalizatio

arXiv (cs.AI) · Jul 8, 2026
4 min
Agon: Competitive Cross-Model RL with Implicit Rival Grading of ReasoningResearch

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen i

arXiv (cs.AI) · Jul 8, 2026
4 min
ECGLight: Compute-Light Framework For Paper ECG Digitization and Myocardial Infarction ScreeningResearch

ECGLight: Compute-Light Framework For Paper ECG Digitization and Myocardial Infarction Screening

Electrocardiography (ECG) is one of the most widely used tests for diagnosing cardiovascular disease. Yet several remote clinics still utilize paper ECG printouts for their analysis due to limited connectivity and computational capacity. As a result, vast numbers of physical ECGs obtained in remote areas still remain incapable of being accessed by contemporary artificial-intelligence (AI)-based decision support as they require high computational resources or strong high-speed internet connectivity. This causes several cases where conditions like acute coronary occlusion (ACS) is over

arXiv (cs.LG) · Jul 8, 2026
4 min
Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained OptimizationResearch

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization

The inverse design of physical systems governed by partial differential equations is computationally demanding due to the high dimensionality and non-convexity of design spaces. Generative models for inverse design often lack robustness and transferability, whereas evolutionary strategies are robust but struggle in high-dimensional spaces. This paper introduces a Neural Operator-enabled Topology-informed Evolutionary Strategy (NOTES) that integrates dimensionality reduction, representation learning, and evolutionary optimization for efficient and transferable inverse design. NOTES co

arXiv (cs.LG) · Jul 8, 2026
3 min
Any-Dimensional Learning by SamplingResearch

Any-Dimensional Learning by Sampling

Many machine learning models are defined for inputs of different sizes, such as point clouds containing different numbers of points, sequences of tokens of different lengths, and graphs on different numbers of nodes. Such models are trained on finitely-many examples of necessarily limited sizes. How well do these models generalize from inputs of small size to larger inputs of size not seen during training? Furthermore, evaluating such models on large inputs is often expensive. How can we sketch large inputs to obtain smaller ones on which the model takes similar values? At the heart

arXiv (cs.LG) · Jul 8, 2026
4 min
How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length GeneralizationResearch

How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization

Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly. We study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution tradeoff and show that, for a data-induced dependency profile of width $W$, the optimal frequency scales as $1/W$. This frequency-matching principle explains controlled o

arXiv (cs.LG) · Jul 8, 2026
4 min
SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI AgentsResearch

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but correct, secure, and maintainable. We introduce SkillCenter, to our knowledge the largest open skill library for agents by total count: 216,938 structured skills across 24 domain bundles. A SkillGate-filtered pipeline contributes 114,565 source-grounded skills from peer-reviewed journals, ArXiv, and over 24,000 technical sources, integrated with 102,373 community skills from GitHub and the ClawHub marketplace. W

arXiv (cs.AI) · Jul 8, 2026
3 min
Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning ProblemsResearch

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems

Group Relative Policy Optimization (GRPO) stalls on a model's hardest problems: when no rollout in a group succeeds, the group-relative advantages vanish and the problem contributes no gradient, wasting the frontier examples we most want to learn from. Prepending a correct prefix of a reference solution raises the success rate, making prefix length a continuous knob on difficulty. Concurrent methods set the knob once; AdaPrefix-GRPO turns it into a feedback controller: throughout training it adjusts how much of the solution each problem gets, holding its success rate near 50%, where

arXiv (cs.LG) · Jul 8, 2026
3 min
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation ModelsResearch

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure fo

arXiv (cs.LG) · Jul 8, 2026
4 min
PeTeR: Post-Training Robustification of Probabilistic CircuitsResearch

PeTeR: Post-Training Robustification of Probabilistic Circuits

Probabilistic circuits (PCs) can model complex joint distributions while supporting exact and efficient computation of many inference queries. However, standard likelihood-based PC learning is vulnerable to overfitting and fragile generalization when confronted with data noise, small sample sizes, or distribution shifts. This can be mitigated using distributionally-robust optimization which consider worst-case distributions within a Wasserstein ball of the empirical distribution, but current methods are limited to training a model from scratch in this framework. Instead, we propose P

arXiv (cs.LG) · Jul 8, 2026
3 min
Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model ScaleResearch

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale

Large language models hallucinate most about entities they have never seen. We ask whether a model's activations betray entity familiarity before a single answer token is generated, and whether that signal predicts the factual reliability of the answers. On four Polish Bielik models (1.5B-11B parameters), we probe four entity domains (athletes, cities, writers, musicians), each with 42 well-known, 42 obscure-but-real, and 42 fabricated entities addressed by a one-sentence question (504 prompts per model). Two unsupervised, single-forward-pass dispersion measures over post-SwiGLU MLP

arXiv (cs.LG) · Jul 8, 2026
4 min
DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect AdaptationResearch

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM}, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our resu

arXiv (cs.AI) · Jul 8, 2026
3 min
Guidance Breaks the Fitted Operator: A Terminal-Fitted Repair for Classifier-Free GuidanceResearch

Guidance Breaks the Fitted Operator: A Terminal-Fitted Repair for Classifier-Free Guidance

Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules. We analyze CFG through an asymptotic-preserving, numerical-analysis lens. Building on a recent result that the deterministic DDIM step is the unique fitted operator for the unguided terminal layer, exact on the final small-sigma stretch of sampling, we show that guidance re-stiffens exactly the discriminative subspace to an ano

arXiv (cs.LG) · Jul 8, 2026
4 min
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research LoopsResearch

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop clo

arXiv (cs.AI) · Jul 8, 2026
4 min
RL Post-Training Builds Compositional Reasoning StrategiesResearch

RL Post-Training Builds Compositional Reasoning Strategies

Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained on a Trace-based reasoning task with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, whi

arXiv (cs.AI) · Jul 8, 2026
4 min
ALER-TI: Aligned Latent Embedding Retrieval for Time Series ImputationResearch

ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation

Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within the corrupted input sequence. This reliance can be limiting in real-world scenarios, where time series often exhibit non-stationary dynamics, weak temporal correlations, and infrequent patterns that are difficult to reconstruct from nearby observations alone. In this paper, we propose ALER-TI, Aligned Latent Embedding Retrieval for Time Series Imputation, a retrieval-augmented framework that explicitly leverages historical patterns to su

arXiv (cs.AI) · Jul 8, 2026
4 min
An optimal control approach for neural network architecture adaptation with a posteriori error estimationResearch

An optimal control approach for neural network architecture adaptation with a posteriori error estimation

This work presents a novel approach for adapting neural network architecture along the depth based on a posteriori error estimation. By formulating neural network training as a continuous-time optimal control problem, we derive rigorous error estimates that quantify how approximation error distributes across network layers. This error decomposition enables a principled depth adaptation strategy: new layers are inserted at locations of maximum estimated error, allowing the network to efficiently capture complex, nonlinear variations in the underlying problem. Our framework introduces

arXiv (cs.LG) · Jul 8, 2026
4 min
QCNN with Rough Path Signature KernelsResearch

QCNN with Rough Path Signature Kernels

Time series analysis plays a vital role across a wide range of scientific and engineering domains but poses substantial computational challenges. A major difficulty arises from the time reparameterization invariance of time series data, which complicates the extraction of meaningful temporal features. In this work, we address the problem of time series classification by exploring the application of quantum computation techniques. We propose a hybrid quantum-classical architecture that integrates recent advances in quantum neural networks with the mathematical framework of path signat

arXiv (cs.AI) · Jul 8, 2026
3 min
Future Confidence Distillation in Large Language ModelsResearch

Future Confidence Distillation in Large Language Models

Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating answer reliability. Existing approaches, however, largely treat confidence as a property of completed responses, overlooking how confidence-related information evolves throughout the answering process. In this work, we investigate confidence from a temporal perspective by comparing pre-solution Feeling-of-Knowing (FOK) and post-solution Judgement-of-Learning

arXiv (cs.AI) · Jul 8, 2026
3 min
Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal CoordinatesResearch

Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal Coordinates

Nonlinear least-squares optimization is central to regression, physics-informed neural networks, and other machine-learning tasks. Such problems have a natural geometric interpretation, model predictions form a manifold in data space, while the chosen parameterization can introduce parameter-effects curvature that becomes a dominant source of nonlinearity. This exposes a limitation of the Levenberg-Marquardt (LM) method, its tangent-space step is applied as a straight update in parameter coordinates. Geodesic acceleration gives a second-order correction, but its removal of parameter-

arXiv (cs.LG) · Jul 8, 2026
4 min
Towards Agentic AI Governance: A Preliminary AssessmentResearch

Towards Agentic AI Governance: A Preliminary Assessment

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, introducing new ethical and governance challenges. This paper presents a systematic review of the emerging literature on agentic AI governance. Our analysis identifies features that distinguish agentic AI from traditional systems and why it warrants targeted governance attention. We synthesize prevailing governance priorities, proposed mechanisms, and

arXiv (cs.AI) · Jul 8, 2026
3 min
Asymmetric Focal Loss Improves Graph Neural Network Prediction of Drug-Drug InteractionsResearch

Asymmetric Focal Loss Improves Graph Neural Network Prediction of Drug-Drug Interactions

Background: Graph neural networks improve computational prediction of polypharmacy side effects, but standard binary cross-entropy training allocates equal capacity to well-classified and difficult examples, potentially missing clinically significant interactions. We evaluated whether an asymmetric focal objective could improve multi-relational drug-drug interaction (DDI) prediction by emphasizing difficult positive interactions. Methods: ClinicalFocal loss was integrated into a relation-aware graph convolutional network using molecular fingerprints, physicochemical descriptors, and

arXiv (cs.LG) · Jul 8, 2026
4 min
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case SynthesisResearch

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end-to-end generation but still struggle to ensure spatiotemporal consistency and physical re

arXiv (cs.AI) · Jul 8, 2026
4 min
Multi-Class vs. Multi-Label BERT for CVE-to-CWE Mapping: How Taxonomy Structure Shapes the ErrorsResearch

Multi-Class vs. Multi-Label BERT for CVE-to-CWE Mapping: How Taxonomy Structure Shapes the Errors

Assigning Common Weakness Enumeration (CWE) categories to Common Vulnerabilities and Exposures (CVE) records remains an important but largely manual step in vulnerability analysis. We study this task as a text classification problem and compare two modelling choices: a \emph{multi-class} formulation that predicts a single CWE per CVE and a \emph{multi-label} formulation that allows multiple assignments. Three transformer encoders (BERT Base, SecureBERT, and CySecBERT) are evaluated on three nested label spaces (83, 47, and 25 classes). Multi-class training achieves higher macro-F1 ac

arXiv (cs.LG) · Jul 8, 2026
4 min
Collaborative Synthetic Data Generation for Knowledge Transfer in Federated LearningResearch

Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning

One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but doing so without sacrificing model quality is non-trivial, particularly when client data distributions diverge. Recent work has addressed this challenge by aggregating client knowledge on the server through the construction of transferable synthetic datasets or distillates. However, most of these methods lack formal privacy guarantees, leaving a gap in jointly achieving low communication, robustness to heterogeneity, and rigorous privacy. We propos

arXiv (cs.AI) · Jul 8, 2026
4 min
PALS: Percentile-Aware Layerwise Sparsity for LLM PruningResearch

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

One-shot pruning methods like Wanda and SparseGPT apply the same sparsity ratio to every layer of a transformer, ignoring known variation in layer importance. We propose PALS (Percentile-Aware Layerwise Sparsity), which adjusts per-layer sparsity based on the 99th percentile of activation magnitudes, bounded to $\pm 5\%$ around the target ratio. On LLaMA-2-7B at 50\% sparsity, PALS achieves 10.96 WikiText-2 perplexity versus 12.92 for uniform Wanda (mean over 9 runs, $p < 0.001$). The benefit is architecture-dependent: LLaMA-3-8B shows marginal gains and Mistral-7B shows none. We als

arXiv (cs.LG) · Jul 8, 2026
3 min
Avoiding unsafe sets when training with Langevin DynamicsResearch

Avoiding unsafe sets when training with Langevin Dynamics

Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics on the loss landscape, and a natural safety question is to bound the probability $ν_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in \mathcal{A}_H)$ that the trajectory lies in a designated failure region $\mathcal{A}_H$. We study this for a smooth, strongly convex loss in $d$ dimensions and a failure region separated from the minimizer by an energy gap. Three bounds emerge. At the end of training, the equilibrium mass $π(\mathcal{A}_H)$ is exponentially small in $d$, with a complementary energy-barrie

arXiv (cs.LG) · Jul 8, 2026
4 min
A Unified Detection Framework for AI-Related Content and ArtifactsResearch

A Unified Detection Framework for AI-Related Content and Artifacts

Artificial intelligence (AI) is a double-edged sword: while it has achieved remarkable success across a wide range of domains, its deployment also calls for effective oversight and regulation, for which the detection of AI-related content and artifacts is perhaps the most direct and cost-effective approach. To this end, we propose a unified detection framework based on Mahalanobis distance scores (MDS), applicable to several important settings, including the detection of large language model (LLM) generated text, hallucination, watermark, and adversarial examples. A key component of

arXiv (cs.LG) · Jul 8, 2026
4 min
Creativity from Friction: Human-AI Interaction for Exploratory Structural DesignResearch

Creativity from Friction: Human-AI Interaction for Exploratory Structural Design

AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and architecture need interactive systems that help users externalise and develop ideas, explore alternatives, and refine partial solutions. The final product of such designs needs to comply with many constraints concerning, e.g., spatial configuration, mechanical behaviour, material quantities, and costs. These constraints create friction in the design process, which can stimulate novel and creative solutions. In this paper, we discuss the misali

arXiv (cs.AI) · Jul 8, 2026
4 min
Stability of Flow Models for Graph SignalsResearch

Stability of Flow Models for Graph Signals

Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While favorable stability properties of Graph Neural Networks (GNNs) have been well documented, it is unclear how structural errors propagate through the dynamics of continuous generative flow models that are gaining traction for graph signal generation. In this paper, we analyze continuous normalized flow models parameterized by GNNs and show that permutation equivariance is preserved for both the resulting continuous-time ordinary differenti

arXiv (cs.AI) · Jul 8, 2026
3 min
Single-Rollout Asynchronous Optimization for Agentic Reinforcement LearningResearch

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framew

arXiv (cs.AI) · Jul 8, 2026
4 min
HIVE: Understanding Post-Hallucination Reasoning in Vision Language ModelsResearch

HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models

Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses on detecting or suppressing hallucinations at generation time, leaving the subsequent reasoning stage largely unexplored. In this work, we study Post Hallucination Reasoning (PHR), the stage in which hallucinated semantics enter the model's inference context and influence downstream predictions. To systematically investigate PHR, we introduce HIVE, Hallucination Inference and Verification Engine, an evaluati

arXiv (cs.AI) · Jul 8, 2026
4 min
Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science WorkflowsResearch

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written skills can encode high-quality guidance, but writing and maintaining them across many data-science task families creates a manual bottleneck. We ask whether LLM-generated skills offer a useful low-curation alternative: do they improve performance over the task prompt alone? We test this quest

arXiv (cs.AI) · Jul 8, 2026
4 min
TimEE: End-to-end Time Series Classification via In-Context LearningResearch

TimEE: End-to-end Time Series Classification via In-Context Learning

Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pretraining on large corpora -- and then fit a task-specific classifier on top. While effective, this decoupling optimizes representation learning independently of the classification objective, requires per-dataset training, and prevents the model from exploiting label information during inference. We introduce TimEE, a 4.5M-parameter foundation model for end-to-end TSC via in-context learning. Given a labeled support set and a query time

arXiv (cs.AI) · Jul 8, 2026
4 min
Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26Research

Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26

Testing is a major effort for the gaming industry, requiring a significant part of development budget and people power. We present a case study on a development version of the ice hockey game EA SPORTS NHL 26, for which human playtesters test the goalie AI for behavioral exploits. To reduce the effort of re-testing the goalie AI after every game or behavior modification in the development phase, we propose Reward-Adaptive Iterative Discovery (RAID), a novel approach to automatically find exploits using an iterative Reinforcement Learning (RL) approach that trains a population of goal

arXiv (cs.AI) · Jul 8, 2026
4 min
Search, Fail, Recover: A Training Framework for Correction-Aware ReasoningResearch

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent, a training and inference framework inspired by the Diligent Learner formulation that represents reasoning as validated search over partial solution chains. A task validator labels generated continuations and failures, and the resulting search trees are converted into supervised targets for three actions: continue, finish, and backtrack, with optional tra

arXiv (cs.AI) · Jul 8, 2026
3 min
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI AgentsResearch

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action was. We introduce an action-graded harm rubric that scores an agent's tool-call trajectory on a seven-level ordinal scale (L0 to L6) according to whether the executed action was reversible, whether it crossed scope to reach another party, and whether it expanded privilege. We compute the scale two ways: a deterministic o

arXiv (cs.AI) · Jul 8, 2026
4 min
Predicting LLM Safety Before Release by Simulating DeploymentResearch

Predicting LLM Safety Before Release by Simulating Deployment

Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how oft...

Papers with Code · Jul 8, 2026
2 min
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive StrategiesResearch

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Googl...

Papers with Code · Jul 8, 2026
2 min
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language ModelResearch

GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and trai...

Papers with Code · Jul 8, 2026
2 min
Efficient Bayesian Deep Ensembles via Analytic Predictive InferenceResearch

Efficient Bayesian Deep Ensembles via Analytic Predictive Inference

We introduce an efficient Bayesian deep ensemble method for predictive regression designed to enhance interpretability while maintaining competitive predicti...

Papers with Code · Jul 7, 2026
1 min
SPEAR: A Simulator for Photorealistic Embodied AI ResearchResearch

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators s...

Papers with Code · Jul 7, 2026
2 min
Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site DatasetsResearch

Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site Datasets

Deep learning-based computer-aided diagnosis (CAD) systems have shown strong performance in breast cancer diagnosis, particularly for classification tasks in...

Papers with Code · Jul 7, 2026
2 min
SASGeo: Stability-Aware Semantic Map Localization for GNSS-Denied UAVs -- A Framework and Synthetic Proof of ConceptResearch

SASGeo: Stability-Aware Semantic Map Localization for GNSS-Denied UAVs -- A Framework and Synthetic Proof of Concept

GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry. Cross-view image retrieval ca...

Papers with Code · Jul 7, 2026
2 min
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language ModelsResearch

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustne...

Papers with Code · Jul 7, 2026
1 min
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application GenerationResearch

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts th...

Papers with Code · Jul 7, 2026
2 min
WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT SynthesisResearch

WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis

Generating CT volumes from MRI and CBCT can improve treatment planning in adaptive radiotherapy while avoiding additional radiation exposure. However, direct...

Papers with Code · Jul 7, 2026
1 min
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at ScaleResearch

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to un...

Papers with Code · Jul 7, 2026
2 min
A toy framework for single and multi-agent human-AI curiosity ecosystemsResearch

A toy framework for single and multi-agent human-AI curiosity ecosystems

This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why an a...

Papers with Code · Jul 7, 2026
1 min
Progressive Reasoning with Primitive Correction for Compositional Zero-Shot LearningResearch

Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning

Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. ...

Papers with Code · Jul 7, 2026
1 min
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgmentResearch

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically i...

Papers with Code · Jul 6, 2026
2 min
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action ModelResearch

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing view-robust Vision-Language-Action (VLA) policies tolerate such camera variations only when the camera extrinsics are explicitly provided, making them fragile and hard to use especially when view robustness is critical. We argue that the policy should not be told where the camera is, but rather figure it out by itself. To this end, we introduce Camera-Centric VLA (CamVLA), a new VLA model that decouples manipu

arXiv (cs.AI) · Jul 6, 2026
4 min
Weak-to-Strong Generalization via Direct On-Policy DistillationResearch

Weak-to-Strong Generalization via Direct On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with th

arXiv (cs.AI) · Jul 6, 2026
4 min
Weak-to-Strong Generalization via Direct On-Policy DistillationResearch

Weak-to-Strong Generalization via Direct On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new...

Papers with Code · Jul 6, 2026
2 min
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty QuantificationResearch

Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification

Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines. Reliable labels are costly, while community labels can be noisy and survey-dependent. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using injected transients and bogus-dominated survey data, remains robust under strong class contamination, and provides calibrated uncertainty quantification. We combine simulated transient injections with a contaminated survey class and train a dual-network mo

arXiv (cs.AI) · Jul 6, 2026
4 min
LLM-as-a-Verifier: A General-Purpose Verification FrameworkResearch

LLM-as-a-Verifier: A General-Purpose Verification Framework

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes

arXiv (cs.AI) · Jul 6, 2026
4 min
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual GenerationResearch

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed multimodal SearchGen-Corpus-1M to support offline, reproducible research. On SearchGen-Bench

arXiv (cs.AI) · Jul 6, 2026
4 min
What Does a Discrete Diffusion Model Learn?Research

What Does a Discrete Diffusion Model Learn?

What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are one object in different coordinates, and reading a neural network in the wrong coordinate changes the process being trained and sampled. Starting with a rigorous derivation of the continuous-time Markov chain (CTMC) ELBO for any noising process, boundary terms included, we prove the \emph{Oracle Distance} theorem: the negative ELBO is exactly equal to the data entropy plus the path KL from the oracle reverse process to the learned one, not merely

arXiv (cs.AI) · Jul 6, 2026
4 min
TabPack: Efficient Hyperparameter Ensembles for Tabular Deep LearningResearch

TabPack: Efficient Hyperparameter Ensembles for Tabular Deep Learning

In deep learning for tabular data, efficient ensembles of multilayer perceptrons (MLPs) have recently emerged as effective and practical architectures. Existing methods of this kind use the same hyperparameters for all underlying MLPs, which requires hyperparameter tuning for achieving the best performance. In this work, we introduce TabPack, an efficient MLP ensemble with strong out-of-the-box performance and reduced reliance on traditional tuning. In a single run, TabPack samples and trains many MLPs with different hyperparameters efficiently in parallel and selects ensemble member

arXiv (cs.LG) · Jul 6, 2026
3 min
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentsResearch

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents

Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continuing the rollout under a compressed context, but incorporating compaction into reinforcement learning remains underexplored. We propose CompactionRL, a reinforcement learning strategy to train long-horizon agentic LLMs with context compaction. Our approach jointly optimizes task execution and summary generatio

arXiv (cs.LG) · Jul 6, 2026
3 min
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon ManipulationResearch

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations. Hierarchical dual-system methods address this but suffer from a gap between high-level planning semantics and low-level execution kinematics. We introduce Cortex, a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from high-level VLM to low-level VLA. Specifically, we standardize manipulat

arXiv (cs.AI) · Jul 6, 2026
4 min
Fitted Occupancy-Ratio Evaluation without Bellman CompletenessResearch

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class i

arXiv (cs.LG) · Jul 6, 2026
4 min
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation TasksResearch

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-free policies? We focus on "Variational Automation" (VA), a class of tasks that have larger variations in object geometry and pose than fixed automation. Model-free policies often struggle to close the reliability gap for VA tasks, which must be executed persistently and reliably in commercial and industrial applications. Motivated by prior work on Task and Motion Planning (TAMP) and the

arXiv (cs.AI) · Jul 6, 2026
4 min
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language ModelsResearch

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality. We introduce SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models from question-answer interactions. SPEARBench constructs controlled dialogue prompts from the Seamless In

arXiv (cs.AI) · Jul 6, 2026
3 min
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution EditingResearch

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-speech spans: the transcript may remain plausible, but the decoded time axis drifts away from the audio. We study this non-speech-induced timestamp drift with self-built gap and long-gap benchmarks across 15 evaluated timestamp-producing ASR and audio-language systems. Naive timestamp-corrected fine-tuning improves alignment but can severely degrade

arXiv (cs.AI) · Jul 6, 2026
4 min
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent ConstraintsResearch

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user's current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. We introduce SovereignPA-Bench, an executable benchmark for evaluating user-owned personal

arXiv (cs.AI) · Jul 6, 2026
4 min
Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP PlanningResearch

Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning

Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. Tree-based search methods such as Monte Carlo Tree Search (MCTS) remain popular, but their branching structure can require sampling budgets that grow exponentially with lookahead depth in the worst case. From a tree perspective, continuous state or action spaces become especially challenging, since the planner must decide where to search in an infinite branching hierarchy. We propose Graph Sparse Sampling (GSS), an online planning algorithm that shares sampled futures

arXiv (cs.AI) · Jul 6, 2026
4 min
Faithfulness to Refusal: A Causal Audit of Neuron SelectorsResearch

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits built on one-shot neuron-row zeroing. We first audit selectors at the language-modeling level: attribution methods substantially outperform activation and magnitude-based baselines at identifying dispensable rows across five LLMs. We then adapt the same intervention into a behavior test by driving it with a con

arXiv (cs.LG) · Jul 6, 2026
4 min
Selective Disclosure Watermarking for Large Language ModelsResearch

Selective Disclosure Watermarking for Large Language Models

Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). Existing approaches include zero-bit schemes for distinguishing synthetic text from human writing and multi-bit schemes for embedding metadata. However, current multi-bit watermarking methods do not allow selective disclosure: verifying any part of the watermark requires revealing the entire embedded message. This lack of control leads to unnecessary information exposure and raises privacy concerns. We propose Hierarchical Vocabulary Routing (HeRo), a watermarking fram

arXiv (cs.AI) · Jul 6, 2026
3 min
Multiplayer Interactive World Models with Representation AutoencodersResearch

Multiplayer Interactive World Models with Representation Autoencoders

We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly avail

arXiv (cs.AI) · Jul 6, 2026
4 min
OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative RefinementResearch

OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement

We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-ready mathematical formulation as well as executable code. Our architecture prioritizes the mathematical modeling step, where dedicated agents extract structures, such as decision variables and constraints, enabling iterative self-correction. We introduce a novel multi-loop validation architecture with four specialized feedback mechanisms, each targeting a distinct failure mode such as misinterpretation, structural defects, mathematica

arXiv (cs.AI) · Jul 6, 2026
4 min
TREK: Distill to Explore, Reinforce to RefineResearch

TREK: Distill to Explore, Reinforce to Refine

Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support. We propose TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion. A key advantage of TREK is its generality: because it only consumes verified output trajectories, it can use an external black-box teacher, a white-box teacher, or the same model given addition

arXiv (cs.AI) · Jul 6, 2026
4 min
TREK: Distill to Explore, Reinforce to RefineResearch

TREK: Distill to Explore, Reinforce to Refine

Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts w...

Papers with Code · Jul 6, 2026
2 min
How Far is Too Far? Defining the Distance Threshold for Verification Siamese NetworksResearch

How Far is Too Far? Defining the Distance Threshold for Verification Siamese Networks

Siamese verification networks are widely used to compare items such as faces, cars, or signatures. In these scenarios, the network is trained to learn an embedding space in which similar objects are mapped closer together, while dissimilar objects are mapped further apart. Two objects are considered to belong to the same class (e.g., the same person in two different images) when the distance between their embeddings falls below a predefined threshold. Defining this threshold, however, is a non-trivial task and typically requires labeled data. In this work, we assume that the distribu

arXiv (cs.LG) · Jul 6, 2026
4 min
Steering Optimisation Trajectories in Diffusion Representation LearningResearch

Steering Optimisation Trajectories in Diffusion Representation Learning

We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut

arXiv (cs.AI) · Jul 6, 2026
3 min
Topological Shape Representation for Aneurysm -- Bifurcation DetectionResearch

Topological Shape Representation for Aneurysm -- Bifurcation Detection

Automated detection of intracranial aneurysms (IAs) from CT angiography (CTA) is severely hindered by high false-positive rates. Convolutional neural networks (CNNs) rely on local pixel intensities, causing systematic confusion between saccular aneurysms and vascular bifurcations -- a problem especially acute for small lesions (<3 mm), where detection sensitivity falls below 60%. We propose a plug-and-play, topology-aware false-positive reduction framework evaluating the Smooth Euler Characteristic Transform (SECT) -- a directional representation encoding global 3D vascular geometry

arXiv (cs.AI) · Jul 6, 2026
4 min
How Much is Left? LLMs Linearly Encode Their Remaining Output LengthResearch

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains. Training minimal-capacity linear probes on frozen hidden states of three open-weight 7-8B models across seven completion-style datasets, we find three converging pieces of evidence. First, total response length is linearly decodable from

arXiv (cs.LG) · Jul 6, 2026
4 min
Evaluating and Understanding Model Editing for Medical Vision Language ModelsResearch

Evaluating and Understanding Model Editing for Medical Vision Language Models

Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise, and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composi

arXiv (cs.AI) · Jul 6, 2026
4 min
Quantum Spectral Anomaly DetectionResearch

Quantum Spectral Anomaly Detection

A core task in quantum anomaly detection is to compute an anomaly score that quantifies how strongly a test quantum state deviates from a given quantum dataset assumed to be normal. Classically, principal component analysis (PCA) for centered data computes the anomaly score by evaluating the test sample relative to the subspace spanned by the selected leading eigenvectors. However, for quantum data that lack a standard centering, explicitly recovering principal eigenvectors, constructing full Gram matrices, or loading quantum-random-access-memory-style data can be more costly than es

arXiv (cs.LG) · Jul 6, 2026
4 min
Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in CancerResearch

Biologically Informed Deep Neural Networks for Multi-Omic Integration, Pathway Activity Inference and Risk Stratification in Cancer

Integrating complex, multi-omics data presents significant challenges. Existing approaches often face a trade-off between model interpretability and representational capacity, with most either relying on post-hoc interpretation or use linear models that may overlook complex interactions. We report Pathway Activity Autoencoders for the multi-omics setting, which embed prior knowledge via pathway-informed architectural constraints, fostering interpretability, while preserving representational power. Our multi-omic framework is applied in the context of breast cancer and is evaluated in

arXiv (cs.LG) · Jul 6, 2026
4 min
Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning PoisoningResearch

Learning Only What Valid Adapters Can Express: Subspace-Constrained Adaptation Against Fine-Tuning Poisoning

Parameter-efficient fine-tuning still leaves a broad space of behavior-changing updates reachable, so a poisoned objective can be represented and optimized. We study an alternative: adaptation constrained to the subspace estimated from a trusted pool of existing task adapters. On flan-t5-large with 196 public LoRA adapters, we show that (1) the functionally relevant content of an adapter lies in a low-dimensional shared subspace, 30 to 38 percent of its weight norm being redundant under the evaluated task distributions; (2) gradient adaptation restricted to 128 coordinates on this su

arXiv (cs.LG) · Jul 6, 2026
4 min
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill EvolutionResearch

MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution

Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability. However, a fixed, hand-authored skill is rarely optimal, and cannot adapt to the diversity of tasks an agent encounters. Self-improving agents address this by rewriting their own skill files from execution traces, yielding meaningful gains on challenging benchmarks. Yet such self-evolution remains non-recursive: it improves only the task skill (what the agent does) while the improvement procedure (how it improv

arXiv (cs.AI) · Jul 6, 2026
4 min
Air Quality Downscaling with Station-Guided Pseudo-SupervisionResearch

Air Quality Downscaling with Station-Guided Pseudo-Supervision

Super-resolving coarse atmospheric fields to local PM$_{2.5}$ variations is uniquely challenged by a mismatch in spatial support: while pixels represent regional averages, ground-truth observations are discrete, unaligned samples of a continuous spatial signal. To bridge this gap, we present a station-guided framework for high-resolution PM$_{2.5}$ downscaling over Europe. Taking coarse CAMS atmospheric composition fields alongside heterogeneous side information (i.e., human activity, land cover, elevation, satellite aerosol observations, and wind fields) our framework jointly super-

arXiv (cs.AI) · Jul 6, 2026
3 min
Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEGResearch

Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG

Schizophrenia is a debilitating neuropsychiatric disorder characterized by profound cortical network dysregulation, for which objective, clinically translatable EEG based biomarkers remain underdeveloped. Existing automated classification pipelines rely predominantly on static power spectral density features inherently blind to amplitude modulation dynamics and cross-frequency coupling, phenomena central to schizophrenia pathophysiology, while adopting epoch level cross validation strategies that introduce temporal data leakage, artificially inflate reported performance. This study i

arXiv (cs.AI) · Jul 6, 2026
4 min
Routing Anonymity and Identifiability of Noisy Quantum HardwareResearch

Routing Anonymity and Identifiability of Noisy Quantum Hardware

Present-day quantum computing is cloud-based, where a user submits a circuit to a service provider's proprietary backend hardware. While providers may wish to hide implementation details, scheduling choices, or even which physical device was used, noisy finite-shot outputs can carry backend-specific fingerprints: information imprinted in the classical output distribution that can reveal the backend identity. So far, such fingerprints have mostly been studied from a benchmarking perspective, with limited attention to privacy considerations for users and providers. This work develops t

arXiv (cs.LG) · Jul 6, 2026
4 min
Advances in Neural Controlled Differential EquationsResearch

Advances in Neural Controlled Differential Equations

Many real-world systems evolve continuously, yet most machine learning models interpret time series as discrete sequences. Continuous-time approaches instead treat time series as samples from an underlying input path, a formulation that naturally accommodates irregularly sampled or oversampled data. Among these, Neural Controlled Differential Equations (NCDEs) are a maximally expressive class of models that parametrise a vector field using a neural network and evolve their hidden state by solving a dynamical system driven by the input path. NCDEs typically use a non-linear vector fie

arXiv (cs.LG) · Jul 6, 2026
4 min
Untrusted Content Masking for Web Agents with Security GuaranteesResearch

Untrusted Content Masking for Web Agents with Security Guarantees

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such as tool-use APIs, this separation arises naturally: agents can reason from interface definitions without ever processing untrusted content. Extending these guarantees to web agents faces a fundamental challenge: to perceive and interact with their environment, web agents must first observe the rendered page, which intermingles trusted content with untrusted content. This structural entanglement removes the

arXiv (cs.LG) · Jul 6, 2026
3 min
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding DistributionsResearch

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications. We introduce ProPS, Prompted Profile Synthesis, a framework for generating distributions of speaker embeddings conditioned on natural language prompts such as "a thirties male speaker with an Indian accent". ProPS converts human-written profile descriptions into sentence embeddings and uses a

arXiv (cs.AI) · Jul 6, 2026
4 min
Adaptive Inference Batching using Policy GradientsResearch

Adaptive Inference Batching using Policy Gradients

Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot adapt to shifting traffic. We investigate whether reinforcement learning (RL) can learn adaptive batching and routing policies that outperform these heuristics, training REINFORCE and PPO agents on a discrete-event simulator validated against queuing theory and production traces (Azure Functions, BurstGPT). We formulate the problem as an MDP over queue state, request type and GPU availabil

arXiv (cs.AI) · Jul 6, 2026
4 min
Shifting from Discrete to Continuous Reference Data: QSM-Derived Horizontal Tree Biomass Distribution for Deep Learning Biomass EstimationResearch

Shifting from Discrete to Continuous Reference Data: QSM-Derived Horizontal Tree Biomass Distribution for Deep Learning Biomass Estimation

Conventional modeling approaches for LiDAR-based above-ground biomass (AGB) estimation rely on discrete plot-level inventory aggregates. This methodology introduces boundary-effect uncertainties that may severely degrade model performance within small field plots. To solve this limitation, we evaluate a Horizontal Biomass Distribution (HBD) reference mapped continuously from Quantitative Structure Models (QSMs). We trained a sparse 3D U-Net on simulated broadleaved forest structures using three AGB reference types: a standard forest inventory (FI) plot-level aggregate, an edge-effect

arXiv (cs.AI) · Jul 6, 2026
4 min
Repurposing CLIP to Localize at Pixel LevelResearch

Repurposing CLIP to Localize at Pixel Level

Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capab...

Papers with Code · Jul 6, 2026
1 min
Three-Phase Evaluation of AI-Assisted Software Development Life CycleResearch

Three-Phase Evaluation of AI-Assisted Software Development Life Cycle

This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect software development productivity, requirement adherence, and de...

Papers with Code · Jul 6, 2026
1 min
Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation PoliciesResearch

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and amb...

Papers with Code · Jul 6, 2026
1 min
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and DefensesResearch

Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses

Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usage information, and...

Papers with Code · Jul 6, 2026
2 min
Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music GenerationResearch

Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music Generation

This paper introduces a quantum-inspired computational framework for harmonic decision-making in music. The proposed approach formulates harmonization as an ...

Papers with Code · Jul 6, 2026
2 min
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMsResearch

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constrain...

Papers with Code · Jul 6, 2026
1 min
Multi-Turn On-Policy Distillation with Prefix ReplayResearch

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teache...

Papers with Code · Jul 6, 2026
2 min
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social DilemmasResearch

Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social ...

Papers with Code · Jul 6, 2026
2 min
Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based MethodologyResearch

Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology

Depression screening from large-scale behavioral data is challenged by fragmented circadian indicators, limited interpretability, and the lack of interventio...

Papers with Code · Jul 6, 2026
2 min
Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition DistributionsResearch

Learning Flexible Generalization in Video Quality Assessment by Bringing Device and Viewing Condition Distributions

Video quality assessment (VQA) plays a critical role in optimizing video delivery systems. While numerous objective metrics have been proposed to approximate...

Papers with Code · Jul 6, 2026
1 min
Measuring What Matters: A Unified Evaluation Framework for GNN ExplainabilityResearch

Measuring What Matters: A Unified Evaluation Framework for GNN Explainability

Graph eXplainable AI (G-XAI) is increasingly important for making Graph Neural Networks interpretable and accountable. While a growing number of explainers a...

Papers with Code · Jul 6, 2026
1 min
Learning to Control LLM Agent Harnesses with Offline Reinforcement LearningResearch

Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model i...

Papers with Code · Jul 5, 2026
2 min
Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical InnovationResearch

Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical Innovation

We advance the hypothesis that human mathematical reasoning, constrained by both the undecidability and the computational intractability of even modest logic...

Papers with Code · Jul 5, 2026
2 min
MechMath Agent Team: LLM Driven Agents for Mathematical ResearchResearch

MechMath Agent Team: LLM Driven Agents for Mathematical Research

AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematica...

Papers with Code · Jul 5, 2026
2 min
Legible-by-Construction: Attention and End-to-End TransformersResearch

Legible-by-Construction: Attention and End-to-End Transformers

A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a self...

Papers with Code · Jul 5, 2026
2 min
Topology-Driven Transferability Estimation for 3D Medical Vision Foundation ModelsResearch

Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on exhau...

Papers with Code · Jul 5, 2026
2 min
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and CybersecurityResearch

Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity

I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this day. The problem ...

Papers with Code · Jul 5, 2026
1 min
MANCE: Manifold Aware Concept ErasureResearch

MANCE: Manifold Aware Concept Erasure

Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because represe...

Papers with Code · Jul 4, 2026
2 min
From Brain Waves to Words: Brain2Qwerty Offers a New Path to Communication Without SurgeryResearch

From Brain Waves to Words: Brain2Qwerty Offers a New Path to Communication Without Surgery

Last year, we introduced Brain2Qwerty v1, research that uses AI to decode brain activity into text without any surgical implant. Now we're sharing the next step: Brain2Qwerty v2, the highest-performing end-to-end pipeline capable of real-time sentence decoding from non-invasive brain recordings, approaching levels of accuracy previously exclusive to techniques that require brain surgery.

Meta AI · Jul 4, 2026
3 min
How Alta Daily Uses Meta’s Segment Anything to Reimagine the Digital ClosetResearch

How Alta Daily Uses Meta’s Segment Anything to Reimagine the Digital Closet

Most people wear only an estimated 20% of the clothes in their closets, but the rest isn’t useless. It's untapped potential: combinations that never quite come together because remembering what you own, let alone imagining how it all fits together, is harder than it sounds.

Meta AI · Jul 4, 2026
4 min
SAM 3.1: Faster and More Accessible Real-Time Video Detection and Tracking With Multiplexing and Global ReasoningResearch

SAM 3.1: Faster and More Accessible Real-Time Video Detection and Tracking With Multiplexing and Global Reasoning

We’ve seen incredible adoption of SAM 3 over the last few months, and during that time, we’ve been working behind the scenes on updates to improve video processing efficiency. Today, we’re pleased to introduce SAM 3.1.

Meta AI · Jul 4, 2026
12 min
Scaling How We Build and Test Our Most Advanced AIResearch

Scaling How We Build and Test Our Most Advanced AI

As we build more capable, personalized AI, reliability, security, and user protections are more important than ever.

Meta AI · Jul 4, 2026
5 min
Introducing Muse Spark: Scaling Towards Personal SuperintelligenceResearch

Introducing Muse Spark: Scaling Towards Personal Superintelligence

Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

Meta AI · Jul 4, 2026
6 min
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene GenerationResearch

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation

Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretatio...

Papers with Code · Jul 4, 2026
1 min
An Interpretable Deep Learning Framework for Discovery and Clinical Validation of Deep Radiomic Signatures in Tumor ClassificationResearch

An Interpretable Deep Learning Framework for Discovery and Clinical Validation of Deep Radiomic Signatures in Tumor Classification

Imaging signatures are quantitative features extracted from medical images that provide clinically meaningful information for tumor diagnosis, characterizati...

Papers with Code · Jul 3, 2026
2 min
Distributed Attacks in Persistent-State AI ControlResearch

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests (PRs) and time its payload for the PR with the best natural cover. To study the resulting dynamics, we introduce Iterative VibeCoding, a setting for AI control, the study of safely deploying capable but potentially untrusted AI. In Iterative VibeCoding, a coding agent builds software over a sequence of PRs in a persistent codeba

arXiv (cs.AI) · Jul 2, 2026
4 min
LACUNA: A Testbed for Evaluating Localization Precision for LLM UnlearningResearch

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art(SOTA) methods often following a localize-first, unlearn-second paradigm that targets specific model parameters. However, existing benchmarks evaluate unlearning solely at the output level, leaving open the question of whether unlearning truly erases knowledge from a model's parameters or merely obfuscates it, a concern reinforced by the success of resurfacing at

arXiv (cs.AI) · Jul 2, 2026
4 min
Program-as-Weights: A Programming Paradigm for Fuzzy FunctionsResearch

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose fuzzy-function programming: compiling such a function from a natural-language specification into a compact, locally-executable neural artifact. We instantiate this paradigm with Program-as-Weights (PAW), in which a 4B compiler trained on FuzzyBench, a 10M-example dataset we release, emits

arXiv (cs.AI) · Jul 2, 2026
3 min
Online Safety Monitoring for LLMsResearch

Online Safety Monitoring for LLMs

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.

arXiv (cs.AI) · Jul 2, 2026
3 min
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context ReasoningResearch

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. In this work, we propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to

arXiv (cs.AI) · Jul 2, 2026
4 min
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent DebatesResearch

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit objective in the prompt, changes what an agent expresses publicly relative to an off-the-record (OTR) channel elicited under the same condition. We introduce a dual-channel debate framework in which agents produce public utterances that enter the shared history alongside OTR responses that are recorded but never shown to the other participant. Across 10 models, 3 scen

arXiv (cs.AI) · Jul 2, 2026
4 min
Reasoning LLM Improves Speaker Recognition in Long-form TV DramasResearch

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field through two primary contributions. (1) We introduce \textbf{DramaSR-532K}, a large-scale benchmark comprising 532K annotated dialogue lines across more than 900 unique characters, necessitating the integration of auditory, linguistic, and visual cues for speaker recognition. (2) We propose \tex

arXiv (cs.AI) · Jul 2, 2026
3 min
DemoPSD: Disagreement-Modulated Policy Self-DistillationResearch

DemoPSD: Disagreement-Modulated Policy Self-Distillation

On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access. However, recent studies have found that the teacher's dense token-level supervision, conditioned on privileged information, can lead to overfitting to in-domain patterns, suppress exploration, and hurt cross-domain generalization, while also introducing a more fundamental issue: *privileged information leakage*, where the student encodes answer-dependent shor

arXiv (cs.AI) · Jul 2, 2026
4 min
Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic PotentialsResearch

Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts on new architectures and datasets have led to increasingly accurate and general models, the choice of optimizer for training has largely remained unexplored, defaulting to Adam and its variants in the community. Here, we implement and systematically compare a class of recently proposed matrix-structured optimizers, including Muon, SOAP, and the hybrid SOAP-Muon, for training NequIP and Allegro MLIP models. We find that these optimizers can substantially outperform Ada

arXiv (cs.AI) · Jul 2, 2026
3 min
Controllable Sim Agents with Behavior LatentsResearch

Controllable Sim Agents with Behavior Latents

Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test autonomous systems without real-world risk. We introduce Controllable Neural Variational Agents (CNeVA), a controllable simulated-agent framework that learns to infer a per-agent Gaussian behavior latent from per-channel discounted returns via a closed-form conjugate variational update, conditioning a rectified-flow trajectory generator trained on a mixed channel

arXiv (cs.LG) · Jul 2, 2026
3 min
G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning ModelsResearch

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes. We propose a neuro-symbolic approach, ``Guiding with Recurrent Reasoning Models'' (G-RRM), which integrates SE-RRMs with symbolic solvers for constraint satisfaction problems. SE-RRMs act as neural solvers that generate full solution proposals and guide classical symbolic solvers, such as backtracking or SAT-based methods like Glucose 4.1 and CaDiCaL 3.0.0, that produce globally correct solutions. Centrally, we investigate when neural guidance wi

arXiv (cs.AI) · Jul 2, 2026
4 min
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token PruningResearch

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning

Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to preserve critical cues under dense instructions and fine-grained queries. In this paper, we investigate this failure and identify two underlying bottlenecks: the widespread dispersion of textual noise that corrupts dense cross-modal scoring, and the feature fragmentation inherent to standard token selection. To address these issues, we propose Entropy-Aware Dense Pruning (EADP), a framework that reformulates pruning as a structured compression pr

arXiv (cs.AI) · Jul 2, 2026
3 min
TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionResearch

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior. Yet existing test generation and update benchmarks often isolate the test from the code change, and rely on static metadata that does not verify whether a test is executable or semantically tied to the code change. This makes it difficult to evaluate whether a test automation agent understands how a code change should propagate into the test suite. We introduce TestEvo-Bench, a benchmark of test and code co-evolution tasks mined from software reposi

arXiv (cs.AI) · Jul 2, 2026
4 min
Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in ForecastingResearch

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an objective, externally resolved benchmark, this pilot shows that the value of human-AI collaboration depends on a specific, measurable form of human capital. Analyzed at the level of the individual forecaster, hybrid performance is trimodal: most people either deferred to the model (matching it) or used it to rubber-stamp a prior guess (performing worse than the model alone), while a minority engaged in genuine complementary reasoning an

arXiv (cs.AI) · Jul 2, 2026
3 min
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAsResearch

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors fr

arXiv (cs.AI) · Jul 2, 2026
4 min
OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion TransformersResearch

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ) is the natural remedy, yet DiT activations shift across timesteps, prompts, and guidance branches, forcing prior methods to re-fit calibration data for every new checkpoint or modality. We present OrbitQuant, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis. In this basis, a randomized permuted block-Hadamard (RPBH) r

arXiv (cs.AI) · Jul 2, 2026
4 min
Neuron-Aware Data Selection for Annotation-Free LLM Self-DistillationResearch

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels. However, these approaches are not without drawbacks: SFT- and GRPO-based variants suffer out-of-domain perfor

arXiv (cs.AI) · Jul 2, 2026
3 min
Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID DataResearch

Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data

Recent research has introduced distributed self-supervised learning (D-SSL) approaches to leverage vast amounts of unlabeled decentralized data. However, D-SSL faces the critical challenge of data heterogeneity, and there is limited theoretical understanding of how different D-SSL frameworks respond to this challenge. To fill this gap, we present a rigorous theoretical analysis of the robustness of D-SSL frameworks under non-IID (non-independent and identically distributed) settings. Our results show that pre-training with Masked Image Modeling (MIM) is inherently more robust to hete

arXiv (cs.LG) · Jul 2, 2026
3 min
Optimal Stabilizer Testing and Learning with Limited Quantum MemoryResearch

Optimal Stabilizer Testing and Learning with Limited Quantum Memory

We study stabilizer state testing and learning with limited coherent quantum memory. Here an algorithm sequentially receives copies of an unknown $n$-qubit state, but may keep only $k$ qubits of coherent quantum memory between measurements. With unrestricted memory, seminal work of Gross, Nezami and Walter showed how to test $n$-qubit stabilizer states using $6$ copies, which is dimension independent, unlike the learning complexity of $Θ(n)$. We show that this testing-vs-learning separation is lost under memory constraints. More concretely we show that (1) The sample complexity of te

arXiv (cs.LG) · Jul 2, 2026
4 min
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive EnvironmentsResearch

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the Evo

arXiv (cs.AI) · Jul 2, 2026
3 min
Extreme Adaptive Transformer for Time Series ForecastingResearch

Extreme Adaptive Transformer for Time Series Forecasting

Time series forecasting remains challenging when the underlying data contain rare but critical extreme events. This issue is particularly important in hydrologic forecasting, where streamflow distributions are often highly skewed and extreme peaks can have substantial impacts on flood monitoring, water resource management, and early warning systems. Although Transformer-based forecasting models have achieved strong performance by modeling long-range temporal dependencies, they typically treat all time points uniformly and may therefore underrepresent rare extreme patterns. In this pa

arXiv (cs.LG) · Jul 2, 2026
3 min
Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational studyResearch

Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study

Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the assumption that more capability yields better software. This study tested that assumption directly. Ninety independent agent runs built the same application, a real time retrospective board, from one detailed specification, each scored on a fixed 14 criterion functional rubric (42 point maximum) and a visual quality review. The runs spanned several model generations, two agent harnesses, two reasoning effort levels, a testing tool, and tw

arXiv (cs.AI) · Jul 2, 2026
4 min
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approachResearch

Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking difficult and rule-based autograders cannot handle partial credit, equivalent solutions, or syntactic variation. This paper evaluates whether four frontier Large Language Models (GPT, Claude Opus, Gemini, and GLM) can approximate expert judgment when grading short Linux/bash command responses. The study adopts a four-level cognitive taxonomy that combines cognitive complexity and operational impact, ranging from information retrieval (L1)

arXiv (cs.AI) · Jul 2, 2026
4 min
WorldSample: Closed-loop Real-robot RL with World ModellingResearch

WorldSample: Closed-loop Real-robot RL with World Modelling

Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve through trial-and-error interaction beyond the states observed in demonstrations. However, deploying RL on real robots remains constrained by high interaction costs, since each physical rollout is costly and reflects only one realized action-outcome path. To address this challenge, we propose WorldSample, a physically grounded data augmentation framework for real-robot RL that closes a real-synthetic loop between physical rollouts, world-model genera

arXiv (cs.AI) · Jul 2, 2026
4 min
QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity RecognitionResearch

QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition

Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for privacy-sensitive robotic sensing applications. However, multi-agent systems generate heterogeneous and non-independent and identically distributed (non-IID) multimodal sensor streams that degrade conventional FL algorithms, while classical fusion modules introduce substantial parameter overhead and communication cost. This paper proposes QFedAgent, a hybrid quantum-classical personalized FL framework for multi-agent activity recognition. The approa

arXiv (cs.AI) · Jul 2, 2026
3 min
Neuron-Aware Active Few-Shot Learning for LLMsResearch

Neuron-Aware Active Few-Shot Learning for LLMs

Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and use as few-shot demonstrations, effectively reducing human annotation costs while promoting high performance. However, existing methods typically rely on output-level signals for sample identification, such as predictive entropy or semantic similarities with test-time data based on external embeddings, which often overlook models' internal dynamics, which could pinpoint specific knowledge gaps. To bridge this gap, we propose NeuFS, a Neuron-Aware Ac

arXiv (cs.AI) · Jul 2, 2026
3 min
LIME: Learning Intent-aware Camera Motion from Egocentric VideoResearch

LIME: Learning Intent-aware Camera Motion from Egocentric Video

Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vision-language navigation translates instructions to base motion and vision-language-action policies map instructions to manipulation actions, language-conditioned camera motion remains comparatively underexplored as a first-class action. We formulate language-conditioned camera motion generation: given a current RGB observation and a free-form natural-language intent, predict a relative target camera pose fo

arXiv (cs.LG) · Jul 2, 2026
4 min
Q-GAIN: A Python Package for Machine Learning and Physically Informed Analysis ApplicationsResearch

Q-GAIN: A Python Package for Machine Learning and Physically Informed Analysis Applications

Here we describe the quantum gas analysis and inference (Q-GAIN) Python package, which enables rapid deployment of machine learning (ML) and physics-informed analysis techniques for cold-atom experiments. Out of the box, Q-GAIN implements classification, object detection, and physics-informed metrics for feature detection in images of atomic Bose-Einstein condensates (BECs). Q-GAIN encourages a natural, module-based workflow: starting with data loading and preprocessing, followed by ML-based feature identification, and ending with conventional analysis techniques. We demonstrate this

arXiv (cs.LG) · Jul 2, 2026
3 min
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan EnvironmentsResearch

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statist

arXiv (cs.AI) · Jul 2, 2026
3 min
Object-centric LeJEPAResearch

Object-centric LeJEPA

Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets. Aligning representations at the level of objects rather than whole scenes promises greater data efficiency, but doing this in a completely self-supervised way, effectively jointly partitioning a scene and representing its objects, is unstable: the two are locked in a cyclic dependency, partitioning requires meaningful representations, while meaningful representations require consistent partitioning. We sid

arXiv (cs.LG) · Jul 2, 2026
3 min
ACID: Action Consistency via Inverse Dynamics for Planning with World ModelsResearch

ACID: Action Consistency via Inverse Dynamics for Planning with World Models

Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard planning cost judges a candidate solely by how close its predicted terminal state lies to the goal, leaving the realizability of the intermediate transitions unchecked -- a predicted trajectory can look convincing while the environment rollout drifts away from it. In this paper, we propose ACID, a decision-time planning framework that introduces cycle action consistency: the action inferred backward from a predicted transition by an inverse dynamics mo

arXiv (cs.AI) · Jul 2, 2026
3 min
Fast Multi-dimensional Refusal Subspaces via RFM-AGOPResearch

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along single linear directions, but recent findings suggest complex behaviours, such as the refusal to answer harmful queries, live in multi-dimensional subspaces. However, existing methods for extracting these subspaces are computationally expensive, which becomes prohibitive on reasoning models who produce long reasoning traces. By adapting the Recursive Feature Machine (RFM) algorithm -- which can be computed effi

arXiv (cs.AI) · Jul 2, 2026
3 min
WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMsResearch

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient GPUs, but operators currently lack the tools to do so without exhaustively profiling each combination. While some predictive models exist, they still require profiling data and struggle to generalize to hardware unseen during training. To address this, we introduce \textit{WattGPU}, featuring two predictive models for mean GPU power draw and Inter-Token Latency (ITL). Our approach leverag

arXiv (cs.LG) · Jul 2, 2026
4 min
DecompRL: Solving Harder Problems by Learning Modular Code GenerationResearch

DecompRL: Solving Harder Problems by Learning Modular Code Generation

How can Large Language Models (LLMs) solve problems they currently cannot? Repeated sampling scales test-time compute but GPU cost grows linearly with attempts, while reinforcement learning (RL) with verifiable rewards improves single-attempt accuracy at the expense of sample diversity. Both strategies ultimately fail when the base policy has near-zero probability of producing a correct solution: no amount of sampling or gradient signal can overcome a search space that is simply too large. We take a different approach: rather than sampling harder, we make the task easier by decomposi

arXiv (cs.LG) · Jul 2, 2026
3 min
Bringing Agentic Search to Earth Observation Data DiscoveryResearch

Bringing Agentic Search to Earth Observation Data Discovery

NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search system, deployed as a public service for the geoscience community, that takes a natural-language research query and returns the matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic search. From the NASA Earth Observation Knowledge Graph (N

arXiv (cs.LG) · Jul 2, 2026
3 min
Transformer Geometry Observatory TGO-II: Representational Similarity ObservatoryResearch

Transformer Geometry Observatory TGO-II: Representational Similarity Observatory

While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal representations throughout training remains insufficiently understood. Existing analyses primarily focus on attention mechanisms and downstream performance, leaving the evolution of representation geometry largely unexplored. In this work, we present Transformer Geometry Observatory-II (TGO-II), a representation geometry analysis framework designed to investigate how Transformer representations evolve during supervised training. TGO-II

arXiv (cs.LG) · Jul 2, 2026
4 min
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent GeometryResearch

The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry

Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is intrinsic or frame-dependent. Constructing within-instance correlation matrices from IPIP-50 responses, we analyze geometry on SPD manifolds under manipulated question orderings in GPT-4o simulating American and Chinese-American personas. We find that persona expression comprises two dissociable components: aggregated features (Big Five scores) degrade under randomization (21% drop) but are frame-

arXiv (cs.LG) · Jul 2, 2026
3 min
Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory GatesResearch

Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates

Quantum Fast-Weight Programmers (QFWPs) store temporal information in dynamically programmed variational-circuit parameters rather than in nonlinear recurrent hidden states, offering a practical route to quantum sequence modeling. Self-Modulating QFWP improves this framework by using input-dependent gates for both new fast-weight updates and the accumulated fast-weight state, but its unbounded old-state multiplier can diverge in long-sequence regimes. We propose a bounded old-state modulation rule that applies a sign-preserving tanh gate only to the recurrent memory branch while leav

arXiv (cs.LG) · Jul 2, 2026
4 min
Self-Gating Attention for Efficient Time Series ForecastingResearch

Self-Gating Attention for Efficient Time Series Forecasting

Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal dependencies across historical timestamps. However, standard self-attention has quadratic time and memory complexity with respect to the look-back length. This cost may limit its use in resource-constrained or high-throughput forecasting systems, where fast and memory-efficient inference is important. Through qualitative and quantitative analyses, we observe that self-attention maps in time series forecasting often contain redundant patt

arXiv (cs.LG) · Jul 2, 2026
4 min
Start building with Nano Banana 2 Lite and Gemini Omni FlashResearch

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Scale your ideas with Nano Banana 2 Lite, our fastest, most cost-efficient Gemini Image model, and Gemini Omni Flash for high-quality video and conversational editing.

Google DeepMind · Jun 30, 2026
7 min
Introducing computer use in Gemini 3.5 FlashResearch

Introducing computer use in Gemini 3.5 Flash

A look at the built-in computer use tool in Gemini 3.5 Flash.

Google DeepMind · Jun 24, 2026
3 min
Google DeepMind and A24 announce first-of-its-kind research partnershipResearch

Google DeepMind and A24 announce first-of-its-kind research partnership

Today, Google DeepMind and A24 are announcing a first-of-its-kind partnership focused on research. The collaboration pairs a world-leading research lab with the industry…

Google DeepMind · Jun 22, 2026
1 min
Securing internal systems against increasingly capable and imperfectly aligned AIResearch

Securing internal systems against increasingly capable and imperfectly aligned AI

Discover our AI Control Roadmap: a defense-in-depth system to securely manage advanced, potentially misaligned AI agents.

Google DeepMind · Jun 18, 2026
6 min
Unlocking UK house-building with AI-accelerated planningResearch

Unlocking UK house-building with AI-accelerated planning

Google DeepMind is working alongside the UK government to co-develop an AI-powered prototype to help cut application decision times by 50%.

Google DeepMind · Jun 16, 2026
5 min
Google DeepMind and partners announce multi-agent safety research funding call.Research

Google DeepMind and partners announce multi-agent safety research funding call.

Google DeepMind and partners are announcing a new technical research funding call of up to $10M for researchers worldwide to strengthen multi-agent safety.

Google DeepMind · Jun 11, 2026
4 min
DiffusionGemma: 4x faster text generationResearch

DiffusionGemma: 4x faster text generation

An overview of DiffusionGemma, an exceptionally fast text generation model with up to 4x faster speeds.

Google DeepMind · Jun 10, 2026
5 min
Gemini’s guided learning: results from a randomized controlled trial in Sierra LeoneResearch

Gemini’s guided learning: results from a randomized controlled trial in Sierra Leone

Google DeepMind shares results from a randomized controlled trial in Sierra Leone, measuring the impact of AI in education on student learning and engagement.

Google DeepMind · Jun 9, 2026
5 min
Powering the future of robotics in EuropeResearch

Powering the future of robotics in Europe

Google DeepMind Accelerator selects 15 robotics companies from across Europe to join the program. Providing 3 months of intensive mentorship and technical support, enabl…

Google DeepMind · Jun 9, 2026
5 min
Fluid, natural voice translation with Gemini 3.5 Live TranslateResearch

Fluid, natural voice translation with Gemini 3.5 Live Translate

Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.

Google DeepMind · Jun 9, 2026
5 min
Introducing Gemma 4 12B: a unified, encoder-free multimodal modelResearch

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

Google DeepMind · Jun 3, 2026
4 min
Google DeepMind & Singapore: National AI partnershipResearch

Google DeepMind & Singapore: National AI partnership

Google DeepMind and Singapore partner to apply frontier AI to address challenges across health, education, sustainability and more through the National Partnerships for AI initiative.

Google DeepMind · May 20, 2026
5 min
AI breakthrough: WeatherNext predicts Hurricane MelissaResearch

AI breakthrough: WeatherNext predicts Hurricane Melissa

Discover how our WeatherNext AI model helps the National Hurricane Center predict Hurricane Melissa's Category 5 landfall in Jamaica

Google DeepMind · May 19, 2026
6 min
Co-ScientistResearch

Co-Scientist

Untangling the mysteries of aging

Google DeepMind · May 19, 2026
2 min
Co-ScientistResearch

Co-Scientist

Fast-tracking infectious disease research

Google DeepMind · May 19, 2026
2 min
Co-ScientistResearch

Co-Scientist

Breakthroughs in liver disease research

Google DeepMind · May 19, 2026
2 min
Co-ScientistResearch

Co-Scientist

Driving creative collaboration in research

Google DeepMind · May 19, 2026
2 min
Co-ScientistResearch

Co-Scientist

Finding new treatments for liver fibrosis

Google DeepMind · May 19, 2026
2 min
Co-ScientistResearch

Co-Scientist

Accelerating cellular aging research

Google DeepMind · May 19, 2026
2 min
Making it easier to understand how content was created and editedResearch

Making it easier to understand how content was created and edited

We're expanding our tools to help you understand how content was created and edited across the web.

Google DeepMind · May 19, 2026
5 min
Gemini for Science: AI experiments and tools for a new era of discoveryResearch

Gemini for Science: AI experiments and tools for a new era of discovery

Gemini for Science is a new collection of science tools and experiments to expand the scale and precision of scientific exploration.

Google DeepMind · May 19, 2026
6 min
Introducing Gemini OmniResearch

Introducing Gemini Omni

Introducing Gemini Omni, which allows you to create anything from any input and edit naturally using conversational language.

Google DeepMind · May 19, 2026
8 min
Simulate real-world places with Project Genie and Street ViewResearch

Simulate real-world places with Project Genie and Street View

We’re connecting Project Genie with nearly 20 years of Google Street View imagery so you can create new worlds anchored in reality.

Google DeepMind · May 19, 2026
4 min
Google Antigravity Blog: Introducing Google Antigravity 2.0Research

Google Antigravity Blog: Introducing Google Antigravity 2.0

Introducing Google Antigravity 2.0

Google DeepMind · May 17, 2026
6 min
We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks.Research

We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks.

The Asia-Pacific region is a global engine for economic growth, but it's also highly vulnerable to climate change. While green technologies are gaining momentum, a recen…

Google DeepMind · May 17, 2026
1 min