Archives AI News

PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection

arXiv:2605.24171v1 Announce Type: new Abstract: Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and…

Multi-Alignment Contrastive Learning for Enzyme–Reaction Retrieval

arXiv:2512.08508v2 Announce Type: replace-cross Abstract: Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent representation-learning methods formulate this problem as enzyme–reaction matching, where paired enzymes and reactions are embedded into…

Rapid mixing in positively weighted restricted Boltzmann machines

arXiv:2604.00963v2 Announce Type: replace-cross Abstract: We show polylogarithmic mixing time bounds for the alternating-scan sampler for positively weighted restricted Boltzmann machines. This is done via analysing the same chain and the Glauber dynamics for ferromagnetic two-spin systems, where we obtain…

Characterizing the Representational Capacity of Neural Processes

arXiv:2605.24210v1 Announce Type: new Abstract: What functions can Neural Processes represent? We analyze the representational capacity of popular NP architectures: Conditional Neural Processes (CNPs), Attentive Neural Processes (ANPs), Transformer Neural Processes (TNPs), and their latent variants. We prove these architectures…

Is TabPFN the Silver Bullet for Insurance Pricing?

arXiv:2605.22892v2 Announce Type: replace-cross Abstract: Modelling claim frequency and severity for non-life insurance pricing predominantly relies on generalised linear models, with gradient-boosted machines as the leading machine learning alternative. Tabular foundation models (TFMs) present a fundamentally different inference paradigm. By…

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

arXiv:2605.24216v1 Announce Type: new Abstract: Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horizon attack patterns. Agents may pursue hidden objectives while maintaining superficially benign behavior, making detection difficult even…

Rao-Blackwellized Score Matching on Manifolds

arXiv:2605.25567v1 Announce Type: cross Abstract: We study denoising score matching (DSM) when the latent distribution is supported on a smooth embedded manifold $M subset mathbb{R}^D$. Under ambient Gaussian corruption, the tangent denoising target contains a singular normal-fiber noise channel whose…

DeGRe: Dense-supervised Generative Reranking for Recommendation

arXiv:2605.25749v1 Announce Type: cross Abstract: In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequences within an exponentially large permutation space. Recent studies have shifted towards end-to-end generative…