Archives AI News

Robust Reasoning Benchmark

arXiv:2604.08571v2 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting. We introduce the Robust Reasoning Benchmark (RRB), a pipeline of 13 deterministic textual…

On the Wasserstein Gradient Flow Interpretation of Drifting Models

arXiv:2605.05118v2 Announce Type: replace Abstract: Recently, Deng et al. (2026) proposed Generative Modeling via Drifting (GMD), a novel framework for generative tasks. This note presents an analysis of GMD through the lens of Wasserstein Gradient Flows (WGF), i.e., the path…

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

arXiv:2605.22200v1 Announce Type: cross Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing…

Reinforcement learning for ion shuttling on trapped-ion quantum computers

arXiv:2605.22463v1 Announce Type: cross Abstract: Scalable trapped-ion quantum computing is commonly realized with modular chips that feature distinct zones with specific functionalities, such as storage, state preparation, and gate execution. To execute a quantum circuit, the ions must be transported…

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

arXiv:2511.07885v4 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local…