Archives AI News

Harnesses for Inference-Time Alignment over Execution Trajectories

arXiv:2605.21516v1 Announce Type: new Abstract: Harness engineering has emerged as an important inference-time technique for large language model (LLM) agents, aiming to improve long-term performance through task decomposition and guided execution. However, more elaborate harnesses are not uniformly better: increasing…

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

arXiv:2605.22200v1 Announce Type: cross Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing…

Reinforcement learning for ion shuttling on trapped-ion quantum computers

arXiv:2605.22463v1 Announce Type: cross Abstract: Scalable trapped-ion quantum computing is commonly realized with modular chips that feature distinct zones with specific functionalities, such as storage, state preparation, and gate execution. To execute a quantum circuit, the ions must be transported…

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

arXiv:2511.07885v4 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local…

TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

arXiv:2605.14738v2 Announce Type: replace Abstract: Recent work has promoted task-aware layer pruning as a way to improve model performance on particular tasks, as shown by TALE. In this paper, we investigate when such improvements occur and why. We show first…

Harnesses for Inference-Time Alignment over Execution Trajectories

arXiv:2605.21516v1 Announce Type: new Abstract: Harness engineering has emerged as an important inference-time technique for large language model (LLM) agents, aiming to improve long-term performance through task decomposition and guided execution. However, more elaborate harnesses are not uniformly better: increasing…

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

arXiv:2605.22200v1 Announce Type: cross Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing…

Reinforcement learning for ion shuttling on trapped-ion quantum computers

arXiv:2605.22463v1 Announce Type: cross Abstract: Scalable trapped-ion quantum computing is commonly realized with modular chips that feature distinct zones with specific functionalities, such as storage, state preparation, and gate execution. To execute a quantum circuit, the ions must be transported…