Archives AI News

MuLoCo: Muon is a practical inner optimizer for DiLoCo

arXiv:2505.23725v2 Announce Type: replace Abstract: DiLoCo is a powerful framework for training large language models (LLMs), enabling larger optimal batch sizes and increased accelerator utilization under networking constraints. However, DiLoCo’s performance has been shown to degrade as the number of…

February 26, 2026

Generative Bayesian Computation as a Scalable Alternative to Gaussian Process Surrogates

arXiv:2602.21408v1 Announce Type: new Abstract: Gaussian process (GP) surrogates are the default tool for emulating expensive computer experiments, but cubic cost, stationarity assumptions, and Gaussian predictive distributions limit their reach. We propose Generative Bayesian Computation (GBC) via Implicit Quantile Networks…

February 26, 2026

Federated Learning in Offline and Online EMG Decoding: A Privacy and Performance Perspective

arXiv:2507.12652v2 Announce Type: replace Abstract: Neural interfaces offer a pathway to intuitive, high-bandwidth interaction, but the sensitive nature of neural data creates significant privacy hurdles for large-scale model training. Federated learning (FL) has emerged as a promising privacy-preserving solution, yet…

February 26, 2026

Benchmarking State Space Models, Transformers, and Recurrent Networks for US Grid Forecasting

arXiv:2602.21415v1 Announce Type: new Abstract: Selecting the right deep learning model for power grid forecasting is challenging, as performance heavily depends on the data available to the operator. This paper presents a comprehensive benchmark of five modern neural architectures: two…

February 26, 2026

Characterization and Learning of Causal Graphs with Latent Confounders and Post-treatment Selection from Interventional Data

arXiv:2509.25800v2 Announce Type: replace Abstract: Interventional causal discovery seeks to identify causal relations by leveraging distributional changes introduced by interventions, even in the presence of latent confounders. Beyond the spurious dependencies induced by latent confounders, we highlight a common yet…

February 26, 2026

New method could increase LLM training efficiency

By leveraging idle computing time, researchers can double the speed of model training while preserving accuracy.

February 26, 2026

Tackling industry’s burdensome bubble problem

MIT researchers uncovered the physics behind bubble-removing membranes that could improve bioreactors, chemical production, and more.

February 26, 2026

Spurious Rewards: Rethinking Training Signals in RLVR

arXiv:2506.10947v2 Announce Type: replace-cross Abstract: We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or even negative correlation with the correct answer. For…

February 26, 2026

Optimizer choice matters for the emergence of Neural Collapse

arXiv:2602.16642v3 Announce Type: replace Abstract: Neural Collapse (NC) refers to the emergence of highly symmetric geometric structures in the representations of deep neural networks during the terminal phase of training. Despite its prevalence, the theoretical understanding of NC remains limited.…

February 26, 2026

Overparameterized Multiple Linear Regression as Hyper-Curve Fitting

arXiv:2404.07849v2 Announce Type: replace-cross Abstract: This work demonstrates that applying a fixed-effect multiple linear regression (MLR) model to an overparameterized dataset is mathematically equivalent to fitting a hyper-curve parameterized by a single scalar. This reformulation shifts the focus from global…

February 26, 2026