Archives AI News

Robust Reasoning Benchmark

arXiv:2604.08571v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their underlying reasoning processes remain highly overfit to standard textual formatting. We propose a perturbation pipeline consisting of 14 techniques to evaluate robustness…

Contribution of task-irrelevant stimuli to drift of neural representations

arXiv:2510.21588v2 Announce Type: replace-cross Abstract: Biological and artificial learners are inherently exposed to a stream of data and experience throughout their lifetimes and must constantly adapt to, learn from, or selectively ignore the ongoing input. Recent findings reveal that, even…

Automatic Self-supervised Learning for Social Recommendations

arXiv:2412.18735v3 Announce Type: replace-cross Abstract: In recent years, researchers have leveraged social relations to enhance recommendation performance. However, most existing social recommendation methods require carefully designed auxiliary social tasks tailored to specific scenarios, which depend heavily on domain knowledge and…

Sim-to-Real Transfer for Muscle-Actuated Robots via Generalized Actuator Networks

arXiv:2604.09487v1 Announce Type: cross Abstract: Tendon drives paired with soft muscle actuation enable faster and safer robots while potentially accelerating skill acquisition. Still, these systems are rarely used in practice due to inherent nonlinearities, friction, and hysteresis, which complicate modeling…

Online Quantile Regression for Nonparametric Additive Models

arXiv:2604.08969v1 Announce Type: cross Abstract: This paper introduces a projected functional gradient descent algorithm (P-FGD) for training nonparametric additive quantile regression models in online settings. This algorithm extends the functional stochastic gradient descent framework to the pinball loss. An advantage…

Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies

arXiv:2604.09189v1 Announce Type: cross Abstract: LLMs internalize safety policies through RLHF, yet these policies are never formally specified and remain difficult to inspect. Existing benchmarks evaluate models against external standards but do not measure whether models understand and enforce their…