Archives AI News

Model-Based Learning of Near-Optimal Finite-Window Policies in POMDPs

arXiv:2604.01024v2 Announce Type: replace Abstract: We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). A common approach to learning under partial observability is to approximate unbounded history dependencies using finite action-observation windows. This induces…

AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning

arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of…

Contextual Online Uncertainty-Aware Preference Learning for Human Feedback

arXiv:2504.19342v4 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a novel statistical framework to simultaneously conduct the online…

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing…

Bayes with No Shame: Admissibility Geometries of Predictive Inference

arXiv:2603.05335v3 Announce Type: replace-cross Abstract: Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate of optimality. We study four criterion-relative geometries: Blackwell risk dominance, anytime-valid admissibility, fixed-level marginal coverage with expected-length efficiency…

Contextual Online Uncertainty-Aware Preference Learning for Human Feedback

arXiv:2504.19342v4 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a novel statistical framework to simultaneously conduct the online…