Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
arXiv:2509.08933v2 Announce Type: replace Abstract: We study the problem of learning the optimal policy in a discounted, infinite-horizon reinforcement learning (RL) setting in the presence of adversarially corrupted rewards. To address this problem, we develop a novel robust variant of…
