Entity

On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits

We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimization or pure exploration paradigms. The goal is to design an adaptive policy that strategically explores the bandit instance in the initial free exploration phase and minimizes the cumulative regret in the subsequent phase. We formalize this regret minimization with free exploration problem and identify an interesting

Paper · arXiv

cs.LG

Authors: Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan
Published: 2026-05-25
Categories: cs.LGcs.AIcs.ITstat.ML

Abstract ↗

via arXiv · 2605.25789