Simple fixes that accommodate switching costs in multi-armed bandits.

التفاصيل البيبلوغرافية
العنوان:	Simple fixes that accommodate switching costs in multi-armed bandits.
المؤلفون:	Teymourian, Ehsan¹ (AUTHOR) teymouex@jmu.edu, Yang, Jian² (AUTHOR)
المصدر:	European Journal of Operational Research. Feb2025, Vol. 320 Issue 3, p616-627. 12p.
مصطلحات موضوعية:	SWITCHING costs, COST analysis, TIME perspective, AMBIGUITY, REGRET
مستخلص:	When switching costs are added to the multi-armed bandit (MAB) problem where the arms' random reward distributions are previously unknown, usually quite different techniques than those for pure MAB are required. We find that two simple fixes on the existing upper-confidence-bound (UCB) policy can work well for MAB with switching costs (MAB-SC). Two cases should be distinguished. One is with positive-gap ambiguity where the performance gap between the leading and lagging arms is known to be at least some δ > 0. For this, our fix is to erect barriers that discourage frivolous arm switchings. The other is with zero-gap ambiguity where absolutely nothing is known. We remedy this by forcing the same arms to be pulled in increasingly prolonged intervals. As usual, the effectivenesses of our fixes are measured by the worst average regrets over long time horizons T. When the barriers are fixed at δ / 2 , we can accomplish a ln (T) -sized regret bound for the positive-gap case. When intervals are such that n of them occupy n 2 periods, we can achieve the best possible T 1 / 2 -sized regret bound for the zero-gap case. Other than UCB, these fixes can be applied to a learning while doing (LWD) heuristic to reach satisfactory results as well. While not yet with the best theoretical guarantees, the LWD-based policies have empirically outperformed those based on UCB and other known alternatives. Numerically competitive policies still include ones resulting from interval-based fixes on Thompson sampling (TS). • We seek good adaptive policies for the MAB problem with switching costs (MAB-SC). • We investigate two different degrees of ambiguity: Positive- and zero-gap cases. • For them, we propose two meta-heuristic fixes to cope with switching costs. • The resulting policies attain near-optimal regrets in the respective ambiguity cases. • Empirically, adopted policies outperform many known policies for the MAB-SC. [ABSTRACT FROM AUTHOR]
	Copyright of European Journal of Operational Research is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
قاعدة البيانات:	Business Source Index

الوصف
تدمد:	03772217
DOI:	10.1016/j.ejor.2024.09.017