EventsMOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published
This submission belongs to the session 03. USEDAT-02: USA-Europe Data Analysis Training Program Workshop, Cambridge, UK-Bilbao, Spain-Miami, USA, 2016 of the event MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed.
Published date
27 Dec, 2016
Citation
Qian Zhou, Xiaofang Zhang, Peng Zhang, Quan Liu, Adaptive Exploration in Stochastic Multi-armed Bandit Problem, in Proceedings of MOL2NET'16, Conference on Molecular, Biomed., Comput. & Network Science and Engineering, 2nd ed., 15 October–20 October 2022, MDPI: Basel, Switzerland, doi: 10.3390/mol2net-02-03848
Share
Email
Facebook
Twitter
LinkedIn

Adaptive Exploration in Stochastic Multi-armed Bandit Problem

Xiaofang Zhang 1,2
Peng Zhang 1
1. School of Computer Science and Technology, Soochow University, Suzhou, Jiangsu, 215006
2. State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023
3. Collaborative Innovation Center of Novel Software Technology and Industrialization, Nanjing, 210000
4. Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, Changchun, 130012
Abstract

The multi-armed bandit (MAB) problem is classic problem of the exploration versus exploitation dilemma in reinforcement learning. As an archetypal MAB problem, the stochastic multi-armed bandit (SMAB) problem is the base of many new MAB problems. To solve the problems of weak theoretical analysis and single information used in existing SMAB methods, this paper presents "the Chosen Number of Arm with Minimal Value" (CNAMV), a method for balancing exploration and exploitation adaptively. Theoretically, the upper bound of CNAMV’s regret is proved, that is the loss due to the fact that the globally optimal policy is not followed all the times. Experimental results show that CNAMV yields greater reward and smaller regret with high efficiency than commonly used methods such as ε-greedy, softmax, or UCB1. Therefore the CNAMV can be an effective SMAB method.

Keywords
reinforcement learning
stochastic multi-armed bandit problem
exploration
exploitation
adaptation
Poster
MOL2NET-02-SUIWML_Adaptive Exploration in Stochastic Multi-armed Bandit Problem.pdf
Immune protection against Trypanosoma cruzi induced by TcVac1 vaccine in a murine model using an intradermal/electroporation protocol
Study on Optimal Control Strategy of Automatic Transmission Based on Policy Search