Stochastic Modeling and Applied Research of Technology

PDF

Keywords

Two-armed Bandit Problem
UCB strategy
Batch Processing
Invariant Description
Variable Batch Size

How to Cite

Ershov, M., & Voroshilov, A. (2026). Stochastic differential equations for limiting description of UCB rule for Gaussian multi-armed bandits. Stochastic Modeling and Applied Research of Technology, 4, 11-21. https://doi.org/10.57753/SMARTY.2026.96.62.002

Abstract

For the Gaussian two-armed bandit, which occurs during the analysis of batch data processing, the case of variable batch size is considered. This optimal control problem has a classical interpretation as a game with nature. The player’s payment function is the mathematical expectation of the regrets of full income caused by incomplete information. The control goal is considered in a minimax setting, and a UCB strategy is used to achieve it. An invariant description of a strategy with a unit control horizon is constructed for the problem under consideration. According to the invariant description, the normalized regrets depend on the number of batches and do not depend on their size. Increasing the batch size during control allows the player to process more data with a constant number of actions (batches). In the area of ‘close’ distributions, a slight increase in normalized regrets was observed, but in ‘far’ distributions, a significant decrease in these was observed.

https://doi.org/10.57753/SMARTY.2026.96.62.002
PDF
Creative Commons Attribution License Logo

This work is licensed under a Creative Commons Attribution 4.0 International License.