TY - JOUR AU - Ershov, Maxim AU - Voroshilov, Albert PY - 2026 Y1 - 2026/08/28 TI - Stochastic differential equations for limiting description of UCB rule for Gaussian multi-armed bandits JF - Stochastic Modeling and Applied Research of Technology JA - SMARTY VL - 4 DO - 10.57753/SMARTY.2026.96.62.002 UR - http://smarty.karelia.website/article/4/2/ SP - 11 LP - 21 SN - 2782-4705 AB - For the Gaussian two-armed bandit, which occurs during the analysis of batch data processing, the case of variable batch size is considered. This optimal control problem has a classical interpretation as a game with nature. The player's payment function is the mathematical expectation of the regrets of full income caused by incomplete information. The control goal is considered in a minimax setting, and a UCB strategy is used to achieve it. An invariant description of a strategy with a unit control horizon is constructed for the problem under consideration. According to the invariant description, the normalized regrets depend on the number of batches and do not depend on their size. Increasing the batch size during control allows the player to process more data with a constant number of actions (batches). In the area of 'close' distributions, a slight increase in normalized regrets was observed, but in 'far' distributions, a significant decrease in these was observed. ER -