Bayesian dithering for learning: Asymptotically optimal policies in dynamic pricing
Woonghee Tim Huh, Michael Jong Kim, Mei-Chun Lin
Production and Operations Management
- 주제동적 가격책정 · 공급망관리
- 방법
- 현상
We consider a dynamic pricing and learning problem where a seller prices multiple products and learns from sales data about unknown demand. We study the parametric demand model in a Bayesian setting. To avoid the classical problem of incomplete learning, we propose dithering policies under which prices are probabilistically selected in a neighborhood surrounding the myopic optimal price. By analyzing the effect of dithering in facilitating learning, we establish regret upper bounds for three typical settings of demand model. We show that the dithering policy achieves an upper bound of order <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:semantics definitionURL="" encoding=""> <mml:mrow> <mml:mi>log</mml:mi> <mml:mi>T</mml:mi> </mml:mrow> <mml:annotation encoding="">$\log T$</mml:annotation> </mml:semantics> </mml:math> when the parameter set is finite. It can be modified to achieve a constant regret bound under an additional assumption. We also prove an upper bound of order <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:semantics definitionURL="" encoding=""> <mml:msqrt> <mml:mrow> <mml:mi>T</mml:mi> <mml:mi>log</mml:mi> <mml:mi>T</mml:mi> </mml:mrow> </mml:msqrt> <mml:annotation encoding="">$\sqrt {T\log T}$</mml:annotation> </mml:semantics> </mml:math> when the parameter set is compact and convex. Each bound matches (up to a logarithmic factor) the existing lower bound of any pricing policy. In this way, we show that dithering policies achieve asymptotically optimal performance in three different parameter settings, which demonstrates dithering as a unified approach to strike the balance between exploration and exploitation.
불러오는 중…
불러오는 중…
불러오는 중…
불러오는 중…
- 저널Production and Operations Management · 31(9) · 3576–3593
- 토픽Advanced Bandit Algorithms Research · Management Science and Operations Research
- DOI10.1111/poms.13786
- 저자Woonghee Tim Huh, Michael Jong Kim, Mei-Chun Lin