IS Atlas
pom·2026년 2월 4일

Revenue Management With Nonparametric Demand Learning and Product Returns

Sheng Ji, Yi Yang

Production and Operations Management

0
피인용
0.0
FWCI
0
IS/마케팅/OM 탑저널 피인용
48
IS/마케팅/OM 탑저널 참고문헌
01Abstract

Product returns are prevalent in practice. Many retailers provide lenient free return policies but with specific return window within which customers are allowed to return products. Motivated by this phenomenon, we consider a single-product online learning and pricing problem with stochastic product returns. A salient feature is that the demand function, depending on price and return window decisions, is initially unknown and must be learned on the fly. The retailer thus faces the classic exploration–exploitation trade-off. Moreover, we consider an inventory constraint, introducing an additional trade-off between earning revenue and managing inventory. We propose a modeling framework to integrate pricing and return window decisions, and develop a deterministic fluid model that serves as the full-information benchmark. To tackle the learning problem, we design a novel nonparametric learning algorithm that seamlessly integrates inverse stochastic gradient descent (SGD) and Upper Confidence Bound (UCB) methods. Under mild assumptions on demand and revenue functions, we establish a regret upper bound for our learning algorithm as <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>O</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msqrt> <mml:mi>W</mml:mi> <mml:mi>T</mml:mi> </mml:msqrt> <mml:mi>log</mml:mi> <mml:mspace width="0.2em"/> <mml:mi>T</mml:mi> <mml:mo stretchy="false">)</mml:mo> </mml:math> , where <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>W</mml:mi> </mml:math> denotes the number of return window candidates and <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline" overflow="scroll"> <mml:mi>T</mml:mi> </mml:math> denotes the time horizon. This result aligns with lower bounds established in both online pricing and multi-armed bandit (MAB) literature. Numerical experiments are conducted to verify the effectiveness and robustness of our algorithm across various environments. From an operational standpoint, retailers can use our learning framework as a decision-support tool to identify the optimal price and return window.

02연구 흐름

불러오는 중…

03비슷한 논문

불러오는 중…

04이후 연구

불러오는 중…

05선행 연구

불러오는 중…

06서지 정보