IS Atlas
misq·2025년 11월 18일

Deep Pareto Reinforcement Learning for Multi-Objective Recommender Systems

Pan Li, Alexander Tuzhilin

MIS Quarterly

2
피인용
3.2
FWCI
0
IS/마케팅/OM 탑저널 피인용
0
IS/마케팅/OM 탑저널 참고문헌
01Abstract

Optimizing multiple objectives simultaneously is an important task for recommendation platforms to improve their performance. However, this task is particularly challenging since the relationships between different objectives are heterogeneous across different consumers and dynamically fluctuate according to different contexts, resulting in a Pareto-frontier in the result of recommendations, where the improvement of any objective comes at the cost of others. Existing multi-objective recommender systems do not systematically consider such dynamic relationships; instead, they balance between these objectives in a static and uniform manner, resulting in only suboptimal recommendation performance. In this paper, we propose a Deep Pareto Reinforcement Learning (DeepPRL) method, where we (1) comprehensively model the complex relationships between multiple recommendation objectives; (2) effectively capture personalized and contextual consumer preferences for each objective; (3) optimize both the short-term and the long-term recommendation performance. As a result, our method achieves significant Pareto-dominance over the state-of-the-art baselines across four offline experiments. Furthermore, we conducted a controlled experiment on Alibaba's video streaming platform, where our method simultaneously improved three conflicting business objectives significantly over the latest production system, demonstrating its tangible economic impact in practice.

02연구 흐름

불러오는 중…

03비슷한 논문

불러오는 중…

04이후 연구

불러오는 중…

05선행 연구

불러오는 중…

06서지 정보