IS Atlas
ms·1978년 7월 1일

Modified Policy Iteration Algorithms for Discounted Markov Decision Problems

Martin L. Puterman, Moon Chirl Shin

Management Science

238
피인용
4.2
FWCI
2
IS/마케팅/OM 탑저널 피인용
20
IS/마케팅/OM 탑저널 참고문헌
01Abstract

In this paper we study a class of modified policy iteration algorithms for solving Markov decision problems. These correspond to performing policy evaluation by successive approximations. We discuss the relationship of these algorithms to Newton-Kantorovich iteration and demonstrate their covergence. We show that all of these algorithms converge at least as quickly as successive approximations and obtain estimates of their rates of convergence. An analysis of the computational requirements of these algorithms suggests that they may be appropriate for solving problems with either large numbers of actions, large numbers of states, sparse transition matrices, or small discount rates. These algorithms are compared to policy iteration, successive approximations, and Gauss-Seidel methods on large randomly generated test problems.

02연구 흐름

불러오는 중…

03비슷한 논문

불러오는 중…

04이후 연구

불러오는 중…

05선행 연구

불러오는 중…

06서지 정보