ms·1972년 12월 1일
Constrained Markov Decision Chains
Cyrus Derman, Arthur F. Veinott
Management Science
29
피인용
0.0
FWCI
0
IS/마케팅/OM 탑저널 피인용
0
IS/마케팅/OM 탑저널 참고문헌
- 주제동적계획과 확률최적화 · 생산·최적화
01Abstract
We consider finite state and action discrete time parameter Markov decision chains. The objective is to provide an algorithm for finding a policy that minimizes the long-run expected average cost when there are linear side conditions on the limit points of the expected state-action frequencies. This problem has been solved previously only for the case where every deterministic stationary policy has at most one ergodic class. This note removes that restriction by applying the Dantzig-Wolfe decomposition principle.
02연구 흐름
불러오는 중…
03비슷한 논문
불러오는 중…
04이후 연구
불러오는 중…
05선행 연구
불러오는 중…
06서지 정보
- 저널Management Science · 19(4-part-1) · 389–390
- 토픽Reinforcement Learning in Robotics · Artificial Intelligence
- DOI10.1287/mnsc.19.4.389
- 저자Cyrus Derman, Arthur F. Veinott