IS Atlas
ms·1975년 7월 1일

On Dynamic Programming with Unbounded Rewards

Steven A. Lippman

Management Science

167
피인용
73.5
FWCI
6
IS/마케팅/OM 탑저널 피인용
7
IS/마케팅/OM 탑저널 참고문헌
01Abstract

Using the technique employed by the author in an earlier paper, the existence of an optimal stationary policy that can be obtained from the usual functional equation is again established in the presence of a bound (not necessarily polynomial) on the one-period reward of a semi-Markov decision process. This is done for both the discounted and the average cost case. In addition to allowing an uncountable state space, the law of motion of the system is rather general in that we permit any state to be reached in a single transition. There is, however, a bound on a weighted moment of the next state reached. Finally, we indicate the applicability of these results.

02연구 흐름

불러오는 중…

03비슷한 논문

불러오는 중…

04이후 연구

불러오는 중…

05선행 연구

불러오는 중…

06서지 정보