We consider zero-sum stochastic games in continuous time with controlled Markov chains and with risk-sensitive average cost criterion. Here the transition and the cost rates may be unbounded. We prove the existence of the value of the game and a saddle-point equilibrium in the class of all stationary strategies under a Lyapunov stability condition. This is accomplished by establishing the existence of a principal eigenpair for the corresponding Hamilton-Jacobi-Isaacs (HJI) equation. This, in turn, is established by using a nonlinear version of Krein-Rutman theorem. We then obtain a characterization of the saddle-point equilibrium in terms of the corresponding HJI equation. Finally, we use a controlled population system to illustrate our results.
| Citation: |
| [1] |
E. Altman, Flow control using theory of zero-sum games, IEEE Trans. Automat. Control, 39 (1994), 814-816.
doi: 10.1109/9.286259.
|
| [2] |
W. J. Anderson, Continuous-Time Markov Chains. An Applications-Oriented Approach, Springer Series in Statistics, Probability and its Applications, Springer-Verlag, New York, 1991.
doi: 10.1007/978-1-4612-3038-0.
|
| [3] |
A. Arapostathis, A counterexample to a nonlinear version of the Kreǐn-Rutman theorem by R. Mahadevan, Nonlinear Anal., 171 (2018), 170-176.
doi: 10.1016/j.na.2018.02.006.
|
| [4] |
A. Arapostathis, A. Biswas and S. Pradhan, On the policy improvement algorithm for ergodic risk-sensitive control, Proceedings of the Royal Society of Edinburgh: Section A Mathematics, 152 (2020), 1305-1330.
doi: 10.1017/prm.2020.61.
|
| [5] |
A. Arapostathis, A. Biswas and S. Saha, Strict monotonicity of principal eigenvalues of elliptic operators in $\mathbb{R}^d$ and risk-sensitive control, J. Math. Pures Appl., 124 (2019), 169-219.
doi: 10.1016/j.matpur.2018.05.008.
|
| [6] |
A. Basu and M. K. Ghosh, Stochastic differential games with multiple modes and application to portfolio optimization, Stoch. Anal. Appl., 25 (2007), 845-867.
doi: 10.1080/07362990701420126.
|
| [7] |
A. Basu and M. K. Ghosh, Zero-sum risk-sensitive stochastic games on a countable state space, Stoch. Processes and Their Appl., 124 (2014), 961-983.
doi: 10.1016/j.spa.2013.09.009.
|
| [8] |
A. Basu and M. K. Ghosh, Nonzero-sum risk-sensitive stochastic games on a countable state space, Math. of Oper. Res., 43 (2018), 516-532.
doi: 10.1287/moor.2017.0870.
|
| [9] |
N. Bauerle and U. Rieder, More risk-sensitive Markov decision processes, Math. of Oper. Res., 39 (2014), 105-120.
doi: 10.1287/moor.2013.0601.
|
| [10] |
N. Bauerle and U. Rieder, Zero-sum risk-sensitive stochastic games, Stoch. Processes and Their Appl., 127 (2017), 622-642.
doi: 10.1016/j.spa.2016.06.020.
|
| [11] |
A. Biswas and S. Pradhan, Ergodic risk-sensitive control of Markov processes on countable state space revisited, ESAIM J. Control Optim Cal. Variations, 28 (2022), Paper No. 26, 50 pp.
doi: 10.1051/cocv/2022018.
|
| [12] |
A. Biswas and S. Saha, Zero-sum stochastic differential games with risk-sensitive cost, Appl. Math. Optim., 81 (2020), 113-140.
doi: 10.1007/s00245-018-9479-8.
|
| [13] |
K. Fan, Minimax theorems, Proc. Natl. Acad. Sci. U.S.A., 39 (1953), 42-47.
doi: 10.1073/pnas.39.1.42.
|
| [14] |
M. K. Ghosh and K. S. Kumar, A stochastic differential game in the orthant, J. Math. Anal. Appl., 265 (2002), 12-37.
doi: 10.1006/jmaa.2001.7679.
|
| [15] |
M. K. Ghosh, K. S. Kumar and C. Pal, Zero-sum risk-sensitive stochastic games for continuous-time Markov chains, Stoch. Anal. Appl., 34 (2016), 835-851.
doi: 10.1080/07362994.2016.1180995.
|
| [16] |
M. K. Ghosh and S. Pradhan, Zero-sum risk-sensitive stochastic differential games in the orthant, ESAIM J. Control Optim. Cal. Variations, 26 (2020), 33 pp.
doi: 10.1051/cocv/2020029.
|
| [17] |
M. K. Ghosh and S. Saha, Risk-sensitive control of continuous-time Markov chains, Stochastic, 86 (2014), 655-675.
doi: 10.1080/17442508.2013.872644.
|
| [18] |
S. Golui and C. Pal, Continuous-time zero-sum games for Markov chains with risk-sensitive finite-horizon cost criterion, Stoch. Anal. Appl., 40 (2022), 78-95.
doi: 10.1080/07362994.2021.1889381.
|
| [19] |
S. Golui, C. Pal and S. Saha, Continuous-time zero-sum games for Markov decision processes with discounted risk-sensitive cost criterion, Dyn. Games and Appl., 12 (2022), 485-512.
doi: 10.1007/s13235-021-00391-2.
|
| [20] |
X. Guo and Y. Huang, Risk-sensitive average continuous-time Markov decision processes with unbounded transition and cost rates, J. Appl. Probab., 58 (2021), 523-550.
doi: 10.1017/jpr.2020.105.
|
| [21] |
X. P. Guo and O. Hernandez-Lerma, Zero-sum games for continuous-time Markov chains with unbounded transition and average payoff rates, J. Appl. Probab., 40 (2003), 327-345.
doi: 10.1017/S0021900200019331.
|
| [22] |
X. P. Guo and O. Hernandez-Lerma, Zero-sum games for continuous-time jump Markov processes in Polish spaces: Discounted payoffs, Adv. in Appl. Probab., 39 (2007), 645-668.
doi: 10.1239/aap/1189518632.
|
| [23] |
X. P. Guo and O. Hernandez-Lerma, Continuous-Time Markov Decision Processes. Theory and Applications, Stochastic Modelling and Applied Probability, 62. Springer-Verlag, Berlin, 2009.
doi: 10.1007/978-3-642-02547-1.
|
| [24] |
X. P. Guo and Z. W. Liao, Risk-sensitive discounted continuous-time Markov decision processes with unbounded rates, SIAM J. Control Optim., 57 (2019), 3857-3883.
doi: 10.1137/18M1222016.
|
| [25] |
X. P. Guo and A. Piunovskiy, Discounted continuous-time Markov decision processes with constraints: Unbounded transition and loss rates, Math. Oper. Res., 36 (2011), 105-132.
doi: 10.1287/moor.1100.0477.
|
| [26] |
X. P. Guo and X. Song, Discounted continuous-time constrained Markov decision processes in polish spaces, Ann. Appl. Probab., 21 (2011), 2016-2049.
doi: 10.1214/10-AAP749.
|
| [27] |
O. Hernandez-Lerma and J. Lasserre, Further Topics on Discrete-Time Markov Control Processes, Springer, New York, 1999.
|
| [28] |
M. Y. Kitaev, Semi-Markov and jump Markov controlled models: Average cost criterion, SIAM Theory Probab. Appl., 30 (1995), 272-288.
|
| [29] |
M. Y. Kitaev and V. V. Rykov, Controlled Queueing Systems, CRC Press, Boca Raton, 1995.
|
| [30] |
M. B. Klompstra, Nash equilibria in risk-sensitive dynamic games, IEEE Trans. Automat. Control., 45 (2007), 1397-1401.
doi: 10.1109/9.867067.
|
| [31] |
A. Neyman, Continuous-time stochastic games, Games and Economic Behaviour, 104 (2017), 92-130.
doi: 10.1016/j.geb.2017.02.004.
|
| [32] |
A. S. Nowak, Notes on risk-sensitive Nash equilibria, Advances in Dynamic Games, Ann. Internat. Soc. Dynam. Games, Birkhauser, 7 (2005), 95-109.
doi: 10.1007/0-8176-4429-6_5.
|
| [33] |
C. Pal and S. Pradhan, Risk-sensitive control of pure jump processes on a general state space, Stochastics, 91 (2019), 155-174.
doi: 10.1080/17442508.2018.1521413.
|
| [34] |
C. Pal and S. Pradhan, Zero-sum games for pure jump processes with risk-sensitive discounted cost criteria, J. Dyn. Games, 9 (2022), 13-25.
doi: 10.3934/jdg.2021020.
|
| [35] |
A. Piunovskiy and Y. Zhang, Discounted continuous-time Markov decision processes with unbounded rates: The convex analytic approach, SIAM J. Control Optim., 49 (2011), 2032-2061.
doi: 10.1137/10081366X.
|
| [36] |
A. Piunovskiy and Y. Zhang, Continuous-Time Markov Decision Processes, Springer, 2020.
doi: 10.1007/978-3-030-54987-9.
|
| [37] |
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, John Wiley & Sons, Inc., New York, 1994.
|
| [38] |
L. S. Shapley, Stochastic games, Proc. Nat. Acad. Sci., 39 (1953), 1095-1100.
doi: 10.1073/pnas.39.10.1095.
|
| [39] |
K. Suresh Kumar and C. Pal, Risk-sensitive control of jump process on denumerable state space with near monotone cost, Appl. Math. Optim., 68 (2013), 311-331.
doi: 10.1007/s00245-013-9208-2.
|
| [40] |
K. Suresh Kumar and C. Pal, Risk-sensitive ergodic control of continuous-time Markov processes with denumerable state space, Stoch. Anal. Appl., 33 (2015), 863-881.
doi: 10.1080/07362994.2015.1050674.
|
| [41] |
Q. D. Wei, Zero-sum games for continuous-time Markov jump processes with risk-sensitive finite-horizon cost criterion, Oper. Res. Lett., 46 (2018), 69-75.
doi: 10.1016/j.orl.2017.11.008.
|
| [42] |
Q. D. Wei and X. Chen, Stochastic games for continuous-time jump processes under finite-horizon payoff criterion, Appl. Math. Optim., 74 (2016), 273-301.
doi: 10.1007/s00245-015-9314-4.
|
| [43] |
Q. D. Wei and X. Chen, Nonzero-sum games for continuous-time jump processes under the expected average payoff criterion, Appl. Math. Optim., 83 (2019), 915-938.
doi: 10.1007/s00245-019-09572-3.
|
| [44] |
Q. D. Wei and X. Chen, Nonzero-sum risk-sensitive average stochastic games: The case of unbounded costs, Dyn. Games and Appl., 11 (2021), 835-862.
doi: 10.1007/s13235-021-00380-5.
|
| [45] |
W. Z. Zhang and X. P. Guo, Nonzero-sum games for continuous-time Markov chains with unbounded transition and average payoff rates, Sci. China Math., 55 (2012), 2405-2416.
doi: 10.1007/s11425-012-4515-7.
|
| [46] |
Y. Zhang, Average optimality for continuous-time Markov decision processes under weak continuity conditions, J. of Appl. Probab., 51 (2014), 954-970.
doi: 10.1239/jap/1421763321.
|