Research article Special Issues

Actor–critic–identifier reinforcement learning with gradient-based integral concurrent learning for continuous-time nonlinear nonzero-sum differential graphical games

  • Published: 30 June 2026
  • In this paper, we investigate multiplayer nonzero-sum differential graphical games for continuous-time nonlinear systems with unknown dynamics. Conventional reinforcement-learning-based optimal control methods for such problems usually suffer from the curse of dimensionality in parameter estimation and high online computational cost. To address these issues, an indirect adaptive control algorithm with distributed sparse identification is developed within an actor–critic–identifier architecture. Specifically, graphical game structural priors are incorporated to reduce the dimension of the unknown parameter space, and a gradient-based integral concurrent learning identifier using a historical data stack is designed for online system identification. In contrast to conventional least-squares-based concurrent learning methods, the proposed identifier avoids covariance matrix inversion and reduces the per-step computational complexity from $ \mathcal{O}(q^2) $ to $ \mathcal{O}(q) $. By combining the identifier with integral reinforcement learning, the proposed method achieves online approximation of the Nash equilibrium solution to the coupled Hamilton–Jacobi–Bellman equations without requiring an exact system model. A Lyapunov-based analysis is further conducted to establish exponential convergence of the identifier parameters under finite excitation and uniformly ultimately bounded stability of the closed-loop system under a persistence-of-excitation condition on the critic regressor. Comparative simulation results for the considered benchmark example show that, while maintaining a weight approximation error on the order of $ 10^{-3} $, the proposed algorithm reduces the convergence time by 75%, thereby improving learning efficiency and alleviating the online computational burden.

    Citation: Lei Guo, Fei Li, Yuan Song. Actor–critic–identifier reinforcement learning with gradient-based integral concurrent learning for continuous-time nonlinear nonzero-sum differential graphical games[J]. Electronic Research Archive, 2026, 34(8): 5496-5524. doi: 10.3934/era.2026246

    Related Papers:

  • In this paper, we investigate multiplayer nonzero-sum differential graphical games for continuous-time nonlinear systems with unknown dynamics. Conventional reinforcement-learning-based optimal control methods for such problems usually suffer from the curse of dimensionality in parameter estimation and high online computational cost. To address these issues, an indirect adaptive control algorithm with distributed sparse identification is developed within an actor–critic–identifier architecture. Specifically, graphical game structural priors are incorporated to reduce the dimension of the unknown parameter space, and a gradient-based integral concurrent learning identifier using a historical data stack is designed for online system identification. In contrast to conventional least-squares-based concurrent learning methods, the proposed identifier avoids covariance matrix inversion and reduces the per-step computational complexity from $ \mathcal{O}(q^2) $ to $ \mathcal{O}(q) $. By combining the identifier with integral reinforcement learning, the proposed method achieves online approximation of the Nash equilibrium solution to the coupled Hamilton–Jacobi–Bellman equations without requiring an exact system model. A Lyapunov-based analysis is further conducted to establish exponential convergence of the identifier parameters under finite excitation and uniformly ultimately bounded stability of the closed-loop system under a persistence-of-excitation condition on the critic regressor. Comparative simulation results for the considered benchmark example show that, while maintaining a weight approximation error on the order of $ 10^{-3} $, the proposed algorithm reduces the convergence time by 75%, thereby improving learning efficiency and alleviating the online computational burden.



    加载中


    [1] Y. Huo, D. Wang, J. Qiao, M. Li, Off-policy model-free learning for multi-player non-zero-sum games with constrained inputs, IEEE Trans. Circuits Syst. I Reg. Pap., 70 (2023), 910–920. https://doi.org/10.1109/TCSI.2022.3221274 doi: 10.1109/TCSI.2022.3221274
    [2] Y. Zhang, B. Zhao, D. Liu, S. Zhang, Adaptive dynamic programming-based event-triggered robust control for multiplayer nonzero-sum games with unknown dynamics, IEEE Trans. Cybern., 53 (2023), 5151–5164. https://doi.org/10.1109/TCYB.2022.3175650 doi: 10.1109/TCYB.2022.3175650
    [3] Q. Wei, L. Zhu, R. Song, P. Zhang, D. Liu, J. Xiao, Model-free adaptive optimal control for unknown nonlinear multiplayer nonzero-sum game, IEEE Trans. Neural Netw. Learn. Syst., 33 (2022), 879–892. https://doi.org/10.1109/TNNLS.2020.3030127 doi: 10.1109/TNNLS.2020.3030127
    [4] D. Liu, S. Xue, B. Zhao, B. Luo, Q. Wei, Adaptive dynamic programming for control: A survey and recent advances, IEEE Trans. Syst. Man Cybern. Syst., 51 (2021), 142–160. https://doi.org/10.1109/TSMC.2020.3042876 doi: 10.1109/TSMC.2020.3042876
    [5] R. S. Sutton, A. G. Barto, Reinforcement Learning: An Introduction, MIT Press, Cambridge, MA, USA, 2018.
    [6] Z. Pan, D. Lei, L. Wang, A knowledge-based two-population optimization algorithm for distributed energy-efficient parallel machines scheduling, IEEE Trans. Cybern., 52 (2022), 5051–5063. https://doi.org/10.1109/TCYB.2020.3026571 doi: 10.1109/TCYB.2020.3026571
    [7] H. Wang, B. R. Sarker, J. Li, J. Li, Adaptive scheduling for assembly job shop with uncertain assembly times based on dual Q-learning, Int. J. Prod. Res., 59 (2021), 5867–5883. https://doi.org/10.1080/00207543.2020.1794075 doi: 10.1080/00207543.2020.1794075
    [8] F. Zhao, Z. Fu, L. Wang, H. Sang, A heterogeneous graph reinforcement learning framework with question-aware neighborhood aggregation and interoption prompt attention for dynamic flexible job shop scheduling problem, IEEE Trans. Ind. Inform., 22 (2026), 2863–2874. https://doi.org/10.1109/TII.2025.3646962 doi: 10.1109/TII.2025.3646962
    [9] D. Liu, H. Li, D. Wang, Online synchronous approximate optimal learning algorithm for multi-player non-zero-sum games with unknown dynamics, IEEE Trans. Syst. Man Cybern. Syst., 44 (2014), 1015–1027. https://doi.org/10.1109/TSMC.2013.2295351 doi: 10.1109/TSMC.2013.2295351
    [10] K. G. Vamvoudakis, F. L. Lewis, S. S. Ge, Neural networks in feedback control systems, in Mechanical Engineers' Handbook, Volume 2: Design, Instrumentation, and Controls, Wiley, (2014), 843–894. https://doi.org/10.1002/9781118985960.meh223
    [11] J. Li, J. Ding, T. Chai, F. L. Lewis, Nonzero-sum game reinforcement learning for performance optimization in large-scale industrial processes, IEEE Trans. Cybern., 50 (2020), 4132–4145. https://doi.org/10.1109/TCYB.2019.2950262 doi: 10.1109/TCYB.2019.2950262
    [12] Q. Zhao, J. Sun, G. Wang, J. Chen, Event-triggered ADP for nonzero-sum games of unknown nonlinear systems, IEEE Trans. Neural Netw. Learn. Syst., 33 (2022), 1905–1913. https://doi.org/10.1109/TNNLS.2021.3071545 doi: 10.1109/TNNLS.2021.3071545
    [13] A. Odekunle, W. Gao, M. Davari, Z. P. Jiang, Reinforcement learning and non-zero-sum game output regulation for multi-player linear uncertain systems, Automatica, 112 (2020), 108672. https://doi.org/10.1016/j.automatica.2019.108672 doi: 10.1016/j.automatica.2019.108672
    [14] J. Sun, H. Zhang, Y. Yan, S. Xu, X. Fan, Optimal regulation strategy for nonzero-sum games of the immune system using adaptive dynamic programming, IEEE Trans. Cybern., 53 (2023), 1475–1484. https://doi.org/10.1109/TCYB.2021.3103820 doi: 10.1109/TCYB.2021.3103820
    [15] L. Guo, H. Zhao, Model-free adaptive optimal control of continuous-time nonlinear non-zero-sum games based on reinforcement learning, IET Control Theory Appl., 17 (2023), 223–239. https://doi.org/10.1049/cth2.12376 doi: 10.1049/cth2.12376
    [16] M. Johnson, R. Kamalapurkar, S. Bhasin, W. E. Dixon, Approximate $N$-player nonzero-sum game solution for an uncertain continuous nonlinear system, IEEE Trans. Neural Netw. Learn. Syst., 26 (2015), 1645–1658. https://doi.org/10.1109/TNNLS.2014.2350835 doi: 10.1109/TNNLS.2014.2350835
    [17] K. G. Vamvoudakis, F. L. Lewis, Multi-player non-zero-sum games: Online adaptive learning solution of coupled Hamilton–Jacobi equations, Automatica, 47 (2011), 1556–1569. https://doi.org/10.1016/j.automatica.2011.03.005 doi: 10.1016/j.automatica.2011.03.005
    [18] T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, Model-based reinforcement learning: A survey, Found. Trends Mach. Learn., 16 (2023), 1–118. https://doi.org/10.1561/2200000086 doi: 10.1561/2200000086
    [19] D. Vrabie, F. L. Lewis, Neural network approach to continuous-time direct adaptive optimal control for partially unknown nonlinear systems, Neural Netw., 22 (2009), 237–246. https://doi.org/10.1016/j.neunet.2009.03.008 doi: 10.1016/j.neunet.2009.03.008
    [20] J. Y. Lee, J. B. Park, Y. H. Choi, Integral reinforcement learning for continuous-time input-affine nonlinear systems with simultaneous invariant explorations, IEEE Trans. Neural Netw. Learn. Syst., 26 (2015), 916–932. https://doi.org/10.1109/TNNLS.2014.2328590 doi: 10.1109/TNNLS.2014.2328590
    [21] L. Guo, H. Zhao, Online adaptive optimal control algorithm based on synchronous integral reinforcement learning with explorations, Neurocomputing, 520 (2023), 250–261. https://doi.org/10.1016/j.neucom.2022.11.055 doi: 10.1016/j.neucom.2022.11.055
    [22] L. Guo, W. Xiong, Y. Song, D. Gan, An efficient model-free adaptive optimal control of continuous-time nonlinear non-zero-sum games based on integral reinforcement learning with exploration, IET Control Theory Appl., 18 (2024), 748–763. https://doi.org/10.1049/cth2.12610 doi: 10.1049/cth2.12610
    [23] S. Bhasin, R. Kamalapurkar, M. Johnson, K. G. Vamvoudakis, F. L. Lewis, W. E. Dixon, A novel actor-critic-identifier architecture for approximate optimal control of uncertain nonlinear systems, Automatica, 49 (2013), 82–92. https://doi.org/10.1016/j.automatica.2012.09.019 doi: 10.1016/j.automatica.2012.09.019
    [24] G. Chowdhary, E. Johnson, Concurrent learning for convergence in adaptive control without persistency of excitation, in 49th IEEE Conference on Decision and Control, IEEE, (2010), 3674–3679. https://doi.org/10.1109/CDC.2010.5717148
    [25] R. Kamalapurkar, J. R. Klotz, W. E. Dixon, Concurrent learning-based approximate feedback-Nash equilibrium solution of $N$-player nonzero-sum differential games, IEEE/CAA J. Autom. Sin., 1 (2014), 239–247. https://doi.org/10.1109/JAS.2014.7004681 doi: 10.1109/JAS.2014.7004681
    [26] F. Tatari, K. G. Vamvoudakis, M. Mazouchi, Optimal distributed learning for disturbance rejection in networked non-linear games under unknown dynamics, IET Control Theory Appl., 13 (2019), 2838–2848. https://doi.org/10.1049/iet-cta.2018.5832 doi: 10.1049/iet-cta.2018.5832
    [27] A. Parikh, R. Kamalapurkar, W. E. Dixon, Integral concurrent learning: Adaptive control with parameter convergence using finite excitation, Int. J. Adapt. Control Signal Process., 33 (2019), 1775–1787. https://doi.org/10.1002/acs.2945 doi: 10.1002/acs.2945
    [28] D. M. Le, O. S. Patil, P. M. Amy, W. E. Dixon, Integral concurrent learning-based accelerated gradient adaptive control of uncertain Euler–Lagrange systems, in 2022 American Control Conference, IEEE, (2022), 806–811. https://doi.org/10.23919/ACC53348.2022.9867710
    [29] W. Si, S. Gao, M. Zhang, T. Wen, Y. Bai, H. Wang, Approximate optimal control for uncertain nonlinear systems: A reinforcement relearning framework, Nonlinear Dyn., 113 (2025), 24821–24848. https://doi.org/10.1007/s11071-025-11393-9 doi: 10.1007/s11071-025-11393-9
    [30] K. G. Vamvoudakis, F. L. Lewis, G. R. Hudas, Multi-agent differential graphical games: Online adaptive learning solution for synchronization with optimality, Automatica, 48 (2012), 1598–1611. https://doi.org/10.1016/j.automatica.2012.05.074 doi: 10.1016/j.automatica.2012.05.074
    [31] V. G. Lopez, F. L. Lewis, Y. Wan, M. Liu, G. A. Hewer, K. Estabridis, Stability and robustness analysis of minmax solutions for differential graphical games, Automatica, 121 (2020), 109177. https://doi.org/10.1016/j.automatica.2020.109177 doi: 10.1016/j.automatica.2020.109177
    [32] M. Liu, Y. Wan, V. G. Lopez, F. L. Lewis, G. A. Hewer, K. Estabridis, Differential graphical game with distributed global Nash solution, IEEE Trans. Control Netw. Syst., 8 (2021), 1371–1382. https://doi.org/10.1109/TCNS.2021.3065654 doi: 10.1109/TCNS.2021.3065654
    [33] V. G. Lopez, F. L. Lewis, M. Liu, Y. Wan, Beyond Nash solutions for differential graphical games, IEEE Trans. Autom. Control, 68 (2023), 5791–5797. https://doi.org/10.1109/TAC.2022.3230734 doi: 10.1109/TAC.2022.3230734
    [34] G. Huang, Z. Zhang, W. Yan, X. Guo, Differential graphical games of multiagent systems with nonzero leader's control input and external disturbances, Int. J. Robust Nonlinear Control, 34 (2024), 8144–8162. https://doi.org/10.1002/rnc.7378 doi: 10.1002/rnc.7378
    [35] B. Lian, W. Xue, F. L. Lewis, A. Davoudi, Nash-minmax strategy for multiplayer multiagent graphical games with reinforcement learning, IEEE Trans. Control Netw. Syst., 12 (2025), 763–775. https://doi.org/10.1109/TCNS.2024.3419823 doi: 10.1109/TCNS.2024.3419823
    [36] Q. Liu, H. Yan, K. Chen, M. Wang, Z. Li, Distributed Nash equilibrium solution for multi-agent game in adversarial environment: A reinforcement learning method, Automatica, 178 (2025), 112342. https://doi.org/10.1016/j.automatica.2025.112342 doi: 10.1016/j.automatica.2025.112342
    [37] Z. Ming, H. Zhang, J. Zhang, X. Xie, A novel actor-critic-identifier architecture for nonlinear multiagent systems with gradient descent method, Automatica, 155 (2023), 111128. https://doi.org/10.1016/j.automatica.2023.111128 doi: 10.1016/j.automatica.2023.111128
    [38] K. S. Narendra, A. M. Annaswamy, Stable Adaptive Systems, Prentice Hall, Englewood Cliffs, NJ, USA, 1989.
    [39] J. J. E. Slotine, W. Li, Applied Nonlinear Control, Prentice Hall, Englewood Cliffs, NJ, USA, 1991.
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(285) PDF downloads(18) Cited by(0)

Article outline

Figures and Tables

Figures(5)  /  Tables(1)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog