Research article Special Issues

Adaptive tracking control with Lyapunov regularization for data-calibrated Takagi–Sugeno fuzzy Markov jump biological systems

  • Published: 16 July 2026
  • MSC : 93C40, 93E35, 92C50, 68T05

  • Cancer–tumour–immune dynamics are nonlinear, uncertain, and subject to abrupt biological or treatment-induced regime changes. This paper develops a data-calibrated adaptive tracking-control framework for Takagi–Sugeno (T–S) fuzzy Markov jump biological systems with unknown local dynamics and unknown transition probabilities. The proposed method combines mode-dependent fitted Q iteration with a Lyapunov-penalized policy update, so that the learned controller uses transition data while remaining close to a bounded stabilizing reference action. For the sampled discounted problem, we establish Bellman contraction, well-posedness of the fitted regression update, an approximate fitted Q error bound, a quantitative Lyapunov-penalization bound, a practical tracking guarantee for the Lyapunov-regularized policy, and positivity preservation under projected states and dimensionless inputs. Public tumour-volume data associated with the FIR atezolizumab non-small-cell lung cancer study are used only to calibrate plausible tumour-growth ranges; a trajectory-level holdout split is applied to avoid in-sample reporting. In the resulting active-treatment simulation regime, the proposed controller gives the best overall balanced performance among the compared strategies, which now include proximal policy optimization (PPO), soft actor–critic (SAC), nominal model predictive control (MPC), local tracking feedback, reference-action control, fixed scheduling, non-switching fitted Q iteration, and no control. In the representative single-seed run, the Lyapunov-regularized fitted Q iteration (LR-FQI) method achieves low weighted cost ($ 51.17 $), low tumour integral squared error (ISE) ($ 1.70 $), short threshold-exceedance time ($ 4 $ days), and moderate total dimensionless input. A Monte Carlo evaluation over 50 independent Markov and noise trajectories gives the lowest mean weighted cost for the proposed controller ($ 47.29\pm19.91 $, mean $ \pm $ standard deviation) and a consistent threshold-exceedance time of $ 4.00\pm0.00 $ days. Assumption 4 is verified numerically over 5000 sampled domain points.

    Citation: Shirali Kadyrov, Ardak Kashkynbayev, Yershat Sapazhanov. Adaptive tracking control with Lyapunov regularization for data-calibrated Takagi–Sugeno fuzzy Markov jump biological systems[J]. AIMS Mathematics, 2026, 11(7): 21167-21189. doi: 10.3934/math.2026860

    Related Papers:

  • Cancer–tumour–immune dynamics are nonlinear, uncertain, and subject to abrupt biological or treatment-induced regime changes. This paper develops a data-calibrated adaptive tracking-control framework for Takagi–Sugeno (T–S) fuzzy Markov jump biological systems with unknown local dynamics and unknown transition probabilities. The proposed method combines mode-dependent fitted Q iteration with a Lyapunov-penalized policy update, so that the learned controller uses transition data while remaining close to a bounded stabilizing reference action. For the sampled discounted problem, we establish Bellman contraction, well-posedness of the fitted regression update, an approximate fitted Q error bound, a quantitative Lyapunov-penalization bound, a practical tracking guarantee for the Lyapunov-regularized policy, and positivity preservation under projected states and dimensionless inputs. Public tumour-volume data associated with the FIR atezolizumab non-small-cell lung cancer study are used only to calibrate plausible tumour-growth ranges; a trajectory-level holdout split is applied to avoid in-sample reporting. In the resulting active-treatment simulation regime, the proposed controller gives the best overall balanced performance among the compared strategies, which now include proximal policy optimization (PPO), soft actor–critic (SAC), nominal model predictive control (MPC), local tracking feedback, reference-action control, fixed scheduling, non-switching fitted Q iteration, and no control. In the representative single-seed run, the Lyapunov-regularized fitted Q iteration (LR-FQI) method achieves low weighted cost ($ 51.17 $), low tumour integral squared error (ISE) ($ 1.70 $), short threshold-exceedance time ($ 4 $ days), and moderate total dimensionless input. A Monte Carlo evaluation over 50 independent Markov and noise trajectories gives the lowest mean weighted cost for the proposed controller ($ 47.29\pm19.91 $, mean $ \pm $ standard deviation) and a consistent threshold-exceedance time of $ 4.00\pm0.00 $ days. Assumption 4 is verified numerically over 5000 sampled domain points.



    加载中


    [1] V. A. Kuznetsov, I. A. Makalkin, M. A. Taylor, A. S. Perelson, Nonlinear dynamics of immunogenic tumors: parameter estimation and global bifurcation analysis, Bull. Math. Biol., 56 (1994), 295–321. https://doi.org/10.1007/BF02460644 doi: 10.1007/BF02460644
    [2] D. Kirschner, J. C. Panetta, Modeling immunotherapy of the tumor–immune interaction, J. Math. Biol., 37 (1998), 235–252. https://doi.org/10.1007/s002850050127 doi: 10.1007/s002850050127
    [3] R. S. Sutton, A. G. Barto, Reinforcement learning: an introduction, MIT Press, 1998.
    [4] S. He, M. Zhang, H. Fang, F. Liu, X. Luan, Z. Ding, Reinforcement learning and adaptive optimization of a class of Markov jump systems with completely unknown dynamic information, Neural Comput. Appl., 32 (2020), 14311–14320. https://doi.org/10.1007/s00521-019-04180-2 doi: 10.1007/s00521-019-04180-2
    [5] H. Fang, Y. Tu, H. Wang, S. He, F. Liu, Z. Ding, et al., Fuzzy-based adaptive optimization of unknown discrete-time nonlinear Markov jump systems with off-policy reinforcement learning, IEEE Trans. Fuzzy Syst., 30 (2022), 5276–5290. https://doi.org/10.1109/TFUZZ.2022.3171844 doi: 10.1109/TFUZZ.2022.3171844
    [6] X. Shi, Y. Li, C. Du, C. Chen, G. Zong, W. Gui, Reinforcement learning-based optimal control for Markov jump systems with completely unknown dynamics, Automatica, 171 (2025), 111886. https://doi.org/10.1016/j.automatica.2024.111886 doi: 10.1016/j.automatica.2024.111886
    [7] Y. Lou, M. Luo, J. Cheng, X. Wang, K. Shi, Double-quantized-based $H_{\infty}$ tracking control of T–S fuzzy semi-Markovian jump systems with adaptive event-triggered, AIMS Math., 8 (2023), 6942–6969. https://doi.org/10.3934/math.2023351 doi: 10.3934/math.2023351
    [8] G. Zong, H. Xie, D. Yang, X. Zhao, Y. Yi, Adaptive fuzzy tracking control for switched nonlinear systems under FDI attacks and input saturation: a flexible transient performance approach, IEEE Trans. Cybern., 54 (2024), 7479–7488. https://doi.org/10.1109/TCYB.2024.3463689 doi: 10.1109/TCYB.2024.3463689
    [9] Y. Wang, G. Zong, Dynamic event-triggered adaptive fixed-time practical tracking control for nonlinear systems through funnel function, IEEE Trans. Autom. Sci. Eng., 22 (2025), 7008–7017. https://doi.org/10.1109/TASE.2024.3458176 doi: 10.1109/TASE.2024.3458176
    [10] Z. Gao, Y. Wang, Z. Ji, Finite-time $H_{\infty}$ control of switched affine non-linear systems based on the Lie derivative and iteration technique, IET Control Theory Appl., 13 (2019), 3148–3154. https://doi.org/10.1049/iet-cta.2018.5902 doi: 10.1049/iet-cta.2018.5902
    [11] M. Sharifi, A. A. Jamshidi, N. N. Sarvestani, An adaptive robust control strategy in a cancer tumor–immune system under uncertainties, IEEE/ACM Trans. Comput. Biol. Bioinf., 16 (2019), 865–873. https://doi.org/10.1109/TCBB.2018.2803175 doi: 10.1109/TCBB.2018.2803175
    [12] H. Jiao, Q. Shen, Y. Shi, P. Shi, Adaptive tracking control for uncertain cancer–tumor–immune systems, IEEE/ACM Trans. Comput. Biol. Bioinf., 18 (2021), 2753–2758. https://doi.org/10.1109/TCBB.2020.3036069 doi: 10.1109/TCBB.2020.3036069
    [13] E. Ahmadi, J. Zarei, R. Razavi-Far, M. Saif, A dual approach for positive T–S fuzzy controller design and its application to cancer treatment under immunotherapy and chemotherapy, Biomed. Signal Process. Control, 58 (2020), 101822. https://doi.org/10.1016/j.bspc.2019.101822 doi: 10.1016/j.bspc.2019.101822
    [14] A. Kashkynbayev, R. Rakkiyappan, Sampled-data output tracking control based on T–S fuzzy model for cancer–tumor–immune systems, Commun. Nonlinear Sci. Numer. Simul., 128 (2024), 107642. https://doi.org/10.1016/j.cnsns.2023.107642 doi: 10.1016/j.cnsns.2023.107642
    [15] T. Takagi, M. Sugeno, Fuzzy identification of systems and its applications to modeling and control, IEEE Trans. Syst., Man, Cybern., SMC-15 (1985), 116–132. https://doi.org/10.1109/TSMC.1985.6313399 doi: 10.1109/TSMC.1985.6313399
    [16] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, arXiv Preprint, 2017. https://doi.org/10.48550/arXiv.1707.06347
    [17] T. Haarnoja, A. Zhou, P. Abbeel, S. Levine, Soft actor–critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor, Proceedings of the 35th International Conference on Machine Learning, 80 (2018), 1861–1870.
    [18] D. Ernst, P. Geurts, L. Wehenkel, Tree-based batch mode reinforcement learning, J. Mach. Learn. Res., 6 (2005), 503–556.
    [19] X. Liu, Q. Li, J. Pan, A deterministic and stochastic model for the system dynamics of tumor–immune responses to chemotherapy, Phys. A: Stat. Mech. Appl., 500 (2018), 162–176. https://doi.org/10.1016/j.physa.2018.02.118 doi: 10.1016/j.physa.2018.02.118
    [20] M. Dassow, S. Djouadi, K. Moussa, Optimal control of a tumor–immune system with a modified Stepanova cancer model, IFAC-PapersOnLine 54 (2021), 227–232. https://doi.org/10.1016/j.ifacol.2021.10.260 doi: 10.1016/j.ifacol.2021.10.260
    [21] M. Farman, A. Ahmad, A. Akgül, M. U. Saleem, K. S. Nisar, V. Vijayakumar, Dynamical behavior of tumor–immune system with fractal-fractional operator, AIMS Math., 7 (2022), 8751–8773. https://doi.org/10.3934/math.2022489 doi: 10.3934/math.2022489
    [22] N. Ghaffari Laleh, C. M. L. Loeffler, J. Grajek, K. Staňková, A. T. Pearson, H. S. Muti, et al., Classical mathematical models for prediction of response to chemotherapy and immunotherapy, PLOS Comput. Biol., 18 (2022), e1009822. https://doi.org/10.1371/journal.pcbi.1009822 doi: 10.1371/journal.pcbi.1009822
    [23] L. A. Zadeh, Fuzzy sets, Inf. Control, 8 (1965), 338–353. https://doi.org/10.1016/S0019-9958(65)90241-X doi: 10.1016/S0019-9958(65)90241-X
    [24] O. L. V. Costa, R. P. Marques, M. D. Fragoso, Discrete-time Markov jump linear systems, London: Springer, 2005. https://doi.org/10.1007/b138575
    [25] O. L. V. Costa, M. D. Fragoso, M. G. Todorov, Continuous-time Markov jump linear systems, Springer Science & Business Media, 2013. https://doi.org/10.1007/978-3-642-34100-7
    [26] R. Bellman, Dynamic programming, Princeton: Princeton University Press, 1957.
    [27] D. P. Bertsekas, Neuro-dynamic programming, In: P. M. Pardalos, O. A. Prokopyev, Encyclopedia of optimization, Springer, 2025. https://doi.org/10.1007/978-3-030-54621-2_440-1
    [28] P. J. Werbos, Approximate dynamic programming for real-time control and neural modeling, In: D. A. White, D. A. Sofge, Handbook of intelligent control: neural, fuzzy, and adaptive approaches, Van Nostrand Reinhold, 1992.
    [29] M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming, New York: John Wiley & Sons, 2014.
    [30] R. Munos, C. Szepesvári, Finite-time bounds for fitted value iteration, J. Mach. Learn. Res., 9 (2008), 815–857.
    [31] Z. Artstein, Stabilization with relaxed controls, Nonlinear Anal.: Theory, Methods Appl., 7 (1983), 1163–1173. https://doi.org/10.1016/0362-546X(83)90049-4 doi: 10.1016/0362-546X(83)90049-4
    [32] E. D. Sontag, A universal construction of Artstein's theorem on nonlinear stabilization, Syst. Control Lett., 13 (1989), 117–123. https://doi.org/10.1016/0167-6911(89)90028-5 doi: 10.1016/0167-6911(89)90028-5
    [33] R. A. Freeman, P. V. Kokotović, Robust nonlinear control design: state-space and Lyapunov techniques, Springer Science & Business Media, 2008.
    [34] M. Nagumo, Über die lage der integralkurven gewöhnlicher differentialgleichungen, Proceedings of the Physico-Mathematical Society of Japan, 3rd Series, 24 (1942), 551–559. https://doi.org/10.11429/ppmsj1919.24.0_551 doi: 10.11429/ppmsj1919.24.0_551
    [35] W. M. Haddad, V. Chellaboina, Q. Hui, Nonnegative and compartmental dynamical systems, Princeton: Princeton University Press, 2010. https://doi.org/10.1515/9781400832248
    [36] A. L. MacLean, E. T. Roussos Torres, J. Kreger, ModelingMDSCs: tumour volume data from the FIR NSCLC atezolizumab study, GitHub repository, 2023. Available from: https://github.com/maclean-lab/ModelingMDSCs.
    [37] D. R. Spigel, J. E. Chaft, S. Gettinger, B. H. Chao, L. Dirix, P. Schmid, et al., FIR: efficacy, safety, and biomarker analysis of a phase Ⅱ open-label study of atezolizumab in PD-L1-selected patients with NSCLC, J. Thorac. Oncol., 13 (2018), 1733–1742. https://doi.org/10.1016/j.jtho.2018.05.004 doi: 10.1016/j.jtho.2018.05.004
    [38] J. Kreger, E. T. Roussos Torres, A. L. MacLean, Myeloid-derived suppressor-cell dynamics control outcomes in the metastatic niche, Cancer Immunol. Res., 11 (2023), 614–628. https://doi.org/10.1158/2326-6066.CIR-22-0617 doi: 10.1158/2326-6066.CIR-22-0617
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(107) PDF downloads(20) Cited by(0)

Article outline

Figures and Tables

Figures(5)  /  Tables(6)

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog