Research article

Double-inertial bilevel optimization for hybrid deep learning: Accelerating chest X-ray classification

  • Published: 13 July 2026
  • MSC : 47H09, 47H10, 47J25, 65K10, 90C25, 68T07

  • Convex bilevel optimization, in which an upper-level problem is optimized over the solution set of a lower-level convex program, naturally arises in machine learning tasks such as hyperparameter tuning, meta-learning, and regularized classification. Existing first-order bilevel solvers typically employ single-inertial acceleration, which can exhibit oscillatory behaviors and premature stagnation. In this paper, we propose the Double-Inertial proximal Gradient Method with Viscosity Approximation (DIPGM-VA), a novel algorithm that combines double-inertial extrapolation with viscosity approximation within a forward–backward splitting framework. We establish the strong convergence of the iterates to the unique solution of the bilevel problem in real Hilbert spaces under standard assumptions. As an application, we develop a hybrid deep learning architecture that integrates EfficientNetV2-B0 with an Extreme Learning Machine (ELM) classifier whose output weights are determined by solving a bilevel optimization problem via DIPGM-VA, and validate it on a pediatric chest X-ray pneumonia benchmark, addressing a leading cause of death among children under five in developing countries. Numerical experiments show that DIPGM-VA outperforms other state-of-the-art bilevel solvers by reaching a lower inner-level objective value within the same iteration budget. Concurrently, the hybrid model delivers a competitive F1-score of 0.9001 and accelerates the training process by an order of magnitude relative to standard end-to-end approaches.

    Citation: Suthep Suantai, Kobkoon Janngam. Double-inertial bilevel optimization for hybrid deep learning: Accelerating chest X-ray classification[J]. AIMS Mathematics, 2026, 11(7): 20502-20534. doi: 10.3934/math.2026834

    Related Papers:

  • Convex bilevel optimization, in which an upper-level problem is optimized over the solution set of a lower-level convex program, naturally arises in machine learning tasks such as hyperparameter tuning, meta-learning, and regularized classification. Existing first-order bilevel solvers typically employ single-inertial acceleration, which can exhibit oscillatory behaviors and premature stagnation. In this paper, we propose the Double-Inertial proximal Gradient Method with Viscosity Approximation (DIPGM-VA), a novel algorithm that combines double-inertial extrapolation with viscosity approximation within a forward–backward splitting framework. We establish the strong convergence of the iterates to the unique solution of the bilevel problem in real Hilbert spaces under standard assumptions. As an application, we develop a hybrid deep learning architecture that integrates EfficientNetV2-B0 with an Extreme Learning Machine (ELM) classifier whose output weights are determined by solving a bilevel optimization problem via DIPGM-VA, and validate it on a pediatric chest X-ray pneumonia benchmark, addressing a leading cause of death among children under five in developing countries. Numerical experiments show that DIPGM-VA outperforms other state-of-the-art bilevel solvers by reaching a lower inner-level objective value within the same iteration budget. Concurrently, the hybrid model delivers a competitive F1-score of 0.9001 and accelerates the training process by an order of magnitude relative to standard end-to-end approaches.



    加载中


    [1] P. Rajpurkar, J. Irvin, K. Zhu, B. Yang, H. Mehta, T. Duan, et al., CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning, arXiv Preprint, 2017. https://doi.org/10.48550/arXiv.1711.05225
    [2] A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, et al., Dermatologist-level classification of skin cancer with deep neural networks, Nature, 542 (2017), 115–118. https://doi.org/10.1038/nature21056 doi: 10.1038/nature21056
    [3] L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, M. Pontil, Bilevel programming for hyperparameter optimization and meta-learning, arXiv Preprint, 2018. https://doi.org/10.48550/arXiv.1806.04910
    [4] A. Shaban, C. A. Cheng, N. Hatch, B. Boots, Truncated back-propagation for bilevel optimization, In: Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS) 2019, Naha, Okinawa, Japan, 89 (2019), 1723–1732. Available from: https://proceedings.mlr.press/v89/shaban19a.html.
    [5] K. Janngam, S. Suantai, Y. J. Cho, A. Kaewkhao, R. Wattanataweekul, A novel inertial viscosity algorithm for bilevel optimization problems applied to classification problems, Mathematics, 11 (2023), 3241. https://doi.org/10.3390/math11143241 doi: 10.3390/math11143241
    [6] S. Sabach, S. Shtern, A first order method for solving convex bilevel optimization problems, SIAM J. Optimiz., 27 (2017), 640–660. https://doi.org/10.1137/16M105592X doi: 10.1137/16M105592X
    [7] J. J. Moreau, Fonctions convexes duales et points proximaux dans un espace hilbertien, C. R. Acad. Sci. Paris, 255 (1962), 2897–2899.
    [8] P. L. Combettes, V. R. Wajs, Signal recovery by proximal forward-backward splitting, Multiscale Model. Sim., 4 (2005), 1168–1200. https://doi.org/10.1137/050626090 doi: 10.1137/050626090
    [9] A. Moudafi, Viscosity approximation methods for fixed-points problems, J. Math. Anal. Appl., 241 (2000), 46–55. https://doi.org/10.1006/jmaa.1999.6615 doi: 10.1006/jmaa.1999.6615
    [10] B. T. Polyak, Some methods of speeding up the convergence of iteration methods, USSR Comp. Math. Math. Phys., 4 (1964), 1–17. https://doi.org/10.1016/0041-5553(64)90137-5 doi: 10.1016/0041-5553(64)90137-5
    [11] Y. Nesterov, A method for solving the convex programming problem with convergence rate $O(1/k^{2})$, Dokl. Akad. Nauk Sssr, 269 (1983), 543–547.
    [12] A. Beck, M. Teboulle, A fast iterative shrinkage-thresholding algorithm for linear inverse problems, SIAM J. Imaging Sci., 2 (2009), 183–202. https://doi.org/10.1137/080716542 doi: 10.1137/080716542
    [13] L. Bussaban, A. Kaewkhao, S. Suantai, Inertial S-iteration forward-backward algorithm for a family of nonexpansive operators with applications to image restoration problems, Filomat, 35 (2021), 771–782. https://doi.org/10.2298/FIL2103771B doi: 10.2298/FIL2103771B
    [14] P. Yatakoat, S. Suantai, A. Hanjing, On some accelerated optimization algorithms based on fixed point and linesearch techniques for convex minimization problems with applications, Adv. Contin. Discret. M., 2022 (2022), 25. https://doi.org/10.1186/s13662-022-03698-5 doi: 10.1186/s13662-022-03698-5
    [15] P. Sae-jia, S. Suantai, A new two-step inertial algorithm for solving convex bilevel optimization problems with application in data classification problems, AIMS Math., 9 (2024), 8476–8496. https://doi.org/10.3934/math.2024412 doi: 10.3934/math.2024412
    [16] R. Wattanataweekul, K. Janngam, S. Suantai, A novel two-step inertial viscosity algorithm for bilevel optimization problems applied to image recovery, Mathematics, 11 (2023), 3518. https://doi.org/10.3390/math11163518 doi: 10.3390/math11163518
    [17] K. Janngam, S. Suantai, R. Wattanataweekul, A novel fixed-point based two-step inertial algorithm for convex minimization in deep learning data classification, AIMS Math., 10 (2025), 6209–6232. https://doi.org/10.3934/math.2025283 doi: 10.3934/math.2025283
    [18] Z. Y. Peng, D. Li, Y. Zhao, R. L. Liang, An accelerated subgradient extragradient algorithm for solving bilevel variational inequality problems involving non-Lipschitz operator, Commun. Nonlinear Sci., 127 (2023), 107549. https://doi.org/10.1016/j.cnsns.2023.107549 doi: 10.1016/j.cnsns.2023.107549
    [19] Y. Shehu, P. T. Vuong, A. Zemkoho, An inertial extrapolation method for convex simple bilevel optimization, Optim. Method. Softw., 36 (2021), 1–19. https://doi.org/10.1080/10556788.2019.1619729 doi: 10.1080/10556788.2019.1619729
    [20] P. Duan, Y. Zhang, Alternated and multi-step inertial approximation methods for solving convex bilevel optimization problems, Optimization, 72 (2023), 2517–2545. https://doi.org/10.1080/02331934.2022.2069022 doi: 10.1080/02331934.2022.2069022
    [21] C. Izuchukwu, M. Aphane, K. O. Aremu, Two-step inertial forward–reflected–anchored–backward splitting algorithm for solving monotone inclusion problems, Comput. Appl. Math., 42 (2023), 351. https://doi.org/10.1007/s40314-023-02485-6 doi: 10.1007/s40314-023-02485-6
    [22] O. S. Iyiola, Y. Shehu, Convergence results of two-step inertial proximal point algorithm, Appl. Numer. Math., 182 (2022), 57–75. https://doi.org/10.1016/j.apnum.2022.07.013 doi: 10.1016/j.apnum.2022.07.013
    [23] D. V. Thong, S. Reich, X. H. Li, P. T. H. Tham, An efficient algorithm with double inertial steps for solving split common fixed point problems and an application to signal processing, Comput. Appl. Math., 44 (2025), 102. https://doi.org/10.1007/s40314-024-03058-x doi: 10.1007/s40314-024-03058-x
    [24] R. Wattanataweekul, K. Janngam, S. Suantai, A novel double inertial viscosity algorithm for convex bilevel optimization problems applied to image restoration problems, Optimization, 74 (2025), 2635–2656. https://doi.org/10.1080/02331934.2024.2398776 doi: 10.1080/02331934.2024.2398776
    [25] S. Suantai, K. Janngam, A novel fixed-point based two-step inertial algorithm for convex bilevel optimization in deep learning data classification, Carpathian J. Math., 42 (2026), 369–392. https://doi.org/10.37193/CJM.2026.02.10 doi: 10.37193/CJM.2026.02.10
    [26] M. Tan, Q. V. Le, EfficientNetV2: Smaller models and faster training, arXiv Preprint, 2021. https://doi.org/10.48550/arXiv.2104.00298
    [27] D. S. Kermany, M. Goldbaum, W. Cai, C. C. S. Valentim, H. Liang, S. L. Baxter, et al., Identifying medical diagnoses and treatable diseases by image-based deep learning, Cell, 172 (2018), 1122–1131. https://doi.org/10.1016/j.cell.2018.02.010 doi: 10.1016/j.cell.2018.02.010
    [28] G. B. Huang, Q. Y. Zhu, C. K. Siew, Extreme learning machine: Theory and applications, Neurocomputing, 70 (2006), 489–501. https://doi.org/10.1016/j.neucom.2005.12.126 doi: 10.1016/j.neucom.2005.12.126
    [29] H. H. Bauschke, P. L. Combettes, Convex analysis and monotone operator theory in Hilbert spaces, 2 Eds., Springer Cham, 2017. https://doi.org/10.1007/978-3-319-48311-5
    [30] K. Goebel, S. Reich, Uniform convexity, hyperbolic geometry, and nonexpansive mappings, New York and Basel: Marcel Dekker, 1984.
    [31] J. B. Baillon, R. E. Bruck, S. Reich, On the asymptotic behavior of nonexpansive mappings and semigroups in Banach spaces, Houston J. Math., 4 (1978), 1–9.
    [32] W. Takahashi, Introduction to nonlinear and Cconvex analysis, Yokohama: Yokohama Publ., 2009.
    [33] K. Nakajo, K. Shimoji, W. Takahashi, Strong convergence theorems by the hybrid method for families of nonexpansive mappings in Hilbert spaces, Taiwan. J. Math., 10 (2006), 339–360. https://doi.org/10.11650/twjm/1500403829 doi: 10.11650/twjm/1500403829
    [34] A. Hanjing, P. Thongpaen, S. Suantai, A new accelerated algorithm with a linesearch technique for convex bilevel optimization problems with applications, AIMS Math., 9 (2024), 22366–22392. https://doi.org/10.3934/math.20241088 doi: 10.3934/math.20241088
    [35] K. Aoyama, W. Takahashi, Strong convergence theorems for a class of nonexpansive mappings in Banach spaces, J. Nonlinear Convex A., 12 (2011), 121–134.
    [36] K. Aoyama, F. Kohsaka, W. Takahashi, Strongly relatively nonexpansive sequences in Banach spaces and applications, J. Fix. Point Theory A., 5 (2009), 201–225. https://doi.org/10.1007/s11784-009-0108-7 doi: 10.1007/s11784-009-0108-7
    [37] S. Saejung, P. Yotkaew, Approximation of zeros of inverse strongly monotone operators in Banach spaces, Nonlinear Anal.-Theor., 75 (2012), 742–750. https://doi.org/10.1016/j.na.2011.09.005 doi: 10.1016/j.na.2011.09.005
    [38] L. Bussaban, A. Kaewkhao, S. Suantai, A parallel inertial S-iteration forward-backward algorithm for regression and classification problems, Carpathian J. Math., 36 (2020), 35–44. https://doi.org/10.37193/cjm.2020.01.04 doi: 10.37193/cjm.2020.01.04
    [39] Y. L. Cun, Y. Bengio, G. Hinton, Deep learning, Nature, 521 (2015), 436–444. https://doi.org/10.1038/nature14539
    [40] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016,770–778. https://doi.org/10.1109/CVPR.2016.90
    [41] S. J. Pan, Q. Yang, A survey on transfer learning, IEEE T. Knowl. Data En., 22 (2010), 1345–1359. https://doi.org/10.1109/TKDE.2009.191 doi: 10.1109/TKDE.2009.191
    [42] J. Yosinski, J. Clune, Y. Bengio, H. Lipson, How transferable are features in deep neural networks? Adv. Neural Inf. Process. Syst., 27 (2014).
    [43] G. B. Huang, H. Zhou, X. Ding, R. Zhang, Extreme learning machine for regression and multiclass classification, IEEE T. Syst. Man Cy. B, 42 (2012), 513–529. https://doi.org/10.1109/TSMCB.2011.2168604 doi: 10.1109/TSMCB.2011.2168604
    [44] W. Deng, Q. Zheng, L. Chen, Regularized extreme learning machine, In: 2009 IEEE Symposium on Computational Intelligence and Data Mining, Nashville, TN, USA, 2009,389–395. https://doi.org/10.1109/CIDM.2009.4938676
    [45] R. Tibshirani, Regression shrinkage and selection via the lasso, J. R. Stat. Soc. B, 58 (1996), 267–288. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x doi: 10.1111/j.2517-6161.1996.tb02080.x
    [46] M. Tan, Q. V. Le, EfficientNet: Rethinking model scaling for convolutional neural networks, In: Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, 97 (2019), 6105–6114. Available from: https://proceedings.mlr.press/v97/tan19a.html.
    [47] S. M. Pizer, E. P. Amburn, J. D. Austin, R. Cromartie, A. Geselowitz, T. Greer, et al., Adaptive histogram equalization and its variations, Comput. Vis. Graph. Image Process., 39 (1987), 355–368. https://doi.org/10.1016/S0734-189X(87)80186-X doi: 10.1016/S0734-189X(87)80186-X
    [48] A. M. Reza, Realization of the contrast limited adaptive histogram equalization (CLAHE) for real-time image enhancement, J. Vlsi Sig. Proc. Syst., 38 (2004), 35–44. https://doi.org/10.1023/B:VLSI.0000028532.53893.82 doi: 10.1023/B:VLSI.0000028532.53893.82
    [49] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, et al., Overcoming catastrophic forgetting in neural networks, P. Natl. A. Sci. USA, 114 (2017), 3521–3526. https://doi.org/10.1073/pnas.1611835114 doi: 10.1073/pnas.1611835114
    [50] I. Loshchilov, F. Hutter, Decoupled weight decay regularization, In: ICLR 2019 International Conference on Learning Representations, 2019. Available from: https://openreview.net/forum?id = Bkg6RiCqY7.
    [51] G. B. Huang, L. Chen, C. K. Siew, Universal approximation using incremental constructive feedforward networks with random hidden nodes, IEEE T. Neural Networ., 17 (2006), 879–892. https://doi.org/10.1109/TNN.2006.875977 doi: 10.1109/TNN.2006.875977
    [52] M. D. McDonnell, T. Vladusich, Enhanced image classification with a fast-learning shallow convolutional neural network, In: 2015 International Joint Conference on Neural Networks (IJCNN), Killarney, Ireland, 2015, 1–7. https://doi.org/10.1109/IJCNN.2015.7280796
    [53] J. Cao, Z. Lin, G. B. Huang, N. Liu, Voting based extreme learning machine, Inform. Sciences, 185 (2012), 66–77. https://doi.org/10.1016/j.ins.2011.09.015 doi: 10.1016/j.ins.2011.09.015
    [54] S. Pang, Z. Yu, M. A. Orgun, A novel end-to-end classifier using domain transferred deep convolutional neural networks for biomedical images, Comput. Meth. Prog. Bio., 140 (2017), 283–293. https://doi.org/10.1016/j.cmpb.2016.12.019 doi: 10.1016/j.cmpb.2016.12.019
    [55] X. Glorot, Y. Bengio, Understanding the difficulty of training deep feedforward neural networks, In: Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) 2010, Chia La guna Resort, Sardinia, Italy, 9 (2010), 249–256. Available from: https://proceedings.mlr.press/v9/glorot10a.html.
    [56] W. J. Youden, Index for rating diagnostic tests, Cancer, 3 (1950), 32–35. https://doi.org/10.1002/1097-0142(1950)3:1%3C32::AID-CNCR2820030106%3E3.0.CO;2-3
    [57] P. Rouzrokh, B. Khosravi, S. Faghani, M. Moassefi, D. V. V. Martinez, B. S. Erickson, et al., Mitigating bias in radiology machine learning: 1. Data handling, Radiol.-Artif. Intell., 4 (2022), e210290. https://doi.org/10.1148/ryai.210290 doi: 10.1148/ryai.210290
    [58] M. Roberts, D. Driggs, M. Thorpe, J. Gilbey, M. Yeung, S. Ursprung, et al., Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans, Nat. Mach. Intell., 3 (2021), 199–217. https://doi.org/10.1038/s42256-021-00307-0 doi: 10.1038/s42256-021-00307-0
    [59] S. Elfwing, E. Uchibe, K. Doya, Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, Neural Networks, 107 (2018), 3–11. https://doi.org/10.1016/j.neunet.2017.12.012 doi: 10.1016/j.neunet.2017.12.012
    [60] P. Ramachandran, B. Zoph, Q. V. Le, Searching for activation functions, arXiv Preprint, 2017. https://doi.org/10.48550/arXiv.1710.05941
    [61] S. Ioffe, C. Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, In: Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 37 (2015), 448–456. Available from: https://proceedings.mlr.press/v37/ioffe15.html.
    [62] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, L. C. Chen, MobileNetV2: Inverted residuals and linear bottlenecks, In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, 4510–4520. https://doi.org/10.1109/CVPR.2018.00474
    [63] J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, 7132–7141. https://doi.org/10.1109/CVPR.2018.00745
    [64] M. Lin, Q. Chen, S. Yan, Network in network, arXiv Preprint, 2014. https://doi.org/10.48550/arXiv.1312.4400
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(95) PDF downloads(11) Cited by(0)

Article outline

Figures and Tables

Figures(3)  /  Tables(10)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog