Research article

A dynamic self‑adaptive optimization framework integrating the metaheuristic strategy for the non‑convex stochastic optimization problem

  • Published: 04 September 2026
  • As a critical branch of optimization theory, adaptive optimization algorithms mitigate the limitations of classical gradient descent, such as slow convergence and vulnerability to local optima, particularly in high-dimensional, non-convex environments. This is achieved by dynamically adjusting learning rates and per-parameter update rules. By integrating metaheuristic strategies with adaptive gradients, a dynamic self-adaptive optimization framework (DSOF) is proposed in this paper. The DSOF introduces a history-informed second-order momentum mechanism, which leverages recent gradient statistics to improve robustness under non-stationary objectives. Additionally, it maintains population diversity through evolutionary crossover and mutation operations to prevent premature convergence, and adaptively contracts or expands parameter bounds to balance exploration and exploitation. Experimental results demonstrated that, under the evaluated settings, the DSOF improves optimization stability by reducing the gradient-variance fluctuation coefficient from 0.418 to 0.382 in the modified national institute of standards and technology classification task, achieves a 100% global optimum hit rate on the 50-dimensional Ackley function, and reduces NAS-Bench-201 search time to 25 hours with only 1.02 GPU-days. These results verify the empirical effectiveness of the DSOF in improving stability, global exploration capability, and computational efficiency.

    Citation: Weiyi Jin, Sha Lu, Libin Liu. A dynamic self‑adaptive optimization framework integrating the metaheuristic strategy for the non‑convex stochastic optimization problem[J]. Electronic Research Archive, 2026, 34(10): 7619-7650. doi: 10.3934/era.2026329

    Related Papers:

  • As a critical branch of optimization theory, adaptive optimization algorithms mitigate the limitations of classical gradient descent, such as slow convergence and vulnerability to local optima, particularly in high-dimensional, non-convex environments. This is achieved by dynamically adjusting learning rates and per-parameter update rules. By integrating metaheuristic strategies with adaptive gradients, a dynamic self-adaptive optimization framework (DSOF) is proposed in this paper. The DSOF introduces a history-informed second-order momentum mechanism, which leverages recent gradient statistics to improve robustness under non-stationary objectives. Additionally, it maintains population diversity through evolutionary crossover and mutation operations to prevent premature convergence, and adaptively contracts or expands parameter bounds to balance exploration and exploitation. Experimental results demonstrated that, under the evaluated settings, the DSOF improves optimization stability by reducing the gradient-variance fluctuation coefficient from 0.418 to 0.382 in the modified national institute of standards and technology classification task, achieves a 100% global optimum hit rate on the 50-dimensional Ackley function, and reduces NAS-Bench-201 search time to 25 hours with only 1.02 GPU-days. These results verify the empirical effectiveness of the DSOF in improving stability, global exploration capability, and computational efficiency.



    加载中


    [1] C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, M. I. Jordan, On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points, J. ACM, 68 (2021), 1–29. https://doi.org/10.1145/3418526 doi: 10.1145/3418526
    [2] K. Xue, C. Qian, L. Xu, X. Fei, Evolutionary gradient descent for non-convex optimization, in Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, (2021), 3221–3227. https://doi.org/10.24963/ijcai.2021/443
    [3] A. Cutkosky, H. Mehta, F. Orabona, Optimal stochastic non-smooth non-convex optimization through online-to-non-convex conversion, in Proceedings of the 40th International Conference on Machine Learning, (2023), 6643–6670.
    [4] K. Verma, A. Maiti, WSAGrad: A novel adaptive gradient based method, Appl. Intell., 53 (2023), 14383–14399. https://doi.org/10.1007/s10489-022-04205-9 doi: 10.1007/s10489-022-04205-9
    [5] G. Ioannou, T. Tagaris, A. Stafylopatis, AdaLip: An adaptive learning rate method per layer for stochastic optimization, Neural Process. Lett., 55 (2023), 6311–6338. https://doi.org/10.1007/s11063-022-11140-w doi: 10.1007/s11063-022-11140-w
    [6] K. Jiang, D. Malik, Y. Li, How does adaptive optimization impact local neural network geometry? in 37th Conference on Neural Information Processing Systems (NeurIPS 2023), (2023), 8305–8384. https://doi.org/10.52202/075280-0366
    [7] F. Kunstner, A. Milligan, R. Yadav, M. Schmidt, A. Bietti, Heavy-tailed class imbalance and why Adam outperforms gradient descent on language models, in 38th Conference on Neural Information Processing Systems (NeurIPS 2024), (2024), 30106–30148. https://doi.org/10.52202/079017-0948
    [8] G. Amulya, A. Jyothi, S. Dasari, E. Sandhya, R. Kumar K V, Performance analysis of ML and DL models: Impact of linear and non-linear optimizers on model efficiency, Adv. Nonlinear Var. Inequalities, 28 (2025), 234–250. https://doi.org/10.52783/anvi.v28.2300 doi: 10.52783/anvi.v28.2300
    [9] N. Hansen, A. Ostermeier, Completely derandomized self-adaptation in evolution strategies, Evol. Comput., 9 (2001), 159–195. https://doi.org/10.1162/106365601750190398 doi: 10.1162/106365601750190398
    [10] D. Wierstra, T. Schaul, T. Glasmachers, Y. Sun, J. Peters, J. Schmidhuber, Natural evolution strategies, J. Mach. Learn. Res., 15 (2014), 949–980.
    [11] I. Loshchilov, F. Hutter, SGDR: Stochastic gradient descent with warm restarts, in International Conference on Learning Representations, (2017), 1–16.
    [12] D. Sarkar, A. Biswas, Complex preference analysis: A score-based evaluation strategy for ranking and comparison of the evolutionary algorithms, Soft Comput., 29, (2025), 1967–1980. https://doi.org/10.1007/s00500-025-10525-y doi: 10.1007/s00500-025-10525-y
    [13] B. Ghimire, A. Mahmood, K. Elleithy, Hybrid parallel ant colony optimization for application to quantum computing to solve large-scale combinatorial optimization problems, Appl. Sci., 13 (2023), 11817. https://doi.org/10.3390/app132111817 doi: 10.3390/app132111817
    [14] Y. Hao, C. Zhao, Y. Zhang, Y. Cao, Z. Li, Constrained multi-objective optimization problems: Methodologies, algorithms and applications, Knowl.-Based Syst., 299 (2024), 111998. https://doi.org/10.1016/j.knosys.2024.111998 doi: 10.1016/j.knosys.2024.111998
    [15] F. Nikbakhtsarvestani, S. Rahnamayan, M. Ebrahimi, Opposition-based multi-objective ADAM optimizer (OMAdam) for training ANN, in 2024 IEEE Congress on Evolutionary Computation (CEC), (2024), 1–10. https://doi.org/10.1109/CEC60901.2024.10612083
    [16] L. Shen, C. Chen, F. Zou, Z. Jie, J. Sun, W. Liu, A unified analysis of AdaGrad with weighted aggregation and momentum acceleration, IEEE Trans. Neural Networks Learn. Syst., 35 (2024), 14482–14490. https://doi.org/10.1109/TNNLS.2023.3279381 doi: 10.1109/TNNLS.2023.3279381
    [17] M. Reyad, A. M. Sarhan, M. Arafa, A modified Adam algorithm for deep neural network optimization, Neural Comput. Appl., 35 (2023), 17095–17112. https://doi.org/10.1007/s00521-023-08568-z doi: 10.1007/s00521-023-08568-z
    [18] H. Sun, L. Shen, Q. Zhong, L. Ding, S. Chen, J. Sun, et al., AdaSAM: Boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks, Neural Networks, 169 (2024), 506–519. https://doi.org/10.1016/j.neunet.2023.10.044 doi: 10.1016/j.neunet.2023.10.044
    [19] E. Levin, J. Kileel, N. Boumal, The effect of smooth parametrizations on nonconvex optimization landscapes, Math. Program., 209 (2025), 63–111. https://doi.org/10.1007/s10107-024-02058-3 doi: 10.1007/s10107-024-02058-3
    [20] R. Islamov, N. Ajroldi, A. Orvieto, A. Lucchi, Loss landscape characterization of neural networks without over-parametrization, in 38th Conference on Neural Information Processing Systems (NeurIPS 2024), (2024), 46680–46727. https://doi.org/10.52202/079017-1481
    [21] W. Zhang, L. Niu, D. Zhang, G. Wang, F. U. D. Farrukh, C. Zhang, HW-Adam: FPGA-based accelerator for adaptive moment estimation, Electronics, 12 (2023), 263. https://doi.org/10.3390/electronics12020263 doi: 10.3390/electronics12020263
    [22] O. Hospodarskyy, V. Martsenyuk, N. Kukharska, A. Hospodarskyy, S. Sverstiuk, Understanding the Adam optimization algorithm in machine learning, in 2nd International Workshop on Computer Information Technologies in Industry 4.0, (2024), 1–14.
    [23] G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, et al., OpenAI Gym, preprint, arXiv: 1606.01540.
    [24] P. Wawrzyński, A cat-like robot real-time learning to run, in Adaptive and Natural Computing Algorithms, (2009), 380–390. https://doi.org/10.1007/978-3-642-04921-7_39
    [25] L. Abualigah, A. Diabat, R. A. Zitar, Orthogonal learning Rosenbrock's direct rotation with the gazelle optimization algorithm for global optimization, Mathematics, 10 (2022), 4509. https://doi.org/10.3390/math10234509 doi: 10.3390/math10234509
    [26] Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition, in Proceedings of the IEEE, 86 (1998), 2278–2324. https://doi.org/10.1109/5.726791
    [27] S. Merity, C. Xiong, J. Bradbury, R. Socher, Pointer sentinel mixture models, in International Conference on Learning Representations, (2017), 1–15.
    [28] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in 2016 IEEE Conference on Computer Vision and Pattern Recognition, (2016), 770–778. https://doi.org/10.1109/CVPR.2016.90
    [29] X. Dong, Y. Yang, NAS-Bench-201: Extending the scope of reproducible neural architecture search, in International Conference on Learning Representations, (2020), 1–16.
    [30] I. Loshchilov, F. Hutter, Decoupled weight decay regularization, in International Conference on Learning Representations, (2019), 1–18.
    [31] B. Liu, L. Wu, L. Chen, K. Liang, J. Zhu, C. Liang, et al., Distributed lion for communication efficient distributed training, in Proceedings of the 38th International Conference on Neural Information Processing Systems, (2024), 18388–18415.
    [32] N. Shazeer, M. Stern, Adafactor: Adaptive learning rates with sublinear memory cost, in Proceedings of the 35th International Conference on Machine Learning, (2018), 4596–4604.
    [33] L. Wright, N. Demeure, Ranger21: A synergistic deep learning optimizer, preprint, arXiv: 2106.13731.
    [34] V. Gupta, T. Koren, Y. Singer, Shampoo: Preconditioned stochastic tensor optimization, in Proceedings of the 35th International Conference on Machine Learning, (2018), 1842–1850.
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(319) PDF downloads(34) Cited by(0)

Article outline

Figures and Tables

Figures(10)  /  Tables(7)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog