Training deep neural networks is often hindered by sharp minima, unstable optimization trajectories, and limited generalization. To tackle these problems, we introduce a novel gradient-free, population-based optimization framework, the Curvature–Interference Quantum Entropic Optimizer (CI-QEO), which integrates the empirical loss, zeroth-order curvature estimation, predictive interference suppression, entropy-driven diversity preservation, and covariance-driven tunnelling into a single energy function. We used a benchmark set of binary and multiclass classification problems, as well as nonlinear regression problems, to test the proposed method with the same network architecture. We found that CI-QEO outperforms strong baseline optimizers in terms of accuracy of binary classification (94.18% vs. 90.58%), accuracy of multiclass classification (91.36% vs. 84.34%), and R² score (0.936 vs. 0.854), with lower test loss (26.46% vs. 37.24%), mean squared error (22.54% vs. 37.24%), predictive interference (64.27% vs. 105.12%), and total optimization energy (73.66% vs. 167.83%). In addition, ablation experiments and statistical significance analysis (p < 0.05) validated each proposed component's contribution to the observed performance gains. The results highlight that CI-QEO is a stable, accurate, and large-scale optimization framework for deep neural network training and a promising alternative to prevalent gradient- and population-based optimization algorithms.
Citation: Irsa Sajjad, Osama Abdulaiz Alamri, Maria Malik, Marwan H. Alhelali, Maysoon A. Sultan. A Quantum-inspired population optimizer integrating curvature stability and predictive consistency for deep learning[J]. AIMS Mathematics, 2026, 11(9): 30374-30402. doi: 10.3934/math.20261204
Training deep neural networks is often hindered by sharp minima, unstable optimization trajectories, and limited generalization. To tackle these problems, we introduce a novel gradient-free, population-based optimization framework, the Curvature–Interference Quantum Entropic Optimizer (CI-QEO), which integrates the empirical loss, zeroth-order curvature estimation, predictive interference suppression, entropy-driven diversity preservation, and covariance-driven tunnelling into a single energy function. We used a benchmark set of binary and multiclass classification problems, as well as nonlinear regression problems, to test the proposed method with the same network architecture. We found that CI-QEO outperforms strong baseline optimizers in terms of accuracy of binary classification (94.18% vs. 90.58%), accuracy of multiclass classification (91.36% vs. 84.34%), and R² score (0.936 vs. 0.854), with lower test loss (26.46% vs. 37.24%), mean squared error (22.54% vs. 37.24%), predictive interference (64.27% vs. 105.12%), and total optimization energy (73.66% vs. 167.83%). In addition, ablation experiments and statistical significance analysis (p < 0.05) validated each proposed component's contribution to the observed performance gains. The results highlight that CI-QEO is a stable, accurate, and large-scale optimization framework for deep neural network training and a promising alternative to prevalent gradient- and population-based optimization algorithms.
| [1] | L. Bottou, Large-scale machine learning with stochastic gradient descent, in Proceedings of COMPSTAT'2010 (eds. Y. Lechevallier and G. Saporta), Physica-Verlag HD, (2010), 177–186. https://doi.org/10.1007/978-3-7908-2604-3_16 |
| [2] |
Y. LeCun, Y. Bengio, G. Hinton, Deep learning, Nature, 521 (2015), 436–444. https://doi.org/10.1038/nature14539 doi: 10.1038/nature14539
|
| [3] | D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in International Conference on Learning Representations, (2015). |
| [4] |
S. Wang, Y. Wang, S. Chen, Z. Zhou, X. Liu, Z. Li, Interactive Siamese network-based roadside perception for multi-vehicle tracking, IEEE Trans. Intell. Transp. Syst., 26 (2025), 22482–22496. https://doi.org/10.1109/TITS.2025.3611287 doi: 10.1109/TITS.2025.3611287
|
| [5] |
B. Xue, Q. Zheng, Z. Li, J. Wang, C. Mu, J. Yang, et al., Perturbation defense ultra high-speed weak target recognition, Eng. Appl. Artif. Intell., 138 (2024), 109420. https://doi.org/10.1016/j.engappai.2024.109420 doi: 10.1016/j.engappai.2024.109420
|
| [6] |
S. Hochreiter, J. Schmidhuber, Flat minima, Neural Comput., 9 (1997), 1–42. https://doi.org/10.1162/neco.1997.9.1.1 doi: 10.1162/neco.1997.9.1.1
|
| [7] |
K. Zhang, Y. Wang, U. A. Bhatti, Y. Zhou, M. Jin, Enhanced ransomware attacks detection using feature selection, sensitivity analysis, and optimized hybrid model, J. Big Data, 12 (2025), 245. https://doi.org/10.1186/s40537-025-01289-1 doi: 10.1186/s40537-025-01289-1
|
| [8] |
L. Chen, H. Wang, Y. Tang, Y. Ma, S. Wen, M. W. Geda, IWOA-optimized deep learning for bearing fault diagnosis under noisy and variable conditions, IEEE T. Instrum. Meas., 74 (2025), 1–18. https://doi.org/10.1109/TIM.2025.3602547 doi: 10.1109/TIM.2025.3602547
|
| [9] | N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. T. P. Tang, On large-batch training for deep learning: Generalization gap and sharp minima, in International Conference on Learning Representations, (2017). |
| [10] |
Z. He, H. Zhong, X. Shi, C. Zhao, J. Wen, M. Shang, Accelerating the tuning process for optimizing DNN operators by ROFT model, Sci. Rep., 15 (2025), 36327. https://doi.org/10.1038/s41598-025-20139-x doi: 10.1038/s41598-025-20139-x
|
| [11] |
A. Beck, M. Teboulle, Mirror descent and nonlinear projected subgradient methods for convex optimization, Oper. Res. Lett., 31 (2003), 167–175. https://doi.org/10.1016/S0167-6377(02)00231-6 doi: 10.1016/S0167-6377(02)00231-6
|
| [12] | J. H. Holland, Adaptation in Natural and Artificial Systems, University of Michigan Press, (1975). |
| [13] | J. Kennedy, R. Eberhart, Particle swarm optimization, in Proceedings of ICNN'95—International Conference on Neural Networks, IEEE, 4 (1995), 1942–1948. https://doi.org/10.1109/ICNN.1995.488968 |
| [14] |
K. H. Han, J. H. Kim, Quantum-inspired evolutionary algorithm for a class of combinatorial optimization, IEEE T. Evol. Comput., 6 (2002), 580–593. https://doi.org/10.1109/TEVC.2002.804320 doi: 10.1109/TEVC.2002.804320
|
| [15] |
T. Wang, M. Liu, H. Li, L. Zhao, C. Jiang, C. Xia, et al., ArchSentry: Enhanced Android malware detection via hierarchical semantic extraction, IEEE T. Netw. Serv. Man., 22 (2025), 2822–2837. https://doi.org/10.1109/TNSM.2025.3559255 doi: 10.1109/TNSM.2025.3559255
|
| [16] |
H. Robbins, S. Monro, A stochastic approximation method, Ann. Math. Stat., 22 (1951), 400–407. https://doi.org/10.1214/aoms/1177729586 doi: 10.1214/aoms/1177729586
|
| [17] |
Y. Tian, Z. Zhu, X. Zhao, X. Chen, W. Huang, X. Zhang, A dynamic and heterogeneous representation for topology optimization using evolutionary algorithms [Research Frontier], IEEE Comput. Intell. M., 20 (2025), 71–82. https://doi.org/10.1109/MCI.2025.3594654 doi: 10.1109/MCI.2025.3594654
|
| [18] | Y. Nesterov, A method for solving the convex programming problem with convergence rate $ O\left(1/{k}^{2}\right) $, Sov. Math. Dokl., 27 (1983), 372–376. |
| [19] |
M. Zhu, J. Yuan, E. Kong, L. Zhao, L. Xiao, D. Gu, Generative adversarial networks with noise optimization and pyramid coordinate attention for robust image denoising, Int. J. Intell. Syst., 2025 (2025), 1546016. https://doi.org/10.1155/int/1546016 doi: 10.1155/int/1546016
|
| [20] | J. Duchi, E. Hazan, Y. Singer, Adaptive subgradient methods for online learning and stochastic optimization, J. Mach. Learn. Res., 12 (2011), 2121–2159. |
| [21] |
T. Luo, R. Hu, Z. He, G. Jiang, H. Xu, Y. Song, et al., DiffW: Multi-encoder based on conditional diffusion model for robust image watermarking, IEEE T. Multimedia, 28 (2026), 837–852. https://doi.org/10.1109/TMM.2025.3632631 doi: 10.1109/TMM.2025.3632631
|
| [22] |
Y. Liu, J. Jie, C. Chen, H. Li, Cooperative optimization of zeroing neural networks (ZNN) in dynamic systems: Evolution, advances, and applications, Neurocomputing, 684 (2026), 133565. https://doi.org/10.1016/j.neucom.2026.133565 doi: 10.1016/j.neucom.2026.133565
|
| [23] |
M. Wan, H. Du, J. Wang, S. Li, H. Shang, H. Bai, et al., HIP-DFPT: Scalable optimization of irregular workloads in quantum perturbation on GPU clusters, IEEE T. Parallel Distr., 37 (2026), 2021–2036. https://doi.org/10.1109/TPDS.2026.3705623 doi: 10.1109/TPDS.2026.3705623
|
| [24] | T. Tieleman, G. Hinton, Divide the gradient by a running average of its recent magnitude, Neural Networks for Machine Learning, Coursera, 2012. |
| [25] | I. Goodfellow, Y. Bengio, A. Courville, Deep Learning, MIT Press, (2016). |
| [26] | H. Li, Z. Xu, G. Taylor, C. Studer, T. Goldstein, Visualizing the loss landscape of neural nets, in Advances in Neural Information Processing Systems, 31 (2018). |
| [27] |
L. He, W. Gong, J. Zhao, Research on fuel injection quantity fluctuation characteristics and optimization improvement of dual-fuel engines, Energy, 324 (2025), 135953. https://doi.org/10.1016/j.energy.2025.135953 doi: 10.1016/j.energy.2025.135953
|
| [28] |
L. Zhu, D. Han, X. Shen, C. Chen, K. C. Li, Enhancing image–text matching through multi-level semantic consistency alignment, Vis. Comput., 41 (2025), 9555–9570. https://doi.org/10.1007/s00371-025-03981-y doi: 10.1007/s00371-025-03981-y
|
| [29] |
F. Meng, H. Xu, Z. Gao, Q. Li, Real-time trajectory planning of unmanned surface vehicles: A constraint-embedded model predictive control approach, Ocean Eng., 363 (2026), 126812. https://doi.org/10.1016/j.oceaneng.2026.126812 doi: 10.1016/j.oceaneng.2026.126812
|
| [30] |
X. Xie, F. Liu, X. Zhang, G. Qin, Systematic feature selection using three-level-fused of three-view uncertainty measures for multi-granularity fuzzy γ coverings, Expert Syst. Appl., 308 (2026), 130999. https://doi.org/10.1016/j.eswa.2025.130999 doi: 10.1016/j.eswa.2025.130999
|
| [31] |
M. Feng, X. Li, J. Luo, W. Dong, Y. Wang, A. Mian, Second-order robust iterative pose optimization for fine-grained cross-view localization, IEEE T. Image Process., 35 (2026), 3157–3171. https://doi.org/10.1109/TIP.2026.3673969 doi: 10.1109/TIP.2026.3673969
|
| [32] | P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, et al., Entropy-SGD: Biasing gradient descent into wide valleys, in International Conference on Learning Representations, (2017). |
| [33] |
S. Bubeck, Convex optimization: Algorithms and complexity, Found. Trends Mach. Learn., 8 (2015), 231–357. https://doi.org/10.1561/2200000050 doi: 10.1561/2200000050
|
| [34] |
J. C. Spall, Multivariate stochastic approximation using a simultaneous perturbation gradient approximation, IEEE T. Autom. Control, 37 (1992), 332–341. https://doi.org/10.1109/9.119632 doi: 10.1109/9.119632
|
| [35] |
Y. Nesterov, V. Spokoiny, Random gradient-free minimization of convex functions, Found. Comput. Math., 17 (2017), 527–566. https://doi.org/10.1007/s10208-015-9296-2 doi: 10.1007/s10208-015-9296-2
|
| [36] | D. E. Goldberg, Genetic Algorithms in Search, Optimization, and Machine Learning, Addison-Wesley, (1989). |
| [37] | R. C. Eberhart, Y. Shi, Particle swarm optimization: Developments, applications and resources, in Proceedings of the 2001 Congress on Evolutionary Computation, IEEE, 1 (2001), 81–86. https://doi.org/10.1109/CEC.2001.934374 |
| [38] | A. Narayanan, M. Moore, Quantum-inspired genetic algorithms, in Proceedings of the IEEE International Conference on Evolutionary Computation, IEEE, (1996), 61–66. https://doi.org/10.1109/ICEC.1996.542334 |
| [39] | T. G. Dietterich, Ensemble methods in machine learning, in Multiple Classifier Systems (eds. J. Kittler and F. Roli), Springer, 1857 (2000), 1–15. https://doi.org/10.1007/3-540-45014-9_1 |
| [40] | H. Robbins, D. Siegmund, A convergence theorem for nonnegative almost supermartingales and some applications, in Optimizing Methods in Statistics (ed. J. S. Rustagi), Academic Press, (1971), 233–257. https://doi.org/10.1016/B978-0-12-604550-5.50015-8. |
| [41] |
I. Sajjad, M. M. AL Sobhi, The quantum-inspired adaptive superposition optimization for neural network training, AIMS Math., 11 (2026), 243–271. https://doi.org/10.3934/math.2026010 doi: 10.3934/math.2026010
|
| [42] |
I. Sajjad, H. M. Alshanbari, M. M. A. Almazah, H. Louati, S. Rauf, Adaptive Grover-driven optimization for quantum-inspired deep learning: A gradient-free training framework, AIMS Math., 10 (2025), 26568–26592. https://doi.org/10.3934/math.20251168 doi: 10.3934/math.20251168
|
| [43] |
I. Sajjad, O. A. Alamri, M. Malik, M. H. Alhelali, Quantum-inspired stochastic geometric hyperstate optimization for explainable deep neural learning under nonlinear manifold dynamics. AIMS Math., 11 (2026), 24419–24444. https://doi.org/10.3934/math.2026986 doi: 10.3934/math.2026986
|