Quantitative structure-activity relationship (QSAR) modeling suffers from high-dimensional feature redundancy and individual algorithm inductive bias. To address these bottlenecks, this paper proposed a two-stage multidimensional integrated feature selection and heterogeneous stacking ensemble framework (MIFS-Stacking). Evaluated on 1974 estrogen receptor alpha (ERα) inhibitor samples from DrugBank, the MIFS module first compressed the 729-dimensional descriptor space to 20 core features using heuristic filtering and a multi-algorithm weighted ensemble scoring scheme across five heterogeneous evaluators. Next, a Level 1 primary layer combining five mathematically diverse base learners (ElasticNet, KNN, SVR, random forest, and XGBoost) captured complex nonlinear patterns, while a Level 2 L2-regularized RidgeCV meta-learner synthesized predictions via out-of-fold (OOF) cross-validation to suppress secondary overfitting. Global hyperparameter optimization was efficiently conducted using the Optuna Bayesian framework. Empirical results demonstrated that MIFS-Stacking achieved superior generalization performance on the test set (R2 = 0.7375, RMSE = 0.7339, MAE = 0.5344), significantly outperforming all standalone baselines. The proposed MIFS-Stacking paradigm offered a robust, scalable, and highly accurate computational strategy for high-dimensional chemoinformatics regression tasks.
Citation: Chengkai Tu, Weifeng Jiang, Ruoyi Liu. A multi-dimensional feature selection and stacking ensemble framework for QSAR prediction[J]. Big Data and Information Analytics, 2026, 10: 177-188. doi: 10.3934/bdia.2026009
Quantitative structure-activity relationship (QSAR) modeling suffers from high-dimensional feature redundancy and individual algorithm inductive bias. To address these bottlenecks, this paper proposed a two-stage multidimensional integrated feature selection and heterogeneous stacking ensemble framework (MIFS-Stacking). Evaluated on 1974 estrogen receptor alpha (ERα) inhibitor samples from DrugBank, the MIFS module first compressed the 729-dimensional descriptor space to 20 core features using heuristic filtering and a multi-algorithm weighted ensemble scoring scheme across five heterogeneous evaluators. Next, a Level 1 primary layer combining five mathematically diverse base learners (ElasticNet, KNN, SVR, random forest, and XGBoost) captured complex nonlinear patterns, while a Level 2 L2-regularized RidgeCV meta-learner synthesized predictions via out-of-fold (OOF) cross-validation to suppress secondary overfitting. Global hyperparameter optimization was efficiently conducted using the Optuna Bayesian framework. Empirical results demonstrated that MIFS-Stacking achieved superior generalization performance on the test set (R2 = 0.7375, RMSE = 0.7339, MAE = 0.5344), significantly outperforming all standalone baselines. The proposed MIFS-Stacking paradigm offered a robust, scalable, and highly accurate computational strategy for high-dimensional chemoinformatics regression tasks.
| [1] |
Hansch C, Fujita T, (1964) p-σ-π analysis. A method for the correlation of biological activity and chemical structure. J Am Chem Soc 86: 1616–1626. https://doi.org/10.1021/ja01062a035 doi: 10.1021/ja01062a035
|
| [2] | Franke R, Gruska A, Devillers J, Chessel D, Dunn III WJ, Wold S, et al. (1995) Multivariate data analysis of chemical and biological data. In: Chemometric Methods in Molecular Design, H. van de Waterbeemd (Ed.). https://doi.org/10.1002/9783527615452.ch4 |
| [3] |
Wishart DS, Knox C, Guo AC, Shrivastava S, Hassanali M, Stothard P, et al., (2006) DrugBank: A comprehensive resource for in silico drug discovery and exploration. Nucleic Acids Res 34: D668–D672. https://doi.org/10.1093/nar/gkj067 doi: 10.1093/nar/gkj067
|
| [4] |
Lin X, Li X, Lin X, (2020) A review on applications of computational methods in drug screening and design. Molecules 25: 1375. https://doi.org/10.3390/molecules25061375 doi: 10.3390/molecules25061375
|
| [5] |
Goh GB, Hodas NO, Vishnu A, (2017) Deep learning for computational chemistry. J Comput Chem 38: 1291–1307. https://doi.org/10.1002/jcc.24764 doi: 10.1002/jcc.24764
|
| [6] |
Paul SM, Mytelka DS, Dunwiddie CT, Persinger CC, Munos BH, Lindborg SR, et al. (2010) How to improve R & D productivity: the pharmaceutical industry's grand challenge. Nat Rev Drug Discov 9: 203–214. https://doi.org/10.1038/nrd3078 doi: 10.1038/nrd3078
|
| [7] |
Zou H, Hastie T, (2005) Regularization and variable selection via the elastic net. J R Stat Soc Ser B Stat Methodol 67: 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x doi: 10.1111/j.1467-9868.2005.00503.x
|
| [8] |
Guyon I, Weston J, Barnhill S, Vapnik V, (2002) Gene selection for cancer classification using support vector machines. Mach Learn 46: 389–422. https://doi.org/10.1023/A:1012487302797 doi: 10.1023/A:1012487302797
|
| [9] |
Breiman L, (2001) Random forests. Mach Learn 45: 5–32. https://doi.org/10.1023/A:1010933404324 doi: 10.1023/A:1010933404324
|
| [10] |
Svetnik V, Liaw A, Tong C, Culberson JC, Sheridan RP, Feuston BP, (2003) Random forest: A classification and regression tool for compound classification and QSAR modeling. J Chem Inf Comput Sci 43: 1947–1958. https://doi.org/10.1021/ci034160g doi: 10.1021/ci034160g
|
| [11] |
Wang JD, Xu YS, Peng F, Xiao BS, (2020) Infrared cooperative localization optimization algorithm based on ridge regression. J Beijing Univ Aeronaut Astronaut 46: 563–570. https://doi.org/10.13700/j.bh.1001-5965.2019.0238 doi: 10.13700/j.bh.1001-5965.2019.0238
|
| [12] |
Smola AJ, Schölkopf B, (2004) A tutorial on support vector regression. Stat Comput 14: 199–222. https://doi.org/10.1023/B:STCO.0000035301.49549.88 doi: 10.1023/B:STCO.0000035301.49549.88
|
| [13] | Chen T, Guestrin C, (2016) XGBoost: A scalable tree boosting system, In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785 |
| [14] |
Zhou JY, He PF, Qiu RF, Chen G, Wu WG, (2021) Research on intrusion detection using a hybrid approach combining random forests and gradient boosting trees. J Software 32: 3254–3265. http://dx. doi.org/10.13328/j.cnki.jos.006062 doi: 10.13328/j.cnki.jos.006062
|
| [15] |
Wolpert DH, (1992) Stacked generalization. Neural Networks 5: 241–259. https://doi.org/10.1016/S0893-6080(05)80023-1 doi: 10.1016/S0893-6080(05)80023-1
|
| [16] |
He B, Gan JY, (2024) Prediction model of anti-breast cancer drug activity based on machine learning. Comput Sci Appl 14: 298–307. https://doi.org/10.12677/CSA.2024.142030 doi: 10.12677/CSA.2024.142030
|
| [17] |
Xu N, La L, (2019) Prediction of new retail coupon usage behavior based on XGBoost. J Southwest China Norm Univ Nat Sci Ed 44: 101–105. https://doi.org/10.13718/j.cnki.xsxb.2019.03.017 doi: 10.13718/j.cnki.xsxb.2019.03.017
|
| [18] |
Wu XY, Wang SH, Zhang YD, (2017) A review of the theory and applications of the k-nearest neighbors algorithm. Comput Eng Appl 53: 1–7. https://doi.org/10.3778/j.issn.1002-8331.1707-0202 doi: 10.3778/j.issn.1002-8331.1707-0202
|