Substructural Analysis (SSA) is a ligand-based virtual screening method that ranks compounds by combining fragment weights, where each weight reflects the contribution of a molecular substructure to the separation of active and inactive compounds. However, conventional SSA weighting schemes are often static and may be less adaptive to heterogeneous and highly imbalanced bioactivity datasets. In this study, we proposed a hybrid SSA framework that used Genetic Algorithms (GA) to optimize fragment-level weights and Genetic Programming (GP) to evolve symbolic weighting equations guided by a Chemical Weighted Root Mean Square Error (CW-RMSE) fitness function. Enrichment factor at the top 1% of the ranked list (EF@1%) was used as the main early-recognition metric because it measured active-compound retrieval relative to random selection in the most practically relevant part of a screening list. The framework was evaluated using 15 ChEMBL-derived activity classes represented by molecular fingerprints. SSA-GP achieved higher median EF@1% than GA-based SSA in 14 of 15 classes, with statistically significant improvements in 13 of these classes using the Mann-Whitney U test (p < 0.05). Notable median EF@1% gains were observed for RNN (16.98 to 69.34), MMP1 (36.67 to 83.33), 5HT1A (3.13 to 46.25), and FXA (15.83 to 57.91), while AT1 was the exception in which GA performed better. These results indicated that SSA-GP, guided by the CW-RMSE fitness function, could improve early active-compound prioritization while retaining the interpretability advantage of symbolic SSA models. Further validation against non-SSA machine-learning baselines and expert chemical interpretation of the generated equations is recommended.
Citation: Mohd Isrul Esa, Nor Samsiah Sani, Salwani Abdullah. Improving substructural analysis with chemical weighted RMSE in symbolic regression[J]. AIMS Medical Science, 2026, 13(3): 159-186. doi: 10.3934/medsci.2026012
Substructural Analysis (SSA) is a ligand-based virtual screening method that ranks compounds by combining fragment weights, where each weight reflects the contribution of a molecular substructure to the separation of active and inactive compounds. However, conventional SSA weighting schemes are often static and may be less adaptive to heterogeneous and highly imbalanced bioactivity datasets. In this study, we proposed a hybrid SSA framework that used Genetic Algorithms (GA) to optimize fragment-level weights and Genetic Programming (GP) to evolve symbolic weighting equations guided by a Chemical Weighted Root Mean Square Error (CW-RMSE) fitness function. Enrichment factor at the top 1% of the ranked list (EF@1%) was used as the main early-recognition metric because it measured active-compound retrieval relative to random selection in the most practically relevant part of a screening list. The framework was evaluated using 15 ChEMBL-derived activity classes represented by molecular fingerprints. SSA-GP achieved higher median EF@1% than GA-based SSA in 14 of 15 classes, with statistically significant improvements in 13 of these classes using the Mann-Whitney U test (p < 0.05). Notable median EF@1% gains were observed for RNN (16.98 to 69.34), MMP1 (36.67 to 83.33), 5HT1A (3.13 to 46.25), and FXA (15.83 to 57.91), while AT1 was the exception in which GA performed better. These results indicated that SSA-GP, guided by the CW-RMSE fitness function, could improve early active-compound prioritization while retaining the interpretability advantage of symbolic SSA models. Further validation against non-SSA machine-learning baselines and expert chemical interpretation of the generated equations is recommended.
active-compound count
average precision
area under the receiver operating characteristic curve
chemical weighted root mean square error
enrichment factor
genetic algorithm
genetic programming
inactive-compound count
ligand-based virtual screening
Molecular ACCess System
number of observations
number of active compounds
number of inactive compounds
root mean square error
substructural analysis
total-compound count
| [1] |
Alves VM, Yasgar A, Wellnitz J, et al. (2023) Lies and liabilities: Computational assessment of high-throughput screening hits to identify artifact compounds. J Med Chem 66: 12828-12839. https://doi.org/10.1021/acs.jmedchem.3c00482
|
| [2] |
Chung HH, Kao CY, Wang TSA, et al. (2021) Reaction tracking and high-throughput screening of active compounds in combinatorial chemistry by tandem mass spectrometry molecular networking. Anal Chem 93: 2456-2463. https://doi.org/10.1021/acs.analchem.0c04481
|
| [3] |
Feng T, Basu P, Sun W, et al. (2019) Optimal design for high-throughput screening via false discovery rate control. Stat Med 38: 2816-2827. https://doi.org/10.1002/sim.8144
|
| [4] |
Muegge I, Bentzien J, Ge Y (2024) Perspectives on current approaches to virtual screening in drug discovery. Expert Opin Drug Discov 19: 1173-1183. https://doi.org/10.1080/17460441.2024.2390511
|
| [5] |
Warr WA, Nicklaus MC, Nicolaou CA, et al. (2022) Exploration of ultralarge compound collections for drug discovery. J Chem Inf Model 62: 2021-2034. https://doi.org/10.1021/acs.jcim.2c00224
|
| [6] |
Shen C, Hu Y, Wang Z, et al. (2021) Beware of the generic machine learning-based scoring functions in structure-based virtual screening. Brief Bioinform 22: bbaa070. https://doi.org/10.1093/bib/bbaa070
|
| [7] |
Choi H, Kang H, Chung KC, et al. (2019) Development and application of a comprehensive machine learning program for predicting molecular biochemical and pharmacological properties. Phys Chem Chem Phys 21: 5189-5199. https://doi.org/10.1039/C8CP07002D
|
| [8] |
Niazi SK, Mariam Z (2023) Recent advances in machine-learning-based chemoinformatics: A comprehensive review. Int J Mol Sci 24: 11488. https://doi.org/10.3390/ijms241411488
|
| [9] |
Ross GA, Morris GM, Biggin PC (2013) One size does not fit all: The limits of structure-based models in drug discovery. J Chem Theory Comput 9: 4266-4274. https://doi.org/10.1021/ct4004228
|
| [10] |
Chauhan SS, Jamal T, Singh A, et al. (2023) Structure-based virtual screening. Cheminformatics, QSAR and Machine Learning Applications for Novel Drug Development . Elsevier 239-262. https://doi.org/10.1016/B978-0-443-18638-7.00016-5
|
| [11] |
Cramer RD, Redi G, Berkoff CE (1974) Substructural analysis. A novel approach to the problem of drug design. J Med Chem 17: 533-535. https://doi.org/10.1021/jm00251a014
|
| [12] |
Machado-Alba JE, Jiménez-Morales AL, Moran-Yela YC, et al. (2020) Adverse drug reactions associated with the use of biological agents. PLoS One 15: e0240276. https://doi.org/10.1371/journal.pone.0240276
|
| [13] |
Villar HO, Hansen MR, Kho R (2007) Substructural analysis in drug discovery. Curr Comput Aided-Drug Des 3: 59-67. https://doi.org/10.2174/157340907780058745
|
| [14] |
Abdo A, Salim N (2011) New fragment weighting scheme for the Bayesian inference network in ligand-based virtual screening. J Chem Inf Model 51: 25-32. https://doi.org/10.1021/ci100232h
|
| [15] |
Holliday JD, Sani N, Willett P (2015) Calculation of substructural analysis weights using a genetic algorithm. J Chem Inf Model 55: 214-221. https://doi.org/10.1021/ci500540s
|
| [16] | Holliday JD, Sani N, Willett P (2018) Ligand–based virtual screening using a genetic algorithm with data fusion. Match Commun Math Co 80: 623-638. |
| [17] |
Wang M, Wu Z, Wang J, et al. (2024) Genetic algorithm-based receptor ligand: A genetic algorithm-guided generative model to boost the novelty and drug-likeness of molecules in a sampling chemical space. J Chem Inf Model 64: 1213-1228. https://doi.org/10.1021/acs.jcim.3c01964
|
| [18] |
Wang Z, Sobey A (2020) A comparative review between Genetic Algorithm use in composite optimization and the state-of-the-art in evolutionary computation. Compos Struct 233: 111739. https://doi.org/10.1016/j.compstruct.2019.111739
|
| [19] |
Greenstein BL, Elsey DC, Hutchison GR (2023) Determining best practices for using genetic algorithms in molecular discovery. J Chem Phys 159: 091501. https://doi.org/10.1063/5.0158053
|
| [20] | Chiang TC, Chang CH, Yu TL (2024) A novel symbolic regressor enhancer using genetic programming. 2024 IEEE Congress on Evolutionary Computation (CEC) : 1-8. https://doi.org/10.1109/CEC60901.2024.10612124 |
| [21] |
Mei Y, Chen Q, Lensen A, et al. (2023) Explainable artificial intelligence by genetic programming: A survey. IEEE T Evolut Comput 27: 621-641. https://doi.org/10.1109/TEVC.2022.3225509
|
| [22] |
Fleck P, Werth B, Affenzeller M (2024) Population dynamics in genetic programming for dynamic symbolic regression. Appl Sci 14: 596. https://doi.org/10.3390/app14020596
|
| [23] |
Tran B, Xue B, Zhang M (2016) Genetic programming for feature construction and selection in classification on high-dimensional data. Memet Comput 8: 3-15. https://doi.org/10.1007/s12293-015-0173-y
|
| [24] |
Darwish SM, Shendi TA, Younes A (2019) Quantum-inspired genetic programming model with application to predict toxicity degree for chemical compounds. Expert Syst 36: 1-15. https://doi.org/10.1111/exsy.12415
|
| [25] |
Deighan DS, Field SE, Capano CD, et al. (2021) Genetic-algorithm-optimized neural networks for gravitational wave classification. Neural Comput Appl 33: 13859-13883. https://doi.org/10.1007/s00521-021-06024-4
|
| [26] |
Stanovov V, Akhmedova S, Semenkin E (2022) The automatic design of parameter adaptation techniques for differential evolution with genetic programming. Knowl-Based Syst 239: 108070. https://doi.org/10.1016/j.knosys.2021.108070
|
| [27] |
Fan Q, Bi Y, Xue B, et al. (2024) A genetic programming-based method for image classification with small training data. Knowl-Based Syst 283: 111188. https://doi.org/10.1016/j.knosys.2023.111188
|
| [28] |
Burlacu B, Yang K, Affenzeller M (2024) Population diversity and inheritance in genetic programming for symbolic regression. Nat Comput 23: 531. https://doi.org/10.1007/s11047-022-09934-x
|
| [29] |
Lee CM, Ahn CW, Kim MJ (2024) Feature optimization and dropout in genetic programming for data-limited image classification. Mathematics 12: 3661. https://doi.org/10.3390/math12233661
|
| [30] |
Makke N, Chawla S (2024) Interpretable scientific discovery with symbolic regression: a review. Artif Intell Rev 57: 2. https://doi.org/10.1007/s10462-023-10622-0
|
| [31] |
Archetti F, Giordani I, Vanneschi L (2010) Genetic programming for QSAR investigation of docking energy. Appl Soft Comput 10: 170-182. https://doi.org/10.1016/j.asoc.2009.06.013
|
| [32] |
Gardiner EJ, Gillet VJ (2015) Perspectives on knowledge discovery algorithms recently introduced in chemoinformatics: Rough set theory, association rule mining, emerging patterns, and formal concept analysis. J Chem Inf Model 55: 1781-1803. https://doi.org/10.1021/acs.jcim.5b00198
|
| [33] |
Orlov AA, Zherebker A, Eletskaya AA, et al. (2019) Examination of molecular space and feasible structures of bioactive components of humic substances by FTICR MS data mining in ChEMBL database. Sci Rep 9: 12066. https://doi.org/10.1038/s41598-019-48000-y
|
| [34] |
Zdrazil B (2025) Fifteen years of ChEMBL and its role in cheminformatics and drug discovery. J Cheminform 17: 32. https://doi.org/10.1186/s13321-025-00963-z
|
| [35] |
Miralavy I, Bricco AR, Gilad AA, et al. (2022) Using genetic programming to predict and optimize protein function. PeerJ Phys Chem 4: e24. https://doi.org/10.7717/peerj-pchem.24
|
| [36] |
Krüger DM, Evers A (2010) Comparison of structure- and ligand-based virtual screening protocols considering hit list complementarity and enrichment factors. ChemMedChem 5: 148-158. https://doi.org/10.1002/cmdc.200900314
|