Speech disorders among children will change their intelligibility and fluency. Dysarthria is one of the speech disorders connected to problems that control the muscles required to speak. Children with dysarthria frequently have low, abnormal breath that initiates problems in producing adequate breath to help speaking. They have lower-pitched, harsh, nasal speech, and extremely poor pronunciation. Furthermore, these problems make the speech of children harder to recognize. Dysarthria is generated by neurological loss and occurs earlier in the lives of children, from neurological loss prolonged earlier, like in cerebral palsy, or earlier childhood over brain damage or nervous disorders. Currently, deep learning (DL)-based methods are applied to recognize dysarthria speech disorder. In this study, we developed an Efficient Framework for the Dysarthria Speech Disorder Recognition Technique of Children Using a Hybrid Deep Learning Model (EFDSDRTC-HDLM). Our aim of the EFDSDRTC-HDLM model was to design an effective recognition technique for identifying dysarthria speech disorders in children using advanced speech processing and classification methods. The proposed framework first converted speech signals into time-frequency images through speech activity detection, framing, windowing, and time-frequency image generation. These visual representations were processed by a Vision Transformer to extract discriminative global features, which were subsequently refined by a hybrid Graph Convolutional Network–Bidirectional Long Short-Term Memory with Multi-Head Attention architecture to capture structural relationships, temporal dependencies, and salient feature representations for accurate dysarthria classification. The efficiency of the proposed framework was evaluated through extensive experiments on a benchmark dataset. Experimental results demonstrated that the proposed model achieves superior scalability and classification performances, outperforming traditional DL models as well as recent transformer-based speech recognition methods across evaluation metrics, accuracy, precision, recall and f-score of 98.73%, 98.4%, 98.73%, and 98.56%, respectively, confirming its effectiveness for pediatric dysarthria speech recognition.
Citation: Fahd N. Al-Wesabi, Wafi Bedewi. Artificial intelligence with vision transformer based hybrid deep neural network for paediatric speech disorder recognition using dysarthric voice data[J]. AIMS Mathematics, 2026, 11(8): 25100-25122. doi: 10.3934/math.20261009
Speech disorders among children will change their intelligibility and fluency. Dysarthria is one of the speech disorders connected to problems that control the muscles required to speak. Children with dysarthria frequently have low, abnormal breath that initiates problems in producing adequate breath to help speaking. They have lower-pitched, harsh, nasal speech, and extremely poor pronunciation. Furthermore, these problems make the speech of children harder to recognize. Dysarthria is generated by neurological loss and occurs earlier in the lives of children, from neurological loss prolonged earlier, like in cerebral palsy, or earlier childhood over brain damage or nervous disorders. Currently, deep learning (DL)-based methods are applied to recognize dysarthria speech disorder. In this study, we developed an Efficient Framework for the Dysarthria Speech Disorder Recognition Technique of Children Using a Hybrid Deep Learning Model (EFDSDRTC-HDLM). Our aim of the EFDSDRTC-HDLM model was to design an effective recognition technique for identifying dysarthria speech disorders in children using advanced speech processing and classification methods. The proposed framework first converted speech signals into time-frequency images through speech activity detection, framing, windowing, and time-frequency image generation. These visual representations were processed by a Vision Transformer to extract discriminative global features, which were subsequently refined by a hybrid Graph Convolutional Network–Bidirectional Long Short-Term Memory with Multi-Head Attention architecture to capture structural relationships, temporal dependencies, and salient feature representations for accurate dysarthria classification. The efficiency of the proposed framework was evaluated through extensive experiments on a benchmark dataset. Experimental results demonstrated that the proposed model achieves superior scalability and classification performances, outperforming traditional DL models as well as recent transformer-based speech recognition methods across evaluation metrics, accuracy, precision, recall and f-score of 98.73%, 98.4%, 98.73%, and 98.56%, respectively, confirming its effectiveness for pediatric dysarthria speech recognition.
| [1] |
M. Shahin, U. Zafar, B. Ahmed, The automatic detection of speech disorders in children: Challenges, opportunities, and preliminary results, IEEE J. STSP, 14 (2020), 400–412. https://doi.org/10.1109/JSTSP.2019.2959393 doi: 10.1109/JSTSP.2019.2959393
|
| [2] |
J. Iuzzini-Seigel, K. M. Allison, R. Stoeckel, A tool for differential diagnosis of childhood apraxia of speech and dysarthria in children: A tutorial, Lang. Speech Hear. Ser., 53(2022), 926–946. https://doi.org/10.1044/2022_LSHSS-21-00164 doi: 10.1044/2022_LSHSS-21-00164
|
| [3] |
G. Tartarisco, R. Bruschetta, S. Summa, L. Ruta, M. Favetta, M. Busà, Artificial intelligence for dysarthria assessment in children with ataxia: A hierarchical approach, IEEE Access, 9 (2021), 166720–166735. https://doi.org/10.1109/ACCESS.2021.3135078 doi: 10.1109/ACCESS.2021.3135078
|
| [4] |
S. H. Lee, M. Kim, H. G. Seo, B. M. Oh, G. Lee, J. H. Leigh, Assessment of dysarthria using one-word speech recognition with hidden Markov models, J. Korean Med. Sci., 34 (2019), e108. https://doi.org/10.3346/jkms.2019.34.e108 doi: 10.3346/jkms.2019.34.e108
|
| [5] |
A. Al-Ali, S. Al-Maadeed, M. Saleh, R. C. Naidu, Z. C. Alex, P. Ramachandran, The detection of dysarthria severity levels using AI models: A review, IEEE Access, 12 (2024), 48223–48238. https://doi.org/10.1109/ACCESS.2024.3382574 doi: 10.1109/ACCESS.2024.3382574
|
| [6] |
A. A. Anthony, C. M. Patil, J. Basavaiah, A review on speech disorders and processing of disordered speech, Wireless Pers Commun, 126 (2022), 1621–1631. https://doi.org/10.1007/s11277-022-09812-w doi: 10.1007/s11277-022-09812-w
|
| [7] |
H. P. Rowe, S. E. Gutz, M. F. Maffei, K. Tomanek, J. R. Green, Characterizing dysarthria diversity for automatic speech recognition: A tutorial from the clinical perspective, Front. Comput. Sci., 4 (2022), 770210. https://doi.org/10.3389/fcomp.2022.770210 doi: 10.3389/fcomp.2022.770210
|
| [8] |
M. Kooi-van Es, C. E. Erasmus, B. J. de Swart, N. B. Voet, P. J. van der Wees, I. J. de Groot, et al., Dysphagia and dysarthria in children with neuromuscular diseases, a prevalence study, J. Neuromuscul. Dis., 7 (2020), 287–295. https://doi.org/10.3233/JND-190436 doi: 10.3233/JND-190436
|
| [9] |
P. McCabe, J. Korkalainen, D. Thomas, Diagnostic uncertainty in childhood motor speech disorders: A review of recent tools and approaches, Curr. Dev. Disord. Rep., 11 (2024), 105–112. https://doi.org/10.1007/s40474-024-00295-x doi: 10.1007/s40474-024-00295-x
|
| [10] |
H. L. Fouad, H. A. Abdulmohsin, A hybrid speech recognition system using deep learning methods, J. Intell. Syst. Internet Things, 15 (2025), 105–121. https://doi.org/10.54216/JISIoT.150109 doi: 10.54216/JISIoT.150109
|
| [11] |
N. A. Aljarallah, A. K. Dutta, A. R. W. Sait, Image classification-driven speech disorder detection using deep learning technique, SLAS Technol., 32 (2025), 100261. https://doi.org/10.1016/j.slast.2025.100261 doi: 10.1016/j.slast.2025.100261
|
| [12] |
D. Zhang, H. Zhang, W. Lu, W. Li, J. Wang, J. Wei, Long-range and non-stationary encoding for dysarthric speech data augmentation, IEEE J. STSP, 19 (2025), 767–782. https://doi.org/10.1109/JSTSP.2025.3562417 doi: 10.1109/JSTSP.2025.3562417
|
| [13] |
A. T. Celin Mariya, P. Vijayalakshmi, T. Nagarajan. K. Mrinalini, Augmentative and alternative speech communication (AASC) aid for people with dysarthria, Comput. Speech Lang., 92 (2025), 101777. https://doi.org/10.1016/j.csl.2025.101777 doi: 10.1016/j.csl.2025.101777
|
| [14] | M. Usha, An IoT and data mining-based tool for early identification of speech disorders in children using advanced algorithms. In: Artificial intelligence based smart and secured applications, ASCIS 2024, 2024,375–384. https://doi.org/10.1007/978-3-031-86290-8_27 |
| [15] |
A. K. Dutta, A. R. W. Sait, A speech disorder detection model using ensemble learning approach, J. Disabil. Res., 3 (2024), 1–8. https://doi.org/10.57197/JDR-2024-0026 doi: 10.57197/JDR-2024-0026
|
| [16] | K. Mittal, K. S. Gill, K. Rajput, V. Singh, Enhancing the diagnosis of speech disorders: An in-depth investigation into dysarthria classification Using the ResNet18 Model. In: 2024 IEEE International Conference on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS), 2024. https://doi.org/10.1109/ICITEICS61368.2024.10625627 |
| [17] |
Y. Lin, L. Wang, Y. Yang, J. Dang, CFDRN: A cognition-inspired feature decomposition and recombination network for dysarthric speech recognition, IEEE/ACM T. Audio SPE, 31 (2023), 3824–3836. https://doi.org/10.1109/TASLP.2023.3319276 doi: 10.1109/TASLP.2023.3319276
|
| [18] |
B. Jolad, R. Khanai, Competitive crow search algorithm-based hierarchical attention network for dysarthric speech recognition, Int. J. Wirel. Mobi. Comput., 25 (2023), 340–352. https://doi.org/10.1504/ijwmc.2023.135384 doi: 10.1504/ijwmc.2023.135384
|
| [19] |
J. Wu, Y. G. Wang, G. J. McLachlan, Informative missingness and its implications in semi-supervised learning, Innovation Inform., 2 (2026), 100033. https://doi.org/10.59717/j.xinn-inform.2026.100033 doi: 10.59717/j.xinn-inform.2026.100033
|
| [20] |
N. Xu, J. Wu, F. Cai, X. A. Li, H. B. Xie, ViT-GCN: A novel hybrid model for accurate pneumonia diagnosis from x-ray images, Biomed. Phys. Eng. Expr., 11 (2025), 045034. https://doi.org/10.1088/2057-1976/adebf4 doi: 10.1088/2057-1976/adebf4
|
| [21] |
S. Aurobindo, R. Prakash, M. Rajeshkumar, Comparative analysis of different time-frequency image representations for the detection and severity classification of dysarthric speech using deep learning, Results Eng., 25 (2025), 104561. https://doi.org/10.1016/j.rineng.2025.104561 doi: 10.1016/j.rineng.2025.104561
|
| [22] |
E. Alattas, J. Clark, A. Al-Aama, S. K. Jarraya, Evaluating features and variations in deepfake videos using the CoAtNet model, J Imaging, 11 (2025), 194. https://doi.org/10.3390/jimaging11060194 doi: 10.3390/jimaging11060194
|
| [23] |
W. Zhou, W. Wang, X. Wang, Research on lightning prediction based on GCN-LSTM model integrating spatiotemporal features, Atmosphere, 16 (2025), 447. https://doi.org/10.3390/atmos16040447 doi: 10.3390/atmos16040447
|
| [24] | Dysarthria and Non-Dysarthria Speech Dataset. Available from: https://www.kaggle.com/datasets/poojag718/dysarthria-and-nondysarthria-speech-dataset. |
| [25] | Noise Reduced UASPEECH Dysarthria Dataset. Available from: https://www.kaggle.com/datasets/aryashah2k/noise-reduced-uaspeech-dysarthria-dataset |
| [26] |
S. V. Vishnika, S. Chandrakala, Investigation of DNN-HMM and lattice-free maximum mutual information approaches for impaired speech recognition, IEEE Access, 9 (2021), 168840–168849. https://doi.org/10.1109/ACCESS.2021.3129847 doi: 10.1109/ACCESS.2021.3129847
|
| [27] |
N. A. Aljarallah, A. K. Dutta, A. R. W. Sait, Image classification-driven speech disorder detection using deep learning technique, SLAS Technol., 32 (2025), 100261. https://doi.org/10.1016/j.slast.2025.100261 doi: 10.1016/j.slast.2025.100261
|
| [28] |
F. Javanmardi, S. R. Kadiri, P. Alku, Pre-trained models for detection and severity level classification of dysarthria from speech, Speech Commun., 158 (2024), 103047. https://doi.org/10.1016/j.specom.2024.103047 doi: 10.1016/j.specom.2024.103047
|