Research article

RDAS: Reliability-driven adaptive supervision for noisy-label facial expression recognition

  • Published: 25 August 2026
  • In-the-wild facial expression recognition (FER) is inherently affected by ambiguous expressions and subjective annotations, making the observed labels only partially reliable. Conventional one-hot supervision nevertheless treats these labels as ground truth and may drive deep networks to memorize corrupted annotations. This paper proposes RDAS, a reliability-driven adaptive supervision framework for noisy-label FER. The central idea was to use label reliability not merely as a sample weight, but as a variable that determines the supervision policy of each training sample. RDAS enhances expression-sensitive representations by combining RGB appearance with structural high-frequency cues, estimates label reliability with an exponential moving average (EMA) teacher through the margin between the annotated class and its strongest competing class, and assigns reliability-specific objectives to high-, medium-, and low-reliability samples. Reliable samples retained hard-label learning, ambiguous samples received teacher-guided soft correction, and unreliable samples were prevented from dominating training with misleading gradients. Experiments on RAF-DB, AffectNet, and FERPlus datasets showed that RDAS improved both standard recognition performance and robustness under synthetic label noise. On RAF-DB, RDAS achieved an average accuracy of $ 89.70\% $ on the original dataset and $ 86.15\% $ under 30% label noise. Ablation studies provided quantitative support for reliability modeling and adaptive supervision, while the visual analyses provided complementary qualitative evidence consistent with the proposed mechanism.

    Citation: Xuefeng Zhao, Yixuan Dong. RDAS: Reliability-driven adaptive supervision for noisy-label facial expression recognition[J]. Electronic Research Archive, 2026, 34(10): 7274-7298. doi: 10.3934/era.2026315

    Related Papers:

  • In-the-wild facial expression recognition (FER) is inherently affected by ambiguous expressions and subjective annotations, making the observed labels only partially reliable. Conventional one-hot supervision nevertheless treats these labels as ground truth and may drive deep networks to memorize corrupted annotations. This paper proposes RDAS, a reliability-driven adaptive supervision framework for noisy-label FER. The central idea was to use label reliability not merely as a sample weight, but as a variable that determines the supervision policy of each training sample. RDAS enhances expression-sensitive representations by combining RGB appearance with structural high-frequency cues, estimates label reliability with an exponential moving average (EMA) teacher through the margin between the annotated class and its strongest competing class, and assigns reliability-specific objectives to high-, medium-, and low-reliability samples. Reliable samples retained hard-label learning, ambiguous samples received teacher-guided soft correction, and unreliable samples were prevented from dominating training with misleading gradients. Experiments on RAF-DB, AffectNet, and FERPlus datasets showed that RDAS improved both standard recognition performance and robustness under synthetic label noise. On RAF-DB, RDAS achieved an average accuracy of $ 89.70\% $ on the original dataset and $ 86.15\% $ under 30% label noise. Ablation studies provided quantitative support for reliability modeling and adaptive supervision, while the visual analyses provided complementary qualitative evidence consistent with the proposed mechanism.



    加载中


    [1] R. W. Picard, Affective Computing, MIT Press, 1997. https://doi.org/10.7551/mitpress/1140.001.0001
    [2] M. Pantic, L. J. M. Rothkrantz, Automatic analysis of facial expressions: The state of the art, IEEE Trans. Pattern Anal. Mach. Intell., 22 (2000), 1424–1445. https://doi.org/10.1109/34.895976 doi: 10.1109/34.895976
    [3] Z. Zeng, M. Pantic, G. I. Roisman, T. S. Huang, A survey of affect recognition methods: Audio, visual, and spontaneous expressions, IEEE Trans. Pattern Anal. Mach. Intell., 31 (2009), 39–58. https://doi.org/10.1109/TPAMI.2008.52 doi: 10.1109/TPAMI.2008.52
    [4] I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, et al., Challenges in representation learning: A report on three machine learning contests, in Neural Information Processing, (2013), 117–124. https://doi.org/10.1007/978-3-642-42051-1_16
    [5] E. Barsoum, C. Zhang, C. C. Ferrer, Z. Zhang, Training deep networks for facial expression recognition with crowd-sourced label distribution, in Proceedings of the ACM International Conference on Multimodal Interaction, (2016), 279–283. https://doi.org/10.1145/2993148.2993165
    [6] S. Li, W. Deng, J. Du, Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2017), 2584–2593. https://doi.org/10.1109/CVPR.2017.277
    [7] A. Mollahosseini, B. Hasani, M. H. Mahoor, AffectNet: A database for facial expression, valence, and arousal computing in the wild, IEEE Trans. Affect. Comput., 10 (2019), 18–31. https://doi.org/10.1109/TAFFC.2017.2740923 doi: 10.1109/TAFFC.2017.2740923
    [8] D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, et al., A closer look at memorization in deep networks, in Proceedings of the International Conference on Machine Learning, (2017), 233–242.
    [9] K. Wang, X. Peng, J. Yang, S. Lu, Y. Qiao, Suppressing uncertainties for large-scale facial expression recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2020), 6897–6906. https://doi.org/10.1109/CVPR42600.2020.00693
    [10] Y. Zhang, C. Wang, X. Ling, W. Deng, Relative uncertainty learning for facial expression recognition, in Advances in Neural Information Processing Systems, 34 (2021), 17616–17627.
    [11] Y. Zhang, C. Wang, W. Deng, Learn from all: Erasing attention consistency for noisy label facial expression recognition, in Proceedings of the European Conference on Computer Vision, (2022), 418–434. https://doi.org/10.1007/978-3-031-19809-0_24
    [12] Y. Tan, H. Xia, S. Song, Robust consistency learning for facial expression recognition under label noise, Vis. Comput., 41 (2025), 2655–2667. https://doi.org/10.1007/s00371-024-03558-1 doi: 10.1007/s00371-024-03558-1
    [13] H. Xia, C. Su, S. Song, Y. Tan, Dual-consistency constraints network for noisy facial expression recognition, Image Vis. Comput., 148 (2024), 105141. https://doi.org/10.1016/j.imavis.2024.105141 doi: 10.1016/j.imavis.2024.105141
    [14] X. Zhang, Y. Lu, H. Yan, J. Huang, Y. Gu, Y. Ji, et al., ReSup: Reliable label noise suppression for facial expression recognition, IEEE Trans. Affect. Comput., 16 (2025), 2006–2019. https://doi.org/10.1109/TAFFC.2025.3549017 doi: 10.1109/TAFFC.2025.3549017
    [15] N. Le, K. Nguyen, Q. Tran, E. Tjiputra, B. Le, A. Nguyen, Uncertainty-aware label distribution learning for facial expression recognition, in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, (2023), 6088–6097. https://doi.org/10.1109/WACV56688.2023.00603
    [16] J. Zheng, B. Li, S. Zhang, S. Wu, L. Cao, S. Ding, Attack can benefit: An adversarial approach to recognizing facial expressions under noisy annotations, in Proceedings of the AAAI Conference on Artificial Intelligence, 37 (2023), 3660–3668. https://doi.org/10.1609/aaai.v37i3.25477
    [17] Y. Li, H. Liu, D. Jiang, J. Liang, Facial expression recognition with label-noisy under dual-branch noise extraction and suppression, IEEE Trans. Affect. Comput., 16 (2025), 1514–1525. https://doi.org/10.1109/TAFFC.2024.3519359 doi: 10.1109/TAFFC.2024.3519359
    [18] J. Ye, D. Liu, C. Wang, H. Huang, L. Ying, L. Zhang, et al., CNLA: Collaborative noisy label adaptive learning for facial expression recognition, Inf. Sci., 719 (2025), 122436. https://doi.org/10.1016/j.ins.2025.122436 doi: 10.1016/j.ins.2025.122436
    [19] J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2018), 7132–7141. https://doi.org/10.1109/CVPR.2018.00745
    [20] S. Woo, J. Park, J. Y. Lee, I. S. Kweon, CBAM: Convolutional block attention module, in Proceedings of the European Conference on Computer Vision, (2018), 3–19. https://doi.org/10.1007/978-3-030-01234-2_1
    [21] K. Wang, X. Peng, J. Yang, D. Meng, Y. Qiao, Region attention networks for pose and occlusion robust facial expression recognition, IEEE Trans. Image Process., 29 (2020), 4057–4069. https://doi.org/10.1109/TIP.2019.2956143 doi: 10.1109/TIP.2019.2956143
    [22] Y. Fan, J. C. K. Lam, V. O. K. Li, Facial expression recognition with deeply-supervised attention network, IEEE Trans. Affective Comput., 13 (2022), 1057–1071. https://doi.org/10.1109/TAFFC.2020.2988264 doi: 10.1109/TAFFC.2020.2988264
    [23] F. Xue, Q. Wang, G. Guo, TransFER: Learning relation-aware facial expression representations with transformers, in Proceedings of the IEEE/CVF International Conference on Computer Vision, (2021), 3601–3610. https://doi.org/10.1109/ICCV48922.2021.00358
    [24] Z. Wen, W. Lin, T. Wang, G. Xu, Distract your attention: Multi-head cross attention network for facial expression recognition, Biomimetics, 8 (2023), 199. https://doi.org/10.3390/biomimetics8020199 doi: 10.3390/biomimetics8020199
    [25] C. Zheng, M. Mendieta, C. Chen, POSTER: A pyramid cross-fusion transformer network for facial expression recognition, in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, (2023), 3146–3155. https://doi.org/10.1109/ICCVW60793.2023.00339
    [26] J. Mao, R. Xu, X. Yin, Y. Chang, B. Nie, A. Huang, POSTER++: A simpler and stronger facial expression recognition network, Pattern Recognit., 157 (2025), 110951. https://doi.org/10.1016/j.patcog.2024.110951 doi: 10.1016/j.patcog.2024.110951
    [27] F. Ma, B. Sun, S. Li, Transformer-augmented network with online label correction for facial expression recognition, IEEE Trans. Affective Comput., 15 (2024), 593–605. https://doi.org/10.1109/TAFFC.2023.3285231 doi: 10.1109/TAFFC.2023.3285231
    [28] Z. Wu, J. Cui, LA-Net: Landmark-aware learning for reliable facial expression recognition under label noise, in Proceedings of the IEEE/CVF International Conference on Computer Vision, (2023), 20698–20707. https://doi.org/10.1109/ICCV51070.2023.01892
    [29] C. Y. Che, H. M. Sun, Y. X. Chen, S. Yang, R. S. Jia, NLFER: Multi-branch attention cross-fusion for robust facial expression recognition amidst noisy labels, Vis. Comput., 42 (2026), 1. https://doi.org/10.1007/s00371-025-04210-2 doi: 10.1007/s00371-025-04210-2
    [30] J. She, Y. Hu, H. Shi, J. Wang, Q. Shen, T. Mei, Dive into ambiguity: Latent distribution mining and pairwise uncertainty estimation for facial expression recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2021), 6248–6257. https://doi.org/10.1109/CVPR46437.2021.00618
    [31] J. Lee, Y. Choi, H. Kim, I. J. Kim, G. P. Nam, Navigating label ambiguity for facial expression recognition in the wild, in Proceedings of the AAAI Conference on Artificial Intelligence, 39 (2025), 4517–4525. https://doi.org/10.1609/aaai.v39i4.32476
    [32] D. Li, W. Xiong, T. Luo, L. Zhang, 3WAUS: A novel three-way adaptive uncertainty-suppressing model for facial expression recognition, Inf. Sci., 677 (2024), 120962. https://doi.org/10.1016/j.ins.2024.120962 doi: 10.1016/j.ins.2024.120962
    [33] Y. Apedo, H. Tao, A weakly supervised pavement crack segmentation based on adversarial learning and transformers, Multimedia Syst., 31 (2025), 266. https://doi.org/10.1007/s00530-025-01850-1 doi: 10.1007/s00530-025-01850-1
    [34] Y. Apedo, H. Tao, W. Gao, C. Xie, S. Zhao, Unsupervised domain adaptation for crack segmentation via cross-domain stylization and dual adversarial feature learning, J. Comput. Civ. Eng., 40 (2026), 04025146. https://doi.org/10.1061/JCCEE5.CPENG-7223 doi: 10.1061/JCCEE5.CPENG-7223
    [35] H. Zhou, H. Tao, Q. Zhang, C. Xie, B. Li, MRKD-PBCL: Multi-level region-wise knowledge distillation and prototype balanced contrastive learning for class incremental semantic segmentation, Neurocomputing, 679 (2026), 133285. https://doi.org/10.1016/j.neucom.2026.133285 doi: 10.1016/j.neucom.2026.133285
    [36] A. Tarvainen, H. Valpola, Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, in Advances in Neural Information Processing Systems, 30 (2017), 1195–1204.
    [37] Y. Zhang, Y. Li, L. Qin, X. Liu, W. Deng, Leave no stone unturned: Mine extra knowledge for imbalanced facial expression recognition, in Advances in Neural Information Processing Systems, 36 (2023), 14414–14426. https://doi.org/10.52202/075280-0634
    [38] Y. Cui, M. Jia, T. Y. Lin, Y. Song, S. Belongie, Class-balanced loss based on effective number of samples, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2019), 9268–9277. https://doi.org/10.1109/CVPR.2019.00949
    [39] B. Zhou, Q. Cui, X. S. Wei, Z. M. Chen, BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2020), 9719–9728. https://doi.org/10.1109/CVPR42600.2020.00974
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(367) PDF downloads(30) Cited by(0)

Article outline

Figures and Tables

Figures(6)  /  Tables(8)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog