In-the-wild facial expression recognition (FER) is inherently affected by ambiguous expressions and subjective annotations, making the observed labels only partially reliable. Conventional one-hot supervision nevertheless treats these labels as ground truth and may drive deep networks to memorize corrupted annotations. This paper proposes RDAS, a reliability-driven adaptive supervision framework for noisy-label FER. The central idea was to use label reliability not merely as a sample weight, but as a variable that determines the supervision policy of each training sample. RDAS enhances expression-sensitive representations by combining RGB appearance with structural high-frequency cues, estimates label reliability with an exponential moving average (EMA) teacher through the margin between the annotated class and its strongest competing class, and assigns reliability-specific objectives to high-, medium-, and low-reliability samples. Reliable samples retained hard-label learning, ambiguous samples received teacher-guided soft correction, and unreliable samples were prevented from dominating training with misleading gradients. Experiments on RAF-DB, AffectNet, and FERPlus datasets showed that RDAS improved both standard recognition performance and robustness under synthetic label noise. On RAF-DB, RDAS achieved an average accuracy of $ 89.70\% $ on the original dataset and $ 86.15\% $ under 30% label noise. Ablation studies provided quantitative support for reliability modeling and adaptive supervision, while the visual analyses provided complementary qualitative evidence consistent with the proposed mechanism.
Citation: Xuefeng Zhao, Yixuan Dong. RDAS: Reliability-driven adaptive supervision for noisy-label facial expression recognition[J]. Electronic Research Archive, 2026, 34(10): 7274-7298. doi: 10.3934/era.2026315
In-the-wild facial expression recognition (FER) is inherently affected by ambiguous expressions and subjective annotations, making the observed labels only partially reliable. Conventional one-hot supervision nevertheless treats these labels as ground truth and may drive deep networks to memorize corrupted annotations. This paper proposes RDAS, a reliability-driven adaptive supervision framework for noisy-label FER. The central idea was to use label reliability not merely as a sample weight, but as a variable that determines the supervision policy of each training sample. RDAS enhances expression-sensitive representations by combining RGB appearance with structural high-frequency cues, estimates label reliability with an exponential moving average (EMA) teacher through the margin between the annotated class and its strongest competing class, and assigns reliability-specific objectives to high-, medium-, and low-reliability samples. Reliable samples retained hard-label learning, ambiguous samples received teacher-guided soft correction, and unreliable samples were prevented from dominating training with misleading gradients. Experiments on RAF-DB, AffectNet, and FERPlus datasets showed that RDAS improved both standard recognition performance and robustness under synthetic label noise. On RAF-DB, RDAS achieved an average accuracy of $ 89.70\% $ on the original dataset and $ 86.15\% $ under 30% label noise. Ablation studies provided quantitative support for reliability modeling and adaptive supervision, while the visual analyses provided complementary qualitative evidence consistent with the proposed mechanism.
| [1] | R. W. Picard, Affective Computing, MIT Press, 1997. https://doi.org/10.7551/mitpress/1140.001.0001 |
| [2] |
M. Pantic, L. J. M. Rothkrantz, Automatic analysis of facial expressions: The state of the art, IEEE Trans. Pattern Anal. Mach. Intell., 22 (2000), 1424–1445. https://doi.org/10.1109/34.895976 doi: 10.1109/34.895976
|
| [3] |
Z. Zeng, M. Pantic, G. I. Roisman, T. S. Huang, A survey of affect recognition methods: Audio, visual, and spontaneous expressions, IEEE Trans. Pattern Anal. Mach. Intell., 31 (2009), 39–58. https://doi.org/10.1109/TPAMI.2008.52 doi: 10.1109/TPAMI.2008.52
|
| [4] | I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, et al., Challenges in representation learning: A report on three machine learning contests, in Neural Information Processing, (2013), 117–124. https://doi.org/10.1007/978-3-642-42051-1_16 |
| [5] | E. Barsoum, C. Zhang, C. C. Ferrer, Z. Zhang, Training deep networks for facial expression recognition with crowd-sourced label distribution, in Proceedings of the ACM International Conference on Multimodal Interaction, (2016), 279–283. https://doi.org/10.1145/2993148.2993165 |
| [6] | S. Li, W. Deng, J. Du, Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2017), 2584–2593. https://doi.org/10.1109/CVPR.2017.277 |
| [7] |
A. Mollahosseini, B. Hasani, M. H. Mahoor, AffectNet: A database for facial expression, valence, and arousal computing in the wild, IEEE Trans. Affect. Comput., 10 (2019), 18–31. https://doi.org/10.1109/TAFFC.2017.2740923 doi: 10.1109/TAFFC.2017.2740923
|
| [8] | D. Arpit, S. Jastrzebski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, et al., A closer look at memorization in deep networks, in Proceedings of the International Conference on Machine Learning, (2017), 233–242. |
| [9] | K. Wang, X. Peng, J. Yang, S. Lu, Y. Qiao, Suppressing uncertainties for large-scale facial expression recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2020), 6897–6906. https://doi.org/10.1109/CVPR42600.2020.00693 |
| [10] | Y. Zhang, C. Wang, X. Ling, W. Deng, Relative uncertainty learning for facial expression recognition, in Advances in Neural Information Processing Systems, 34 (2021), 17616–17627. |
| [11] | Y. Zhang, C. Wang, W. Deng, Learn from all: Erasing attention consistency for noisy label facial expression recognition, in Proceedings of the European Conference on Computer Vision, (2022), 418–434. https://doi.org/10.1007/978-3-031-19809-0_24 |
| [12] |
Y. Tan, H. Xia, S. Song, Robust consistency learning for facial expression recognition under label noise, Vis. Comput., 41 (2025), 2655–2667. https://doi.org/10.1007/s00371-024-03558-1 doi: 10.1007/s00371-024-03558-1
|
| [13] |
H. Xia, C. Su, S. Song, Y. Tan, Dual-consistency constraints network for noisy facial expression recognition, Image Vis. Comput., 148 (2024), 105141. https://doi.org/10.1016/j.imavis.2024.105141 doi: 10.1016/j.imavis.2024.105141
|
| [14] |
X. Zhang, Y. Lu, H. Yan, J. Huang, Y. Gu, Y. Ji, et al., ReSup: Reliable label noise suppression for facial expression recognition, IEEE Trans. Affect. Comput., 16 (2025), 2006–2019. https://doi.org/10.1109/TAFFC.2025.3549017 doi: 10.1109/TAFFC.2025.3549017
|
| [15] | N. Le, K. Nguyen, Q. Tran, E. Tjiputra, B. Le, A. Nguyen, Uncertainty-aware label distribution learning for facial expression recognition, in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, (2023), 6088–6097. https://doi.org/10.1109/WACV56688.2023.00603 |
| [16] | J. Zheng, B. Li, S. Zhang, S. Wu, L. Cao, S. Ding, Attack can benefit: An adversarial approach to recognizing facial expressions under noisy annotations, in Proceedings of the AAAI Conference on Artificial Intelligence, 37 (2023), 3660–3668. https://doi.org/10.1609/aaai.v37i3.25477 |
| [17] |
Y. Li, H. Liu, D. Jiang, J. Liang, Facial expression recognition with label-noisy under dual-branch noise extraction and suppression, IEEE Trans. Affect. Comput., 16 (2025), 1514–1525. https://doi.org/10.1109/TAFFC.2024.3519359 doi: 10.1109/TAFFC.2024.3519359
|
| [18] |
J. Ye, D. Liu, C. Wang, H. Huang, L. Ying, L. Zhang, et al., CNLA: Collaborative noisy label adaptive learning for facial expression recognition, Inf. Sci., 719 (2025), 122436. https://doi.org/10.1016/j.ins.2025.122436 doi: 10.1016/j.ins.2025.122436
|
| [19] | J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2018), 7132–7141. https://doi.org/10.1109/CVPR.2018.00745 |
| [20] | S. Woo, J. Park, J. Y. Lee, I. S. Kweon, CBAM: Convolutional block attention module, in Proceedings of the European Conference on Computer Vision, (2018), 3–19. https://doi.org/10.1007/978-3-030-01234-2_1 |
| [21] |
K. Wang, X. Peng, J. Yang, D. Meng, Y. Qiao, Region attention networks for pose and occlusion robust facial expression recognition, IEEE Trans. Image Process., 29 (2020), 4057–4069. https://doi.org/10.1109/TIP.2019.2956143 doi: 10.1109/TIP.2019.2956143
|
| [22] |
Y. Fan, J. C. K. Lam, V. O. K. Li, Facial expression recognition with deeply-supervised attention network, IEEE Trans. Affective Comput., 13 (2022), 1057–1071. https://doi.org/10.1109/TAFFC.2020.2988264 doi: 10.1109/TAFFC.2020.2988264
|
| [23] | F. Xue, Q. Wang, G. Guo, TransFER: Learning relation-aware facial expression representations with transformers, in Proceedings of the IEEE/CVF International Conference on Computer Vision, (2021), 3601–3610. https://doi.org/10.1109/ICCV48922.2021.00358 |
| [24] |
Z. Wen, W. Lin, T. Wang, G. Xu, Distract your attention: Multi-head cross attention network for facial expression recognition, Biomimetics, 8 (2023), 199. https://doi.org/10.3390/biomimetics8020199 doi: 10.3390/biomimetics8020199
|
| [25] | C. Zheng, M. Mendieta, C. Chen, POSTER: A pyramid cross-fusion transformer network for facial expression recognition, in Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, (2023), 3146–3155. https://doi.org/10.1109/ICCVW60793.2023.00339 |
| [26] |
J. Mao, R. Xu, X. Yin, Y. Chang, B. Nie, A. Huang, POSTER++: A simpler and stronger facial expression recognition network, Pattern Recognit., 157 (2025), 110951. https://doi.org/10.1016/j.patcog.2024.110951 doi: 10.1016/j.patcog.2024.110951
|
| [27] |
F. Ma, B. Sun, S. Li, Transformer-augmented network with online label correction for facial expression recognition, IEEE Trans. Affective Comput., 15 (2024), 593–605. https://doi.org/10.1109/TAFFC.2023.3285231 doi: 10.1109/TAFFC.2023.3285231
|
| [28] | Z. Wu, J. Cui, LA-Net: Landmark-aware learning for reliable facial expression recognition under label noise, in Proceedings of the IEEE/CVF International Conference on Computer Vision, (2023), 20698–20707. https://doi.org/10.1109/ICCV51070.2023.01892 |
| [29] |
C. Y. Che, H. M. Sun, Y. X. Chen, S. Yang, R. S. Jia, NLFER: Multi-branch attention cross-fusion for robust facial expression recognition amidst noisy labels, Vis. Comput., 42 (2026), 1. https://doi.org/10.1007/s00371-025-04210-2 doi: 10.1007/s00371-025-04210-2
|
| [30] | J. She, Y. Hu, H. Shi, J. Wang, Q. Shen, T. Mei, Dive into ambiguity: Latent distribution mining and pairwise uncertainty estimation for facial expression recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2021), 6248–6257. https://doi.org/10.1109/CVPR46437.2021.00618 |
| [31] | J. Lee, Y. Choi, H. Kim, I. J. Kim, G. P. Nam, Navigating label ambiguity for facial expression recognition in the wild, in Proceedings of the AAAI Conference on Artificial Intelligence, 39 (2025), 4517–4525. https://doi.org/10.1609/aaai.v39i4.32476 |
| [32] |
D. Li, W. Xiong, T. Luo, L. Zhang, 3WAUS: A novel three-way adaptive uncertainty-suppressing model for facial expression recognition, Inf. Sci., 677 (2024), 120962. https://doi.org/10.1016/j.ins.2024.120962 doi: 10.1016/j.ins.2024.120962
|
| [33] |
Y. Apedo, H. Tao, A weakly supervised pavement crack segmentation based on adversarial learning and transformers, Multimedia Syst., 31 (2025), 266. https://doi.org/10.1007/s00530-025-01850-1 doi: 10.1007/s00530-025-01850-1
|
| [34] |
Y. Apedo, H. Tao, W. Gao, C. Xie, S. Zhao, Unsupervised domain adaptation for crack segmentation via cross-domain stylization and dual adversarial feature learning, J. Comput. Civ. Eng., 40 (2026), 04025146. https://doi.org/10.1061/JCCEE5.CPENG-7223 doi: 10.1061/JCCEE5.CPENG-7223
|
| [35] |
H. Zhou, H. Tao, Q. Zhang, C. Xie, B. Li, MRKD-PBCL: Multi-level region-wise knowledge distillation and prototype balanced contrastive learning for class incremental semantic segmentation, Neurocomputing, 679 (2026), 133285. https://doi.org/10.1016/j.neucom.2026.133285 doi: 10.1016/j.neucom.2026.133285
|
| [36] | A. Tarvainen, H. Valpola, Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, in Advances in Neural Information Processing Systems, 30 (2017), 1195–1204. |
| [37] | Y. Zhang, Y. Li, L. Qin, X. Liu, W. Deng, Leave no stone unturned: Mine extra knowledge for imbalanced facial expression recognition, in Advances in Neural Information Processing Systems, 36 (2023), 14414–14426. https://doi.org/10.52202/075280-0634 |
| [38] | Y. Cui, M. Jia, T. Y. Lin, Y. Song, S. Belongie, Class-balanced loss based on effective number of samples, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2019), 9268–9277. https://doi.org/10.1109/CVPR.2019.00949 |
| [39] | B. Zhou, Q. Cui, X. S. Wei, Z. M. Chen, BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2020), 9719–9728. https://doi.org/10.1109/CVPR42600.2020.00974 |