Research article Special Issues

FDC-pose: Enhancing multi-person pose estimation in complex scenes via dynamic synergy mechanisms

  • Published: 09 July 2026
  • Real-time multi-person pose estimation suffers from inter-instance occlusion and severe scale variations. To resolve these bottlenecks, we present focal-dynamic context (FDC)-pose, an efficient single-stage architecture tailored for crowded scenes. The framework incorporates three sequential innovations: 1) a cascaded pyramid network-focal omni-dimensional dynamic convolution (CPN-FocalOD) backbone combining context modulation with omni-dimensional dynamic filtering to disentangle overlapping silhouettes; 2) a cross-scale fusion network featuring content-aware upsampling (DySample) to suppress quantization errors and preserve distal joint details; and 3) a minimum point distance intersection over union (MPDIoU) regression loss enforcing point-level geometric constraints to eliminate occlusion-induced localization jitter. Evaluations on the common objects in context (COCO) 2017 dataset demonstrate that FDC-pose yields absolute gains of 3.4% in AP, 3.0% in AP50, and 4.2% in the strict AP$ _{75} $ metric relative to the baseline. Parallel evaluations on the CrowdPose dataset show that our framework presents a prominent robustness gain, yielding an absolute improvement of 4.5% in AP.

    Citation: Mengye Lin, Zhiming Cai, Yixin Zhang, Jiangchao Zhang, Zukun Xu, Huabin He. FDC-pose: Enhancing multi-person pose estimation in complex scenes via dynamic synergy mechanisms[J]. Electronic Research Archive, 2026, 34(9): 5861-5886. doi: 10.3934/era.2026260

    Related Papers:

  • Real-time multi-person pose estimation suffers from inter-instance occlusion and severe scale variations. To resolve these bottlenecks, we present focal-dynamic context (FDC)-pose, an efficient single-stage architecture tailored for crowded scenes. The framework incorporates three sequential innovations: 1) a cascaded pyramid network-focal omni-dimensional dynamic convolution (CPN-FocalOD) backbone combining context modulation with omni-dimensional dynamic filtering to disentangle overlapping silhouettes; 2) a cross-scale fusion network featuring content-aware upsampling (DySample) to suppress quantization errors and preserve distal joint details; and 3) a minimum point distance intersection over union (MPDIoU) regression loss enforcing point-level geometric constraints to eliminate occlusion-induced localization jitter. Evaluations on the common objects in context (COCO) 2017 dataset demonstrate that FDC-pose yields absolute gains of 3.4% in AP, 3.0% in AP50, and 4.2% in the strict AP$ _{75} $ metric relative to the baseline. Parallel evaluations on the CrowdPose dataset show that our framework presents a prominent robustness gain, yielding an absolute improvement of 4.5% in AP.



    加载中


    [1] S. Ren, K. He, R. Girshick, J. Sun, Faster R-CNN: Towards real-time object detection with region proposal networks, IEEE Trans. Pattern Anal. Mach. Intell., 39 (2017), 1137–1149. https://doi.org/10.1109/TPAMI.2016.2577031 doi: 10.1109/TPAMI.2016.2577031
    [2] C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, et al., Deep learning-based human pose estimation: A survey, ACM Comput. Surv., 56 (2023), 1–37. https://doi.org/10.1145/3603618 doi: 10.1145/3603618
    [3] M. Ma, X. He, X. Bai, A survey of deep learning-based human pose estimation: Methods, datasets, and evaluation metrics, in 2025 2nd International Conference on Digital Image Processing and Computer Applications (DIPCA), (2025), 146–152. https://doi.org/10.1109/DIPCA65051.2025.11042554
    [4] Y. Chen, Z. Zhang, C. Yuan, B. Li, Y. Deng, W. Hu, Channel-wise topology refinement graph convolution for skeleton-based action recognition, preprint, arXiv:2107.12213.
    [5] Y. Wang, P. Liu, H. Kang, D. Wu, D. Miao, ICFNet: Interactive-complementary fusion network for monocular 3D human pose estimation, Neurocomputing, 616 (2025), 128947. https://doi.org/10.1016/j.neucom.2024.128947 doi: 10.1016/j.neucom.2024.128947
    [6] A. Newell, K. Yang, J. Deng, Stacked hourglass networks for human pose estimation, in Computer Vision-ECCV 2016, 9912 (2016), 483–499. https://doi.org/10.1007/978-3-319-46484-8_29
    [7] B. Xiao, H. Wu, Y. Wei, Simple baselines for human pose estimation and tracking, in Computer Vision-ECCV 2018, 11210 (2018), 472–487. https://doi.org/10.1007/978-3-030-01231-1_29
    [8] Y. Chen, Z. Wang, Y. Peng, Z. Zhang, G. Yu, J. Sun, Cascaded pyramid network for multi-person pose estimation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 7103–7112. https://doi.org/10.1109/CVPR.2018.00742
    [9] K. Sun, B. Xiao, D. Liu, J. Wang, Deep high-resolution representation learning for human pose estimation, in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2019), 5686–5696. https://doi.org/10.1109/CVPR.2019.00584
    [10] Y. Xu, J. Zhang, Q. Zhang, D. Tao, ViTPose: Simple vision transformer baselines for human pose estimation, preprint, arXiv:2204.12484.
    [11] D. Yang, Y. Ge, N. Xu, R. Shi, DDCEFormer: Dual-domain cross enhanced transformer for 3D human pose estimation, Neurocomputing, 681 (2026), 133333. https://doi.org/10.1016/j.neucom.2026.133333 doi: 10.1016/j.neucom.2026.133333
    [12] C. Yu, B. Xiao, C. Gao, L. Yuan, L. Zhang, N. Sang, et al., Lite-HRNet: A lightweight high-resolution network, in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2021), 10435–10445. https://doi.org/10.1109/CVPR46437.2021.01030
    [13] Y. Chen, Y. Tian, M. He, Monocular human pose estimation: A survey of deep learning-based methods, Comput. Vision Image Understanding, 192 (2020), 102897. https://doi.org/10.1016/j.cviu.2019.102897 doi: 10.1016/j.cviu.2019.102897
    [14] Z. Geng, K. Sun, B. Xiao, Z. Zhang, J. Wang, Bottom-up human pose estimation via disentangled keypoint regression, preprint, arXiv:2104.02300.
    [15] Z. Cao, G. Hidalgo, T. Simon, S. Wei, Y. Sheikh, OpenPose: Realtime multi-person 2D pose estimation using part affinity fields, IEEE Trans. Pattern Anal. Mach. Intell., 43 (2021), 172–186. https://doi.org/10.1109/TPAMI.2019.2929257 doi: 10.1109/TPAMI.2019.2929257
    [16] C. Wang, A. Bochkovskiy, H. M. Liao, YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2023), 7464–7475. https://doi.org/10.1109/CVPR52729.2023.00721
    [17] T. Jiang, P. Lu, L. Zhang, N. Ma, R. Han, C. Lyu, et al., RTMPose: Real-time multi-person pose estimation based on MMPose, preprint, arXiv:2303.07399.
    [18] Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, Z. Liu, Dynamic convolution: Attention over convolution kernels, in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2020), 11027–11036. https://doi.org/10.1109/CVPR42600.2020.01104
    [19] C. Li, A. Zhou, A. Yao, Omni-dimensional dynamic convolution, preprint, arXiv:2209.07947.
    [20] J. Wang, K. Chen, R. Xu, Z. Liu, C. C. Loy, D. Lin, CARAFE: Content-aware reassembly of features, preprint, arXiv:1905.02188.
    [21] S. Liu, D. Huang, Y. Wang, Learning spatial fusion for single-shot object detection, preprint, arXiv:1911.09516.
    [22] Z. Zheng, P. Wang, D. Ren, W. Liu, R. Ye, Q. Hu, et al., Enhancing geometric factors in model learning and inference for object detection and instance segmentation, IEEE Trans. Cybern., 52 (2022), 8574–8586. https://doi.org/10.1109/TCYB.2021.3095305 doi: 10.1109/TCYB.2021.3095305
    [23] G. Lan, Y. Wu, F. Hu, Q. Hao, Vision-based human pose estimation via deep learning: A survey, IEEE Trans. Hum.-Mach. Syst., 53 (2023), 253–268. https://doi.org/10.1109/THMS.2022.3219242 doi: 10.1109/THMS.2022.3219242
    [24] D. Maji, S. Nagori, M. Mathew, D. Poddar, YOLO-Pose: Enhancing YOLO for multi person pose estimation using object keypoint similarity loss, preprint, arXiv:2204.06806.
    [25] J. Chen, X. Wang, Z. Guo, X. Zhang, J. Sun, Dynamic region-aware convolution, in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2021), 8060–8069. https://doi.org/10.1109/CVPR46437.2021.00797
    [26] J. Yang, C. Li, X. Dai, L. Yuan, J. Gao, Focal modulation networks, preprint, arXiv:2203.11926.
    [27] S. Liu, L. Qi, H. Qin, J. Shi, J. Jia, Path aggregation network for instance segmentation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 8759–8768. https://doi.org/10.1109/CVPR.2018.00913
    [28] S. Huang, Z. Lu, R. Cheng, C. He, FaPN: Feature-aligned pyramid network for dense image prediction, in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), (2021), 844–853. https://doi.org/10.1109/ICCV48922.2021.00090
    [29] W. Liu, H. Lu, H. Fu, Z. Cao, Learning to upsample by learning to sample, in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), (2023), 6004–6014. https://doi.org/10.1109/ICCV51070.2023.00554
    [30] H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, S. Savarese, Generalized intersection over union: A metric and a loss for bounding box regression, in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2019), 658–666. https://doi.org/10.1109/CVPR.2019.00075
    [31] X. Li, W. Wang, L. Wu, S. Chen, X. Hu, J. Li, et al., Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection, preprint, arXiv:2006.04388.
    [32] S. Ma, Y. Xu, MPDIoU: A loss for efficient and accurate bounding box regression, preprint, arXiv:2307.07662.
    [33] J. Li, S. Bian, A. Zeng, C. Wang, B. Pang, W. Liu, et al., Human pose regression with residual log-likelihood estimation, preprint, arXiv:2107.11291.
    [34] Y. Ma, L. Xu, Y. Zhang, T. Zhang, X. Luo, Steganalysis feature selection with multidimensional evaluation and dynamic threshold allocation, IEEE Trans. Circuits Syst. Video Technol., 34 (2024), 1954–1969. https://doi.org/10.1109/TCSVT.2023.3295364 doi: 10.1109/TCSVT.2023.3295364
    [35] Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, S. Xie, A ConvNet for the 2020s, in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), 11966–11976. https://doi.org/10.1109/CVPR52688.2022.01167
    [36] G. Huang, Y. Sun, Z. Liu, D. Sedra, K. Q. Weinberger, Deep networks with stochastic depth, in Computer Vision-ECCV 2016, 9908 (2016), 646–661. https://doi.org/10.1007/978-3-319-46493-0_39
    [37] Y. Ma, L. Xu, Q. Zhang, Y. Zhang, X. Xin, X. Luo, EIS-OBEA: Enhanced image steganalysis via opposition-based evolutionary algorithm, IEEE Trans. Inf. Forensics Secur., 20 (2025), 3616–3631. https://doi.org/10.1109/TIFS.2025.3549692 doi: 10.1109/TIFS.2025.3549692
    [38] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, et al., DETRs beat YOLOs on real-time object detection, preprint, arXiv:2304.08069.
    [39] R. Khanam, M. Hussain, YOLOv11: An overview of the key architectural enhancements, preprint, arXiv:2410.17725.
    [40] H. Fang, J. Li, H. Tang, C. Xu, H. Zhu, Y. Xiu, et al., AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time, IEEE Trans. Pattern Anal. Mach. Intell., 45 (2023), 7157–7173. https://doi.org/10.1109/TPAMI.2022.3222784 doi: 10.1109/TPAMI.2022.3222784
    [41] T. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, et al., Microsoft COCO: Common objects in context, preprint, arXiv:1405.0312.
    [42] J. Li, C. Wang, H. Zhu, Y. Mao, H. Fang, C. Lu, CrowdPose: Efficient crowded scenes pose estimation and a new benchmark, preprint, arXiv:1812.00324.
    [43] F. Yu, D. Wang, E. Shelhamer, T. Darrell, Deep layer aggregation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 2403–2412. https://doi.org/10.1109/CVPR.2018.00255
    [44] C. Neff, A. Sheth, S. Furgurson, J. Middleton, H. Tabkhi, EfficientHRNet: Efficient and scalable high-resolution networks for real-time multi-person 2D human pose estimation, J. Real-Time Image Process., 18 (2021), 1037–1049. https://doi.org/10.1007/s11554-021-01132-9 doi: 10.1007/s11554-021-01132-9
    [45] H. Wang, J. Liu, J. Tang, G. Wu, Lightweight super-resolution head for human pose estimation, in Proceedings of the 31st ACM International Conference on Multimedia, (2023), 2353–2361. https://doi.org/10.1145/3581783.3612236
    [46] J. Terven, D. Córdova-Esparza, J. Romero-González, A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS, Mach. Learn. Knowl. Extr., 5 (2023), 1680–1716. https://doi.org/10.3390/make5040083 doi: 10.3390/make5040083
    [47] Y. Tang, K. Han, J. Guo, C. Xu, C. Xu, Y. Wang, GhostNetV2: Enhance cheap operation with long-range attention, preprint, arXiv:2211.12905.
    [48] S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, et al., ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders, in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2023), 16133–16142. https://doi.org/10.1109/CVPR52729.2023.01548
    [49] S. Mehta, M. Rastegari, MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer, preprint, arXiv:2110.02178.
    [50] M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, et al., EdgeNeXt: Efficiently amalgamated CNN-transformer architecture for mobile vision applications, preprint, arXiv:2206.10589.
    [51] X. Ding, X. Zhang, J. Han, G. Ding, Scaling up your kernels to $31\times31$: Revisiting large kernel design in CNNs, in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), 11953–11965. https://doi.org/10.1109/CVPR52688.2022.01166
    [52] J. He, S. Erfani, X. Ma, J. Bailey, Y. Chi, X. Hua, Alpha-IoU: A family of power intersection over union losses for bounding box regression, preprint, arXiv:2110.13675.
    [53] H. Zhang, S. Zhang, Shape-IoU: More accurate metric considering bounding box shape and scale, preprint, arXiv:2312.17663.
    [54] Z. Gevorgyan, SIoU loss: More powerful learning for bounding box regression, preprint, arXiv:2205.12740.
    [55] Y. Zhang, W. Ren, Z. Zhang, Z. Jia, L. Wang, T. Tan, Focal and efficient IOU loss for accurate bounding box regression, preprint, arXiv:2101.08158.
    [56] Z. Tong, Y. Chen, Z. Xu, R. Yu, Wise-IoU: Bounding box regression loss with dynamic focusing mechanism, preprint, arXiv:2301.10051.
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(432) PDF downloads(40) Cited by(0)

Article outline

Figures and Tables

Figures(9)  /  Tables(6)

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog