Real-time multi-person pose estimation suffers from inter-instance occlusion and severe scale variations. To resolve these bottlenecks, we present focal-dynamic context (FDC)-pose, an efficient single-stage architecture tailored for crowded scenes. The framework incorporates three sequential innovations: 1) a cascaded pyramid network-focal omni-dimensional dynamic convolution (CPN-FocalOD) backbone combining context modulation with omni-dimensional dynamic filtering to disentangle overlapping silhouettes; 2) a cross-scale fusion network featuring content-aware upsampling (DySample) to suppress quantization errors and preserve distal joint details; and 3) a minimum point distance intersection over union (MPDIoU) regression loss enforcing point-level geometric constraints to eliminate occlusion-induced localization jitter. Evaluations on the common objects in context (COCO) 2017 dataset demonstrate that FDC-pose yields absolute gains of 3.4% in AP, 3.0% in AP50, and 4.2% in the strict AP$ _{75} $ metric relative to the baseline. Parallel evaluations on the CrowdPose dataset show that our framework presents a prominent robustness gain, yielding an absolute improvement of 4.5% in AP.
Citation: Mengye Lin, Zhiming Cai, Yixin Zhang, Jiangchao Zhang, Zukun Xu, Huabin He. FDC-pose: Enhancing multi-person pose estimation in complex scenes via dynamic synergy mechanisms[J]. Electronic Research Archive, 2026, 34(9): 5861-5886. doi: 10.3934/era.2026260
Real-time multi-person pose estimation suffers from inter-instance occlusion and severe scale variations. To resolve these bottlenecks, we present focal-dynamic context (FDC)-pose, an efficient single-stage architecture tailored for crowded scenes. The framework incorporates three sequential innovations: 1) a cascaded pyramid network-focal omni-dimensional dynamic convolution (CPN-FocalOD) backbone combining context modulation with omni-dimensional dynamic filtering to disentangle overlapping silhouettes; 2) a cross-scale fusion network featuring content-aware upsampling (DySample) to suppress quantization errors and preserve distal joint details; and 3) a minimum point distance intersection over union (MPDIoU) regression loss enforcing point-level geometric constraints to eliminate occlusion-induced localization jitter. Evaluations on the common objects in context (COCO) 2017 dataset demonstrate that FDC-pose yields absolute gains of 3.4% in AP, 3.0% in AP50, and 4.2% in the strict AP$ _{75} $ metric relative to the baseline. Parallel evaluations on the CrowdPose dataset show that our framework presents a prominent robustness gain, yielding an absolute improvement of 4.5% in AP.
| [1] |
S. Ren, K. He, R. Girshick, J. Sun, Faster R-CNN: Towards real-time object detection with region proposal networks, IEEE Trans. Pattern Anal. Mach. Intell., 39 (2017), 1137–1149. https://doi.org/10.1109/TPAMI.2016.2577031 doi: 10.1109/TPAMI.2016.2577031
|
| [2] |
C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, et al., Deep learning-based human pose estimation: A survey, ACM Comput. Surv., 56 (2023), 1–37. https://doi.org/10.1145/3603618 doi: 10.1145/3603618
|
| [3] | M. Ma, X. He, X. Bai, A survey of deep learning-based human pose estimation: Methods, datasets, and evaluation metrics, in 2025 2nd International Conference on Digital Image Processing and Computer Applications (DIPCA), (2025), 146–152. https://doi.org/10.1109/DIPCA65051.2025.11042554 |
| [4] | Y. Chen, Z. Zhang, C. Yuan, B. Li, Y. Deng, W. Hu, Channel-wise topology refinement graph convolution for skeleton-based action recognition, preprint, arXiv:2107.12213. |
| [5] |
Y. Wang, P. Liu, H. Kang, D. Wu, D. Miao, ICFNet: Interactive-complementary fusion network for monocular 3D human pose estimation, Neurocomputing, 616 (2025), 128947. https://doi.org/10.1016/j.neucom.2024.128947 doi: 10.1016/j.neucom.2024.128947
|
| [6] | A. Newell, K. Yang, J. Deng, Stacked hourglass networks for human pose estimation, in Computer Vision-ECCV 2016, 9912 (2016), 483–499. https://doi.org/10.1007/978-3-319-46484-8_29 |
| [7] | B. Xiao, H. Wu, Y. Wei, Simple baselines for human pose estimation and tracking, in Computer Vision-ECCV 2018, 11210 (2018), 472–487. https://doi.org/10.1007/978-3-030-01231-1_29 |
| [8] | Y. Chen, Z. Wang, Y. Peng, Z. Zhang, G. Yu, J. Sun, Cascaded pyramid network for multi-person pose estimation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 7103–7112. https://doi.org/10.1109/CVPR.2018.00742 |
| [9] | K. Sun, B. Xiao, D. Liu, J. Wang, Deep high-resolution representation learning for human pose estimation, in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2019), 5686–5696. https://doi.org/10.1109/CVPR.2019.00584 |
| [10] | Y. Xu, J. Zhang, Q. Zhang, D. Tao, ViTPose: Simple vision transformer baselines for human pose estimation, preprint, arXiv:2204.12484. |
| [11] |
D. Yang, Y. Ge, N. Xu, R. Shi, DDCEFormer: Dual-domain cross enhanced transformer for 3D human pose estimation, Neurocomputing, 681 (2026), 133333. https://doi.org/10.1016/j.neucom.2026.133333 doi: 10.1016/j.neucom.2026.133333
|
| [12] | C. Yu, B. Xiao, C. Gao, L. Yuan, L. Zhang, N. Sang, et al., Lite-HRNet: A lightweight high-resolution network, in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2021), 10435–10445. https://doi.org/10.1109/CVPR46437.2021.01030 |
| [13] |
Y. Chen, Y. Tian, M. He, Monocular human pose estimation: A survey of deep learning-based methods, Comput. Vision Image Understanding, 192 (2020), 102897. https://doi.org/10.1016/j.cviu.2019.102897 doi: 10.1016/j.cviu.2019.102897
|
| [14] | Z. Geng, K. Sun, B. Xiao, Z. Zhang, J. Wang, Bottom-up human pose estimation via disentangled keypoint regression, preprint, arXiv:2104.02300. |
| [15] |
Z. Cao, G. Hidalgo, T. Simon, S. Wei, Y. Sheikh, OpenPose: Realtime multi-person 2D pose estimation using part affinity fields, IEEE Trans. Pattern Anal. Mach. Intell., 43 (2021), 172–186. https://doi.org/10.1109/TPAMI.2019.2929257 doi: 10.1109/TPAMI.2019.2929257
|
| [16] | C. Wang, A. Bochkovskiy, H. M. Liao, YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors, in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2023), 7464–7475. https://doi.org/10.1109/CVPR52729.2023.00721 |
| [17] | T. Jiang, P. Lu, L. Zhang, N. Ma, R. Han, C. Lyu, et al., RTMPose: Real-time multi-person pose estimation based on MMPose, preprint, arXiv:2303.07399. |
| [18] | Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, Z. Liu, Dynamic convolution: Attention over convolution kernels, in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2020), 11027–11036. https://doi.org/10.1109/CVPR42600.2020.01104 |
| [19] | C. Li, A. Zhou, A. Yao, Omni-dimensional dynamic convolution, preprint, arXiv:2209.07947. |
| [20] | J. Wang, K. Chen, R. Xu, Z. Liu, C. C. Loy, D. Lin, CARAFE: Content-aware reassembly of features, preprint, arXiv:1905.02188. |
| [21] | S. Liu, D. Huang, Y. Wang, Learning spatial fusion for single-shot object detection, preprint, arXiv:1911.09516. |
| [22] |
Z. Zheng, P. Wang, D. Ren, W. Liu, R. Ye, Q. Hu, et al., Enhancing geometric factors in model learning and inference for object detection and instance segmentation, IEEE Trans. Cybern., 52 (2022), 8574–8586. https://doi.org/10.1109/TCYB.2021.3095305 doi: 10.1109/TCYB.2021.3095305
|
| [23] |
G. Lan, Y. Wu, F. Hu, Q. Hao, Vision-based human pose estimation via deep learning: A survey, IEEE Trans. Hum.-Mach. Syst., 53 (2023), 253–268. https://doi.org/10.1109/THMS.2022.3219242 doi: 10.1109/THMS.2022.3219242
|
| [24] | D. Maji, S. Nagori, M. Mathew, D. Poddar, YOLO-Pose: Enhancing YOLO for multi person pose estimation using object keypoint similarity loss, preprint, arXiv:2204.06806. |
| [25] | J. Chen, X. Wang, Z. Guo, X. Zhang, J. Sun, Dynamic region-aware convolution, in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2021), 8060–8069. https://doi.org/10.1109/CVPR46437.2021.00797 |
| [26] | J. Yang, C. Li, X. Dai, L. Yuan, J. Gao, Focal modulation networks, preprint, arXiv:2203.11926. |
| [27] | S. Liu, L. Qi, H. Qin, J. Shi, J. Jia, Path aggregation network for instance segmentation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 8759–8768. https://doi.org/10.1109/CVPR.2018.00913 |
| [28] | S. Huang, Z. Lu, R. Cheng, C. He, FaPN: Feature-aligned pyramid network for dense image prediction, in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), (2021), 844–853. https://doi.org/10.1109/ICCV48922.2021.00090 |
| [29] | W. Liu, H. Lu, H. Fu, Z. Cao, Learning to upsample by learning to sample, in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), (2023), 6004–6014. https://doi.org/10.1109/ICCV51070.2023.00554 |
| [30] | H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, S. Savarese, Generalized intersection over union: A metric and a loss for bounding box regression, in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2019), 658–666. https://doi.org/10.1109/CVPR.2019.00075 |
| [31] | X. Li, W. Wang, L. Wu, S. Chen, X. Hu, J. Li, et al., Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection, preprint, arXiv:2006.04388. |
| [32] | S. Ma, Y. Xu, MPDIoU: A loss for efficient and accurate bounding box regression, preprint, arXiv:2307.07662. |
| [33] | J. Li, S. Bian, A. Zeng, C. Wang, B. Pang, W. Liu, et al., Human pose regression with residual log-likelihood estimation, preprint, arXiv:2107.11291. |
| [34] |
Y. Ma, L. Xu, Y. Zhang, T. Zhang, X. Luo, Steganalysis feature selection with multidimensional evaluation and dynamic threshold allocation, IEEE Trans. Circuits Syst. Video Technol., 34 (2024), 1954–1969. https://doi.org/10.1109/TCSVT.2023.3295364 doi: 10.1109/TCSVT.2023.3295364
|
| [35] | Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, S. Xie, A ConvNet for the 2020s, in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), 11966–11976. https://doi.org/10.1109/CVPR52688.2022.01167 |
| [36] | G. Huang, Y. Sun, Z. Liu, D. Sedra, K. Q. Weinberger, Deep networks with stochastic depth, in Computer Vision-ECCV 2016, 9908 (2016), 646–661. https://doi.org/10.1007/978-3-319-46493-0_39 |
| [37] |
Y. Ma, L. Xu, Q. Zhang, Y. Zhang, X. Xin, X. Luo, EIS-OBEA: Enhanced image steganalysis via opposition-based evolutionary algorithm, IEEE Trans. Inf. Forensics Secur., 20 (2025), 3616–3631. https://doi.org/10.1109/TIFS.2025.3549692 doi: 10.1109/TIFS.2025.3549692
|
| [38] | Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, et al., DETRs beat YOLOs on real-time object detection, preprint, arXiv:2304.08069. |
| [39] | R. Khanam, M. Hussain, YOLOv11: An overview of the key architectural enhancements, preprint, arXiv:2410.17725. |
| [40] |
H. Fang, J. Li, H. Tang, C. Xu, H. Zhu, Y. Xiu, et al., AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time, IEEE Trans. Pattern Anal. Mach. Intell., 45 (2023), 7157–7173. https://doi.org/10.1109/TPAMI.2022.3222784 doi: 10.1109/TPAMI.2022.3222784
|
| [41] | T. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, et al., Microsoft COCO: Common objects in context, preprint, arXiv:1405.0312. |
| [42] | J. Li, C. Wang, H. Zhu, Y. Mao, H. Fang, C. Lu, CrowdPose: Efficient crowded scenes pose estimation and a new benchmark, preprint, arXiv:1812.00324. |
| [43] | F. Yu, D. Wang, E. Shelhamer, T. Darrell, Deep layer aggregation, in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2018), 2403–2412. https://doi.org/10.1109/CVPR.2018.00255 |
| [44] |
C. Neff, A. Sheth, S. Furgurson, J. Middleton, H. Tabkhi, EfficientHRNet: Efficient and scalable high-resolution networks for real-time multi-person 2D human pose estimation, J. Real-Time Image Process., 18 (2021), 1037–1049. https://doi.org/10.1007/s11554-021-01132-9 doi: 10.1007/s11554-021-01132-9
|
| [45] | H. Wang, J. Liu, J. Tang, G. Wu, Lightweight super-resolution head for human pose estimation, in Proceedings of the 31st ACM International Conference on Multimedia, (2023), 2353–2361. https://doi.org/10.1145/3581783.3612236 |
| [46] |
J. Terven, D. Córdova-Esparza, J. Romero-González, A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS, Mach. Learn. Knowl. Extr., 5 (2023), 1680–1716. https://doi.org/10.3390/make5040083 doi: 10.3390/make5040083
|
| [47] | Y. Tang, K. Han, J. Guo, C. Xu, C. Xu, Y. Wang, GhostNetV2: Enhance cheap operation with long-range attention, preprint, arXiv:2211.12905. |
| [48] | S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, et al., ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders, in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2023), 16133–16142. https://doi.org/10.1109/CVPR52729.2023.01548 |
| [49] | S. Mehta, M. Rastegari, MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer, preprint, arXiv:2110.02178. |
| [50] | M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, et al., EdgeNeXt: Efficiently amalgamated CNN-transformer architecture for mobile vision applications, preprint, arXiv:2206.10589. |
| [51] | X. Ding, X. Zhang, J. Han, G. Ding, Scaling up your kernels to $31\times31$: Revisiting large kernel design in CNNs, in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), (2022), 11953–11965. https://doi.org/10.1109/CVPR52688.2022.01166 |
| [52] | J. He, S. Erfani, X. Ma, J. Bailey, Y. Chi, X. Hua, Alpha-IoU: A family of power intersection over union losses for bounding box regression, preprint, arXiv:2110.13675. |
| [53] | H. Zhang, S. Zhang, Shape-IoU: More accurate metric considering bounding box shape and scale, preprint, arXiv:2312.17663. |
| [54] | Z. Gevorgyan, SIoU loss: More powerful learning for bounding box regression, preprint, arXiv:2205.12740. |
| [55] | Y. Zhang, W. Ren, Z. Zhang, Z. Jia, L. Wang, T. Tan, Focal and efficient IOU loss for accurate bounding box regression, preprint, arXiv:2101.08158. |
| [56] | Z. Tong, Y. Chen, Z. Xu, R. Yu, Wise-IoU: Bounding box regression loss with dynamic focusing mechanism, preprint, arXiv:2301.10051. |