Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Lightweight Pyramid Network-Based Multi-Modal Feature Fusion for Crowd Counting

Han Shi, Rongkai Wang, Deqing Li

College of Information and Electronic Technology, Jiamusi University, Jiamusi, 154007, China

Received: May 31, 2026
Accepted: August 17, 2026
Publication Date: September 11, 2026

上傳圖片

Proposed crowd counting structure

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202612_35.025

Download PDF

Accurate crowd counting under complex scenes remains challenging due to scale variation, occlusion, and background noise. This paper introduces a lightweight pyramid network that fuses multi-modal features for robust crowd density estimation. The architecture integrates VGG-16, a multi-scale feature fusion module (MFFM) and a convolutional block attention module (CBAM) to suppress noise and enhance head-region focus. A dilated convolution module (DCM) further refines features into background attention and density maps, which are element-wise multiplied to produce the final prediction. A joint loss function combining Euclidean and adaptive Bayesian losses improves density map accuracy. Evaluated on ShanghaiTech, UCF-CC-50, and NWPU-Crowd datasets, the proposed method achieves state-of-the-art MAE and RMSE across sparse and dense scenes. Ablation studies confirm the effectiveness of MFFM, CBAM, and DCM modules. The model demonstrates strong generalization and robustness, even under low-light, occlusion, and background clutter.

Keywords: Crowd counting, Lightweight pyramid network, Multi-modal feature fusion, Dilated convolution module.

  1. [1] G. Gao, J. Gao, Q. Liu, Q. Wang, and Y. Wang, (2025) “A survey of deep learning methods for density estimation and crowd counting” Vicinagearth 2(1): 2. DOI: https://doi.org/10.1007/s44336-024-00011-8.
  2. [2] Y. Yu, F. Zhu, J. Qian, H. Fujita, J. Yu, K. Zeng, and E. Chen, (2025) “CrowdFPN: crowd counting via scale-enhanced and location-aware feature pyramid network: Y. Yu et al.” Applied Intelligence 55(5): 359. DOI: https://doi.org/10.1007/s10489-025-06263-1.
  3. [3] H. Lin, X. Hong, Z. Ma, Y. Wang, and D. Meng, (2024) “Multidimensional measure matching for crowd counting” IEEE Transactions on Neural Networks and Learning Systems 36(5): 9112-9126. DOI: https://doi.org/10.1109/TNNLS.2024.3435854.
  4. [4] M. Ling, J. Chen, Y. Liu, W. Fang, and X. Geng, (2025) “Dual-branch adjacent connection and channel mixing network for video crowd counting” Pattern Recognition 167: 111709. DOI: https://doi.org/10.1016/j.patcog.2025.111709.
  5. [5] S. Yin, L. Wang, T. Chen, H. Huang, J. Gao, J. Zhang, M. Liu, P. Li, and C. Xu, (2026) “LKAFormer: A lightweight kolmogorov-arnold transformer model for image semantic segmentation” ACM Transactions on Intelligent Systems and Technology 17(3): 1-24. DOI: https://doi.org/10.1145/3759254.
  6. [6] Y. Hu, Y. Liu, G. Cao, and J. Wang, (2025) “CrowdCL: unsupervised crowd counting network via contrastive learning” IEEE Internet of Things Journal 12(12): 21704-21719. DOI: https://doi.org/10.1109/JIOT.2025.3547898.
  7. [7] Y. Qian, L. Zhang, Z. Guo, X. Hong, O. Arandjelović, and C. R. Donovan, (2025) “Perspective-assisted prototype-based learning for semi-supervised crowd counting” Pattern Recognition 158: 111073. DOI: https://doi.org/10.1016/j.patcog.2024.111073.
  8. [8] Z. Zou, Y. Cheng, X. Qu, S. Ji, X. Guo, and P. Zhou, (2019) “Attend to count: Crowd counting with adaptive capacity multi-scale CNNs” Neurocomputing 367: 75-83. DOI: https://doi.org/10.1016/j.neucom.2019.08.009.
  9. [9] Y. Li, X. Zhang, and D. Chen, (2018) “Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition: 1091-1100. DOI: https://doi.org/10.1109/CVPR.2018.00120.
  10. [10] T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, (2017) “Feature pyramid networks for object detection” Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR): 936-944. DOI: https://doi.org/10.1109/CVPR.2017.106.
  11. [11] P. Ding, H. Qian, Y. Zhou, S. Yan, S. Feng, and S. Yu, (2023) “Real-time efficient semantic segmentation network based on improved ASPP and parallel fusion module in complex scenes” Journal of Real-Time Image Processing 20(3): 41. DOI: https://doi.org/10.1007/s11554-023-01298-4.
  12. [12] S. Yin, H. Li, A. A. Laghari, T. R. Gadekallu, G. A. Sampedro, and A. Almadhor, (2024) “An anomaly detection model based on deep auto-encoder and capsule graph convolution via sparrow search algorithm in 6G Internet of Everything” IEEE Internet of Things Journal 11(18): 29402-29411. DOI: https://doi.org/10.1109/JIOT.2024.3353337.
  13. [13] L. Chen, H. Yao, J. Fu, and C. T. Ng, (2023) “The classification and localization of crack using lightweight convolutional neural network with CBAM” Engineering Structures 275: 115291. DOI: https://doi.org/10.1016/j.engstruct.2022.115291.
  14. [14] Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, (2016) “Single-image crowd counting via multi-column convolutional neural network” Proceedings of the IEEE conference on computer vision and pattern recognition: 589-597. DOI: https://doi.org/10.1109/CVPR.2016.70.
  15. [15] D. Babu Sam, S. Surya, and R. Venkatesh Babu, (2017) “Switching convolutional neural network for crowd counting” Proceedings of the IEEE conference on computer vision and pattern recognition: 5744-5752. DOI: https://doi.org/10.1109/CVPR.2017.429.
  16. [16] X. Tian and H. Hiraishi, (2025) “Design of crowd counting system based on improved CSRNet” Artificial Life and Robotics 30(1): 3-11. DOI: https://doi.org/10.1007/s10015-024-00993-0.
  17. [17] J. Yi, F. Chen, Z. Shen, Y. Xiang, S. Xiao, and W. Zhou, (2023) “An effective lightweight crowd counting method based on an encoder-decoder network for internet of video things” IEEE Internet of Things Journal 11(2): 3082-3094. DOI: https://doi.org/10.1109/JIOT.2023.3294727.
  18. [18] J. Xu, Z. Zhang, X. Li, W. Li, and K. Yu, (2024) “Attention Mixture Network for Crowd Counting via Binarization Transfer” Proceedings of the 2nd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice: 45-53. DOI: https://doi.org/10.1145/3688867.3690172.
  19. [19] L. Zhou, P. Wang, W. Li, J. Leng, and B. Lei, (2022) “Semantic-refined spatial pyramid network for crowd counting” Pattern Recognition Letters 159: 9-15. DOI: https://doi.org/10.1016/j.patrec.2022.04.029.
  20. [20] J. Yu and H. Hu, (2025) “Multiscale regional calibration network for crowd counting” Scientific Reports 15(1): 2866. DOI: https://doi.org/10.1038/s41598-025-86247-w.
  21. [21] K. Liu, Z. Dou, F. Wang, X. Xia, and J. Sang, (2025) “Cross-level attention multi-scale context-enhanced crowd counting network for transportation cyber-physical systems” IEEE Transactions on Intelligent Transportation Systems 26(9): 14250-14263. DOI: https://doi.org/10.1109/TITS.2025.3566718.
  22. [22] B. Yan, Y. Li, L. Dong, Z. Ren, H. Liu, X. Gao, and W. Cheng, (2025) “Crowd counting with WiFi sensing based on iterative attentional feature fusion” Computer Communications 241: 108245. DOI: https://doi.org/10.1016/j.comcom.2025.108245.
  23. [23] R. Ma, Y. Hou, C. Li, H. Jia, and X. Xie, (2025) “Scene-adaptive unsupervised crowd counting for video surveillance” IEEE Transactions on Circuits and Systems for Video Technology 35(7): 6910-6925. DOI: https://doi.org/10.1109/TCSVT.2025.3540850.