{"id":11677,"date":"2026-09-11T17:12:22","date_gmt":"2026-09-11T09:12:22","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=11677"},"modified":"2026-09-11T19:45:05","modified_gmt":"2026-09-11T11:45:05","slug":"jase-202612-35-025","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202612-35-025","title":{"rendered":"Lightweight Pyramid Network-Based Multi-Modal Feature Fusion for Crowd Counting"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=11162\" data-type=\"page\" data-id=\"11162\">Volume 35<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-09-11T17:12:22+08:00\">2026-09-11<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Han Shi, Rongkai Wang<a href=\"mailto:wangrk2024@126.com\"><i class=\"fa fa-envelope\"><\/i><\/a>, Deqing Li<a href=\"mailto:byoungholee@qq.com\"><i class=\"fa fa-envelope\"><\/i><\/a><\/p>\n\n\n\n<p style=\"font-size:14px\">College of Information and Electronic Technology, Jiamusi University, Jiamusi, 154007, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: May 31, 2026<br>Accepted: August 17, 2026<br>Publication Date: September 11, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/09\/35_025.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Proposed crowd counting structure <\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:  <a href=\"\/jase\/wp-content\/uploads\/2026\/09\/V35.0025.txt\" data-type=\"attachment\" data-id=\"11722\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202612_35.025\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202612_35.025<\/a>  <\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/09\/025_2026_1452_V35.pdf\" data-type=\"attachment\" data-id=\"11678\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Accurate crowd counting under complex scenes remains challenging due to scale variation, occlusion, and background noise. This paper introduces a lightweight pyramid network that fuses multi-modal features for robust crowd density estimation. The architecture integrates VGG-16, a multi-scale feature fusion module (MFFM) and a convolutional block attention module (CBAM) to suppress noise and enhance head-region focus. A dilated convolution module (DCM) further refines features into background attention and density maps, which are element-wise multiplied to produce the final prediction. A joint loss function combining Euclidean and adaptive Bayesian losses improves density map accuracy. Evaluated on ShanghaiTech, UCF-CC-50, and NWPU-Crowd datasets, the proposed method achieves state-of-the-art MAE and RMSE across sparse and dense scenes. Ablation studies confirm the effectiveness of MFFM, CBAM, and DCM modules. The model demonstrates strong generalization and robustness, even under low-light, occlusion, and background clutter.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;Crowd counting, Lightweight pyramid network, Multi-modal feature fusion, Dilated convolution module.<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_53180f64958e6cc2\" class=\"markdown markdown-main-panel md-content enable-luminous-fast-follows enable-updated-hr-color stronger\" dir=\"ltr\" aria-busy=\"false\" aria-live=\"polite\">\n<ol>\n<li data-path-to-node=\"0\">[1] G. Gao, J. Gao, Q. Liu, Q. Wang, and Y. Wang, (2025) &#8220;A survey of deep learning methods for density estimation and crowd counting&#8221; Vicinagearth 2(1): 2. DOI: https:\/\/doi.org\/10.1007\/s44336-024-00011-8.<\/li>\n<li data-path-to-node=\"0\">[2] Y. Yu, F. Zhu, J. Qian, H. Fujita, J. Yu, K. Zeng, and E. Chen, (2025) &#8220;CrowdFPN: crowd counting via scale-enhanced and location-aware feature pyramid network: Y. Yu et al.&#8221; Applied Intelligence 55(5): 359. DOI: https:\/\/doi.org\/10.1007\/s10489-025-06263-1.<\/li>\n<li data-path-to-node=\"0\">[3] H. Lin, X. Hong, Z. Ma, Y. Wang, and D. Meng, (2024) &#8220;Multidimensional measure matching for crowd counting&#8221; IEEE Transactions on Neural Networks and Learning Systems 36(5): 9112-9126. DOI: https:\/\/doi.org\/10.1109\/TNNLS.2024.3435854.<\/li>\n<li data-path-to-node=\"0\">[4] M. Ling, J. Chen, Y. Liu, W. Fang, and X. Geng, (2025) &#8220;Dual-branch adjacent connection and channel mixing network for video crowd counting&#8221; Pattern Recognition 167: 111709. DOI: https:\/\/doi.org\/10.1016\/j.patcog.2025.111709.<\/li>\n<li data-path-to-node=\"0\">[5] S. Yin, L. Wang, T. Chen, H. Huang, J. Gao, J. Zhang, M. Liu, P. Li, and C. Xu, (2026) &#8220;LKAFormer: A lightweight kolmogorov-arnold transformer model for image semantic segmentation&#8221; ACM Transactions on Intelligent Systems and Technology 17(3): 1-24. DOI: https:\/\/doi.org\/10.1145\/3759254.<\/li>\n<li data-path-to-node=\"0\">[6] Y. Hu, Y. Liu, G. Cao, and J. Wang, (2025) &#8220;CrowdCL: unsupervised crowd counting network via contrastive learning&#8221; IEEE Internet of Things Journal 12(12): 21704-21719. DOI: https:\/\/doi.org\/10.1109\/JIOT.2025.3547898.<\/li>\n<li data-path-to-node=\"0\">[7] Y. Qian, L. Zhang, Z. Guo, X. Hong, O. Arandjelovi\u0107, and C. R. Donovan, (2025) &#8220;Perspective-assisted prototype-based learning for semi-supervised crowd counting&#8221; Pattern Recognition 158: 111073. DOI: https:\/\/doi.org\/10.1016\/j.patcog.2024.111073.<\/li>\n<li data-path-to-node=\"0\">[8] Z. Zou, Y. Cheng, X. Qu, S. Ji, X. Guo, and P. Zhou, (2019) &#8220;Attend to count: Crowd counting with adaptive capacity multi-scale CNNs&#8221; Neurocomputing 367: 75-83. DOI: https:\/\/doi.org\/10.1016\/j.neucom.2019.08.009.<\/li>\n<li data-path-to-node=\"0\">[9] Y. Li, X. Zhang, and D. Chen, (2018) &#8220;Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes&#8221; Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition: 1091-1100. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00120.<\/li>\n<li data-path-to-node=\"0\">[10] T.-Y. Lin, P. Doll\u00e1r, R. Girshick, K. He, B. Hariharan, and S. Belongie, (2017) &#8220;Feature pyramid networks for object detection&#8221; Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR): 936-944. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.106.<\/li>\n<li data-path-to-node=\"0\">[11] P. Ding, H. Qian, Y. Zhou, S. Yan, S. Feng, and S. Yu, (2023) &#8220;Real-time efficient semantic segmentation network based on improved ASPP and parallel fusion module in complex scenes&#8221; Journal of Real-Time Image Processing 20(3): 41. DOI: https:\/\/doi.org\/10.1007\/s11554-023-01298-4.<\/li>\n<li data-path-to-node=\"0\">[12] S. Yin, H. Li, A. A. Laghari, T. R. Gadekallu, G. A. Sampedro, and A. Almadhor, (2024) &#8220;An anomaly detection model based on deep auto-encoder and capsule graph convolution via sparrow search algorithm in 6G Internet of Everything&#8221; IEEE Internet of Things Journal 11(18): 29402-29411. DOI: https:\/\/doi.org\/10.1109\/JIOT.2024.3353337.<\/li>\n<li data-path-to-node=\"0\">[13] L. Chen, H. Yao, J. Fu, and C. T. Ng, (2023) &#8220;The classification and localization of crack using lightweight convolutional neural network with CBAM&#8221; Engineering Structures 275: 115291. DOI: https:\/\/doi.org\/10.1016\/j.engstruct.2022.115291.<\/li>\n<li data-path-to-node=\"0\">[14] Y. Zhang, D. Zhou, S. Chen, S. Gao, and Y. Ma, (2016) &#8220;Single-image crowd counting via multi-column convolutional neural network&#8221; Proceedings of the IEEE conference on computer vision and pattern recognition: 589-597. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.70.<\/li>\n<li data-path-to-node=\"0\">[15] D. Babu Sam, S. Surya, and R. Venkatesh Babu, (2017) &#8220;Switching convolutional neural network for crowd counting&#8221; Proceedings of the IEEE conference on computer vision and pattern recognition: 5744-5752. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.429.<\/li>\n<li data-path-to-node=\"0\">[16] X. Tian and H. Hiraishi, (2025) &#8220;Design of crowd counting system based on improved CSRNet&#8221; Artificial Life and Robotics 30(1): 3-11. DOI: https:\/\/doi.org\/10.1007\/s10015-024-00993-0.<\/li>\n<li data-path-to-node=\"0\">[17] J. Yi, F. Chen, Z. Shen, Y. Xiang, S. Xiao, and W. Zhou, (2023) &#8220;An effective lightweight crowd counting method based on an encoder-decoder network for internet of video things&#8221; IEEE Internet of Things Journal 11(2): 3082-3094. DOI: https:\/\/doi.org\/10.1109\/JIOT.2023.3294727.<\/li>\n<li data-path-to-node=\"0\">[18] J. Xu, Z. Zhang, X. Li, W. Li, and K. Yu, (2024) &#8220;Attention Mixture Network for Crowd Counting via Binarization Transfer&#8221; Proceedings of the 2nd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice: 45-53. DOI: https:\/\/doi.org\/10.1145\/3688867.3690172.<\/li>\n<li data-path-to-node=\"0\">[19] L. Zhou, P. Wang, W. Li, J. Leng, and B. Lei, (2022) &#8220;Semantic-refined spatial pyramid network for crowd counting&#8221; Pattern Recognition Letters 159: 9-15. DOI: https:\/\/doi.org\/10.1016\/j.patrec.2022.04.029.<\/li>\n<li data-path-to-node=\"0\">[20] J. Yu and H. Hu, (2025) &#8220;Multiscale regional calibration network for crowd counting&#8221; Scientific Reports 15(1): 2866. DOI: https:\/\/doi.org\/10.1038\/s41598-025-86247-w.<\/li>\n<li data-path-to-node=\"0\">[21] K. Liu, Z. Dou, F. Wang, X. Xia, and J. Sang, (2025) &#8220;Cross-level attention multi-scale context-enhanced crowd counting network for transportation cyber-physical systems&#8221; IEEE Transactions on Intelligent Transportation Systems 26(9): 14250-14263. DOI: https:\/\/doi.org\/10.1109\/TITS.2025.3566718.<\/li>\n<li data-path-to-node=\"0\">[22] B. Yan, Y. Li, L. Dong, Z. Ren, H. Liu, X. Gao, and W. Cheng, (2025) &#8220;Crowd counting with WiFi sensing based on iterative attentional feature fusion&#8221; Computer Communications 241: 108245. DOI: https:\/\/doi.org\/10.1016\/j.comcom.2025.108245.<\/li>\n<li data-path-to-node=\"0\">[23] R. Ma, Y. Hou, C. Li, H. Jia, and X. Xie, (2025) &#8220;Scene-adaptive unsupervised crowd counting for video surveillance&#8221; IEEE Transactions on Circuits and Systems for Video Technology 35(7): 6910-6925. DOI: https:\/\/doi.org\/10.1109\/TCSVT.2025.3540850.<\/li>\n<\/ol>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1956,6],"tags":[2105],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation: BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202612_35.025 Download PDF Accurate crowd counting under complex scenes remains challenging due to scale&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/11677"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=11677"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=11677"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=11677"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}