Qingwu Shi, Ziteng Wang, and Han Shi
College of Information and Electronic Technology, Jiamusi University, Jiamusi, 154007, China
Received: Septempter 29, 2025
Accepted: July 04, 2026
Publication Date: August 12, 2026
Structure of the adaptive feature fusion network (AFFN)
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202611_34.043
Saliency object detection (SOD) aims to identify the most visually prominent regions in an image that attract humanattention. Despite significant advancements in deep learning-based SOD methods, existing approaches still face challenges in effectively integrating multi-scale, multi-semantic, and multi-spatial features, leading to incomplete saliency prediction, blurred boundaries, and poor generalization to complex scenes. To address these issues, this paper proposes a novel SOD framework that combines multiple-perspective feature fusion (MPFF) and deep adversarial networks (DAN), named MPFF-DAN. First, a multi-perspective feature extraction module is designed to capture complementary information from three critical perspectives: (1) the spatial perspective (low-level features with precise spatial localization); (2) the semantic perspective (high-level features with strong category-aware representation); (3) the scale perspective (multi-scale features to adapt to objects of varying sizes). Second, an adaptive feature fusion network (AFFN) is proposed to dynamically weight and aggregate the multi-perspective features, leveraging a dual-attention mechanism (channel attention + spatial attention) to enhance the discrimination of salient regions while suppressing background noise. Third, a deep adversarial network is integrated into the framework, where a generator (based on an improved U-Net) generates high quality saliency maps, and a discriminator (a multi-scale convolutional neural network) distinguishes between the generated saliency maps and ground-truth masks. The adversarial training paradigm drives the generator to produce more realistic and boundary-preserving saliency results. Extensive experiments are conducted on
five benchmark datasets using four evaluation metrics. Quantitative and qualitative results demonstrate that MPFF-DAN outperforms 15 state-of-the-art (SOTA) methods. It maintains high efficiency with a computational complexity.
Keywords: Saliency object detection; Feature fusion; Multiple perspectives; Adversarial networks; Attention mechanism
- [1] A. Borji, M.-M. Cheng, Q. Hou, H. Jiang, and J. Li, (2019) “Salient object detection: A survey” Computational visual media 5(2): 117–150. DOI: 10.1007/s41095-019-0149-9.
- [2] M. Ahmadi, N. Karimi, and S. Samavi, (2021) “Context-aware saliency detection for image retargeting using convolutional neural networks” Multimedia Tools and Applications 80(8): 11917–11941. DOI: 10.1007/s11042-020-10185-0.
- [3] J. Han, S. He, X. Qian, and D. Wang, (2013) “An Object-Oriented Visual Saliency Detection Framework Based on Sparse Coding Representations” IEEE Transactions on Circuits Systems for Video Technology 23(12): 2009–2021. DOI: 10.1109/TCSVT.2013.2242594.
- [4] N. Ding, C. Zhang, and A. Eskandarian, (2023) “SalienDet: A Saliency-based Feature Enhancement Algorithm for Object Detection for Autonomous Driving” IEEE Transactions on Intelligent Vehicles 9(1): 2624–2635. DOI: 10.1109/TIV.2023.3287359.
- [5] R. Han, X. Liu, and T. Chen. “Yolo-SG: Salience-guided detection of small objects in medical images”. In: 2022 IEEE International conference on image processing (ICIP). IEEE. 2022, 4218–4222. DOI: 10.1109/ICIP46576.2022.9898077.
- [6] G. Yuan, J. Song, and J. Li, (2025) “IF-USOD: Multi-modal information fusion interactive feature enhancement architecture for underwater salient object detection” Information Fusion 117: 102806. DOI: 10.1016/j.inffus.2024.102806.
- [7] B. Wang, M. Yang, P. Cao, and Y. Liu, (2025) “A novel embedded cross framework for high-resolution salient object detection: B. Wang et al.” Applied Intelligence 55(4): 277. DOI: 10.1007/s10489-024-06073-x.
- [8] R. Achanta, S. Hemami, F. Estrada, and S. Susstrunk. “Frequency-tuned salient region detection, 2009”. In: IEEE Conference on CVPR, 1597–1604. DOI: 10.1109/CVPR.2009.5206596.
- [9] T. Guo and X. Xu, (2021) “Salient object detection from low contrast images based on local contrast enhancing and non-local feature learning” The Visual Computer 37(8): 2069–2081. DOI: 10.1007/s00371-020-01964-9.
- [10] C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang. “Saliency detection via graph-based manifold ranking”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2013, 3166–3173. DOI: 10.1109/CVPR.2013.407.
- [11] M. W. Ahmed, A. Alazeb, N. Al Mudawi, T. Sadiq, B. Alabdullah, A. Algarni, A. Jalal, et al., (2025) “Perception of Natural Scenes: Objects Detection and Segmentations using Saliency Map with AlexNet.” International Arab Journal of Information Technology (IAJIT) 22(3): DOI: 10.34028/iajit/22/3/4.
- [12] Y. Zhang, (2025) “Image Denoising Based On Deep Feature Fusion And U-Net Network” 28(10): 2277–2285. DOI: 10.6180/jase.202510_28(10).0020.
- [13] Y. Tang, X. Wu, and W. Bu. “Deeply-supervised recurrent convolutional neural network for saliency detection”. In: Proceedings of the 24th ACM international conference on Multimedia. 2016, 397–401. DOI: 10.1145/2964284.2967250.
- [14] K. Alahmadi, S. Alharbi, and X. Wang, (2025) “Integrating dense layers with residual connections into transformers for enhanced sentiment classification” The Journal of Supercomputing 81(16): 1–28. DOI: 10.1007/s11227-025-07971-8.
- [15] X. Qin, Z. Zhang, C. Huang, C. Gao, M. Dehghan, and M. Jagersand. “Basnet: Boundary-aware salient object detection”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019, 7479–7489. DOI: 10.1109/CVPR.2019.00766.
- [16] M. Ma, C. Xia, and J. Li. “Pyramidal feature shrinking for salient object detection”. In: Proceedings of the AAAI conference on artificial intelligence. 35. 3. 2021, 2311–2318. DOI: 10.1609/aaai.v35i3.16331.
- [17] M.-M. Cheng, S.-H. Gao, A. Borji, Y.-Q. Tan, Z. Lin, and M. Wang, (2021) “A highly efficient model to study the semantics of salient object detection” IEEE Transactions on Pattern Analysis and Machine Intelligence 44(11): 8006–8021. DOI: 10.1109/TPAMI.2021.3107956.
- [18] R. Elakkiya, K. S. S. Teja, L. Jegatha Deborah, C. Bisogni, and C. Medaglia, (2022) “Imaging based cervical cancer diagnostics using small object detection-generative adversarial networks” Multimedia Tools and Applications 81(1): 191–207. DOI: 10.1007/s11042-021-10627-3.
- [19] H. Qiu and K. Zhang, (2025) “Multi-scale adaptive graph convolution-based thick cloud removal method for optical remote sensing images” International Journal of Information and Communication Technology 26(15): 57–77. DOI: 10.1504/IJICT.2025.146369.
- [20] W. Deng, L. Zheng, Q. Ye, G. Kang, Y. Yang, and J. Jiao. “Image-image domain adaptation with preserved self-similarity and domain-dissimilarity for person re-identification”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2018, 994–1003. DOI: 10.1109/CVPR.2018.00110.
- [21] J. Wang, M. Jung, S. Yin, and H. Li, (2025) “Adaptive Multi-Scale Gated Convolution and Context-Aware Attention Network for Accurate Small Object Detection [J]” International Journal of Computational Methods and Experimental Measurements 13(3): 576–587. DOI: 10.56578/ijcmem130308.
- [22] D.-P. Fan, C. Gong, Y. Cao, B. Ren, M.-M. Cheng, and A. Borji, (2018) “Enhanced-alignment measure for binary foreground map evaluation” arXiv preprint arXiv:1805.10421: DOI: 10.48550/arXiv.1805.10421.
