Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

A Bi-Mamba Fusion Framework with KAN Pattern Mining for Multi-modal Sensitive Content Detection

Ming Lian and Teng Li

School of Artificial Intelligence, Anhui University, Hefei 230039 China

Received: May 7, 2026
Accepted: June 4, 2026
Publication Date: July 3, 2026

上傳圖片

The illustration of MamKAN. It consists of a multi-modal feature extraction with textual inversion, a bidirectional Mamba fusion, and an adaptive pattern mining 

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202610_33.030  

Download PDF

The meteoric expansion of social media platforms has established multi-modal data as the dominant paradigm for global digital communication. However, this evolution has concurrently facilitated the rapid dissemination of sensitive and harmful material, including hate speech, extremist propaganda, and violent imagery, which presents a severe challenge to online security and public psychological well-being. Modern automated detection systems frequently falter because they fail to bridge the heterogeneous semantic gap between visual and linguistic modalities effectively, often relying on late-stage fusion or computationally expensive Transformer architectures that struggle with long-range dependencies. Furthermore, the conventional use of multi-layer perceptrons (MLPs) as predictors often fails to capture the irregular and complex decision boundaries inherent in high-dimensional multi-modal manifolds. To address these critical limitations, we propose MamKAN, a sophisticated multi-modal fusion framework that integrates State Space Models with spline-based neural architectures. Our methodology innovates at three distinct stages: First, we implement an implicit modal interaction strategy using textual inversion, which maps visual features into pseudo-word tokens within the text embedding space to achieve robust early-stage semantic alignment. Second, we introduce a bidirectional Mamba fusion module based on state space models; this mechanism achieves linear computational complexity while simultaneously capturing comprehensive vision-to-text and text-to-vision contextual dependencies through forward and backward scanning. Finally, we replace traditional MLPs with a Kolmogorov-Arnold Network (KAN) predictor. By employing learnable B-spline activation functions on the network edges, the KAN module adaptively fits complex non-linear distributions, significantly enhancing discriminative precision. Comprehensive experiments conducted on the HMC and HarMeme benchmark datasets demonstrate that
MamKAN consistently and significantly outperforms state-of-the-art models across three metrics.

Keywords: State space models; Kolmogorov-Arnold network; Sensitive content detection

  1. [1] D. Povedano Álvarez, A. L. Sandoval Orozco, J. P. García-Miguel, and L. J. García Villalba, (2023) “Learning strategies for sensitive content detection” Electronics 12(11): 2496. DOI: 10.3390/electronics12112496.
  2. [2] X. Meng and Y. Xu, (2019) “Research on sensitive content detection in social networks” CCF Transactions on Networking 2(2): 126–135. DOI: 10.1007/s42045-019-00021-x.
  3. [3] J. Gao, M. Liu, P. Li, J. Zhang, and Z. Chen, (2024) “Deep Multiview Adaptive Clustering With Semantic Invariance” IEEE Transactions on Neural Networks and Learning Systems 35(9): 12965–12978. DOI: 10.1109/TNNLS.2023.3265699.
  4. [4] J. Gao, M. Liu, P. Li, A. A. Laghari, A. R. Javed, N. Victor, and T. R. Gadekallu, (2024) “Deep Incomplete Multiview Clustering via Information Bottleneck for Pattern Mining of Data in Extreme-Environment IoT” IEEE Internet of Things Journal 11(1): 26700–26712. DOI: 10.1109/JIOT.2023.3325272.
  5. [5] T. Markov, C. Zhang, S. Agarwal, F. E. Nekoul, T. Lee, S. Adler, A. Jiang, and L. Weng. “A holistic approach to undesired content detection in the real world”. In: Proceedings of the AAAI conference on artificial intelligence. 37. 12. 2023, 15009–15018. DOI: 10.1609/aaai.v37i12.26752.
  6. [6] G. Burbi, A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo. “Mapping memes to words for multimodal hateful meme classification”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, 2832–2836.
  7. [7] F. Wu, B. Gao, X. Pan, L. Li, Y. Ma, S. Liu, and Z. Liu, (2024) “Fuser: An enhanced multimodal fusion framework with congruent reinforced perceptron for hateful memes detection” Information Processing & Management 61(4): 103772. DOI: 10.1016/j.ipm.2024.103772.
  8. [8] J. Paul, S. Mallick, A. Mitra, A. Roy, and J. Sil, (2025) “Multi-modal Twitter Data Analysis for Identifying Offensive Posts Using a Deep Cross-Attention–based Transformer Framework” ACM Transactions on Knowledge Discovery from Data 19(3): 1–30. DOI: doi.org/10.1145/3713077.
  9. [9] S. B. Shah, S. Shiwakoti, M. Chaudhary, and H. Wang. “MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification”. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 17320–17332.
  10. [10] M. Tzelepi and V. Mezaris. “Improving multimodal hateful meme detection exploiting LMM-generated knowledge”. In: Proceedings of the Computer Vision and Pattern Recognition Conference. 2025, 202–211.
  11. [11] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. “Learning transferable visual models from natural language supervision”. In: International conference on machine learning. 2021, 8748–8763.
  12. [12] G. A. Koushik, D. Kanojia, and H. Treharne. “Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content”. In: Companion Proceedings of the ACM on Web Conference 2025. 2025, 2014–2023.
  13. [13] S. B. Shah, S. Shiwakoti, T. Bhuiyan, M. A. Moni, S. Thapa, and U. Naseem. “Entity-Aware Optimal Transport and Residual Attention for Multimodal Content Moderation”. In: Companion Proceedings of the ACM on Web Conference 2025. 2025, 2306–2313.
  14. [14] A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo. “Zero-shot composed image retrieval with textual inversion”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2023, 15338–15347.
  15. [15] S. Somvanshi, A. A. Javed, M. M. Islam, D. Pandit, and S. Das, (2025) “A survey on kolmogorov-arnold network” ACM Computing Surveys 58(2): 1–35. DOI: doi.org/10.1145/3743128.
  16. [16] Z. Zhan, X. Mao, H. Liu, and S. Yu, (2025) “STGL: Self-Supervised Spatio-Temporal Graph Learning for Traffic Forecasting” Journal of Artificial Intelligence Research 2(1): 1–8. DOI: 10.70891/JAIR.2025.040001.
  17. [17] W. Zhang and J. Wang, (2024) “English text sentiment analysis network based on CNN and U-Net” IFS/ACM Transactions on Machine Learning 1(1): 13–18. DOI: 10.70891/JSE.2024.100009.