{"id":8920,"date":"2026-07-03T17:43:09","date_gmt":"2026-07-03T09:43:09","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=8920"},"modified":"2026-07-03T18:34:06","modified_gmt":"2026-07-03T10:34:06","slug":"jase-202610-33-030","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202610-33-030","title":{"rendered":"A Bi-Mamba Fusion Framework with KAN Pattern Mining for Multi-modal Sensitive Content Detection"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=7886\" data-type=\"page\" data-id=\"7886\">Volume 33<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-07-03T17:43:09+08:00\">2026-07-03<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Ming Lian<a href=\"mailto:lianming@mail.ustc.edu.cn\"><i class=\"fa fa-envelope\"><\/i><\/a> and Teng Li<\/p>\n\n\n\n<p style=\"font-size:14px\">School of Artificial Intelligence, Anhui University, Hefei 230039 China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: May 7, 2026<br>Accepted:\u00a0June 4, 2026<br>Publication Date:\u00a0July 3, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/07\/33_030.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">The illustration of\u00a0MamKAN. It consists of a\u00a0multi-modal\u00a0feature\u00a0extraction\u00a0with\u00a0textual\u00a0inversion, a\u00a0bidirectional\u00a0Mamba fusion, and an adaptive pattern\u00a0mining\u00a0<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a rel=\"noreferrer noopener\" href=\"\/jase\/wp-content\/uploads\/2026\/06\/V33.0026.txt\" data-type=\"attachment\" data-id=\"8786\" target=\"_blank\">BibTeX <\/a>| <a href=\"http:\/\/dx.doi.org\/10.6180\/jase.202610_33.030\" target=\"_blank\" rel=\"noreferrer noopener\">http:\/\/dx.doi.org\/10.6180\/jase.202610_33.030<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/07\/030_2026_1171_V33.pdf\" data-type=\"attachment\" data-id=\"8925\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>The meteoric expansion of social media platforms has established multi-modal data as the dominant paradigm for global digital communication. However, this evolution has concurrently facilitated the rapid dissemination of sensitive and harmful material, including hate speech, extremist propaganda, and violent imagery, which presents a severe challenge to online security and public psychological well-being. Modern automated detection systems frequently falter because they fail to bridge the heterogeneous semantic gap between visual and linguistic modalities effectively, often relying on late-stage fusion or computationally expensive Transformer architectures that struggle with long-range dependencies. Furthermore, the conventional use of multi-layer perceptrons (MLPs) as predictors often fails to capture the irregular and complex decision boundaries inherent in high-dimensional multi-modal manifolds. To address these critical limitations, we propose MamKAN, a sophisticated multi-modal fusion framework that integrates State Space Models with spline-based neural architectures. Our methodology innovates at three distinct stages: First, we implement an implicit modal interaction strategy using textual inversion, which maps visual features into pseudo-word tokens within the text embedding space to achieve robust early-stage semantic alignment. Second, we introduce a bidirectional Mamba fusion module based on state space models; this mechanism achieves linear computational complexity while simultaneously capturing comprehensive vision-to-text and text-to-vision contextual dependencies through forward and backward scanning. Finally, we replace traditional MLPs with a Kolmogorov-Arnold Network (KAN) predictor. By employing learnable B-spline activation functions on the network edges, the KAN module adaptively fits complex non-linear distributions, significantly enhancing discriminative precision. Comprehensive experiments conducted on the HMC and HarMeme benchmark datasets demonstrate that<br>MamKAN consistently and significantly outperforms state-of-the-art models across three metrics.<\/p>\n\n\n\n<p><em>Keywords:\u00a0State space models; Kolmogorov-Arnold network; Sensitive content detection<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_442dd220420d5a90\" class=\"markdown markdown-main-panel stronger enable-updated-hr-color\" dir=\"ltr\" aria-live=\"polite\" aria-busy=\"false\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_bb65d10f6a8ddbc4\" class=\"markdown markdown-main-panel stronger enable-updated-hr-color\" dir=\"ltr\" aria-live=\"polite\" aria-busy=\"false\">\n<ol>\n<li data-path-to-node=\"0\">[1] D. Povedano \u00c1lvarez, A. L. Sandoval Orozco, J. P. Garc\u00eda-Miguel, and L. J. Garc\u00eda Villalba, (2023) \u201cLearning strategies for sensitive content detection\u201d Electronics 12(11): 2496. DOI: 10.3390\/electronics12112496.<\/li>\n<li data-path-to-node=\"0\">[2] X. Meng and Y. Xu, (2019) \u201cResearch on sensitive content detection in social networks\u201d CCF Transactions on Networking 2(2): 126\u2013135. DOI: 10.1007\/s42045-019-00021-x.<\/li>\n<li data-path-to-node=\"0\">[3] J. Gao, M. Liu, P. Li, J. Zhang, and Z. Chen, (2024) \u201cDeep Multiview Adaptive Clustering With Semantic Invariance\u201d IEEE Transactions on Neural Networks and Learning Systems 35(9): 12965\u201312978. DOI: 10.1109\/TNNLS.2023.3265699.<\/li>\n<li data-path-to-node=\"0\">[4] J. Gao, M. Liu, P. Li, A. A. Laghari, A. R. Javed, N. Victor, and T. R. Gadekallu, (2024) \u201cDeep Incomplete Multiview Clustering via Information Bottleneck for Pattern Mining of Data in Extreme-Environment IoT\u201d IEEE Internet of Things Journal 11(1): 26700\u201326712. DOI: 10.1109\/JIOT.2023.3325272.<\/li>\n<li data-path-to-node=\"0\">[5] T. Markov, C. Zhang, S. Agarwal, F. E. Nekoul, T. Lee, S. Adler, A. Jiang, and L. Weng. \u201cA holistic approach to undesired content detection in the real world\u201d. In: Proceedings of the AAAI conference on artificial intelligence. 37. 12. 2023, 15009\u201315018. DOI: 10.1609\/aaai.v37i12.26752.<\/li>\n<li data-path-to-node=\"0\">[6] G. Burbi, A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo. \u201cMapping memes to words for multimodal hateful meme classification\u201d. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2023, 2832\u20132836.<\/li>\n<li data-path-to-node=\"0\">[7] F. Wu, B. Gao, X. Pan, L. Li, Y. Ma, S. Liu, and Z. Liu, (2024) \u201cFuser: An enhanced multimodal fusion framework with congruent reinforced perceptron for hateful memes detection\u201d Information Processing &amp; Management 61(4): 103772. DOI: 10.1016\/j.ipm.2024.103772.<\/li>\n<li data-path-to-node=\"0\">[8] J. Paul, S. Mallick, A. Mitra, A. Roy, and J. Sil, (2025) \u201cMulti-modal Twitter Data Analysis for Identifying Offensive Posts Using a Deep Cross-Attention\u2013based Transformer Framework\u201d ACM Transactions on Knowledge Discovery from Data 19(3): 1\u201330. DOI: doi.org\/10.1145\/3713077.<\/li>\n<li data-path-to-node=\"0\">[9] S. B. Shah, S. Shiwakoti, M. Chaudhary, and H. Wang. \u201cMemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification\u201d. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 17320\u201317332.<\/li>\n<li data-path-to-node=\"0\">[10] M. Tzelepi and V. Mezaris. \u201cImproving multimodal hateful meme detection exploiting LMM-generated knowledge\u201d. In: Proceedings of the Computer Vision and Pattern Recognition Conference. 2025, 202\u2013211.<\/li>\n<li data-path-to-node=\"0\">[11] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. \u201cLearning transferable visual models from natural language supervision\u201d. In: International conference on machine learning. 2021, 8748\u20138763.<\/li>\n<li data-path-to-node=\"0\">[12] G. A. Koushik, D. Kanojia, and H. Treharne. \u201cTowards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content\u201d. In: Companion Proceedings of the ACM on Web Conference 2025. 2025, 2014\u20132023.<\/li>\n<li data-path-to-node=\"0\">[13] S. B. Shah, S. Shiwakoti, T. Bhuiyan, M. A. Moni, S. Thapa, and U. Naseem. \u201cEntity-Aware Optimal Transport and Residual Attention for Multimodal Content Moderation\u201d. In: Companion Proceedings of the ACM on Web Conference 2025. 2025, 2306\u20132313.<\/li>\n<li data-path-to-node=\"0\">[14] A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo. \u201cZero-shot composed image retrieval with textual inversion\u201d. In: Proceedings of the IEEE\/CVF international conference on computer vision. 2023, 15338\u201315347.<\/li>\n<li data-path-to-node=\"0\">[15] S. Somvanshi, A. A. Javed, M. M. Islam, D. Pandit, and S. Das, (2025) \u201cA survey on kolmogorov-arnold network\u201d ACM Computing Surveys 58(2): 1\u201335. DOI: doi.org\/10.1145\/3743128.<\/li>\n<li data-path-to-node=\"0\">[16] Z. Zhan, X. Mao, H. Liu, and S. Yu, (2025) \u201cSTGL: Self-Supervised Spatio-Temporal Graph Learning for Traffic Forecasting\u201d Journal of Artificial Intelligence Research 2(1): 1\u20138. DOI: 10.70891\/JAIR.2025.040001.<\/li>\n<li data-path-to-node=\"0\">[17] W. Zhang and J. Wang, (2024) \u201cEnglish text sentiment analysis network based on CNN and U-Net\u201d IFS\/ACM Transactions on Machine Learning 1(1): 13\u201318. DOI: 10.70891\/JSE.2024.100009.<\/li>\n<\/ol>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1483,6],"tags":[1644],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202610_33.030\u00a0\u00a0 Download PDF The meteoric expansion of social media platforms has established multi-modal data&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/8920"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8920"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8920"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8920"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}