{"id":11725,"date":"2026-09-13T22:41:46","date_gmt":"2026-09-13T14:41:46","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=11725"},"modified":"2026-09-13T23:15:44","modified_gmt":"2026-09-13T15:15:44","slug":"jase-202612-35-034","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202612-35-034","title":{"rendered":"Cross Modal Fusion Network: Integrating Visual Language Features for Context Aware Visual Communication Design"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=11162\" data-type=\"page\" data-id=\"11162\">Volume 35<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-09-13T22:41:46+08:00\">2026-09-13<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Ying Zhang<a href=\"mailto:ying_zhang67@outlook.com\"><i class=\"fa fa-envelope\"><\/i><\/a><\/p>\n\n\n\n<p style=\"font-size:14px\">Xinxiang Vocational and Technical College Henan, 453000, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: May 09, 2026<br>Accepted: August 05, 2026<br>Publication Date: September 13, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/09\/35_034.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Overall architecture of the proposed model <\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:  <a href=\"\/jase\/wp-content\/uploads\/2026\/09\/V35.0034.txt\" data-type=\"attachment\" data-id=\"11749\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202612_35.034\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202612_35.034<\/a>  <\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/09\/034_2026_1192_V35.pdf\" data-type=\"attachment\" data-id=\"11730\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Multimodal information integration is essential for developing context-aware visual communication designs, especially as digital media increasingly relies on flexible and efficient interaction. Traditional approaches often treat text and images as separate elements, limiting contextual coherence and weakening message clarity. This study proposes a cross-modal fusion network that integrates visual and linguistic information to enhance context-aware visual communication. Using the Scene Parse 150 dataset, preprocessing involved tokenization, image resizing, and feature extraction through a convolutional neural network. The model introduces a Self-Attention Based Generative Adversarial-tuned Robustly Optimized BERT Pretraining Approach (SAGA-ROBERTa), where RoBERTa encodes textual descriptions to capture semantic richness and guide a generative adversarial network in producing or refining visual design layouts. A multi-channel fusion module with self-attention further captures intrinsic relationships between modalities, ensuring strong semantic alignment. The discriminator evaluates visual realism and coherence with textual intent. Experimental results demonstrate that SAGA-ROBERTa achieves notable improvements, including an Average IoU of 68.6% and pixel accuracy of 94.5%, outperforming conventional unimodal methods. The findings highlight the potential of cross-modal deep learning frameworks to support more adaptive, emotionally resonant, and semantically precise visual communication across diverse media contexts.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;Cross-Modal Fusion, Visual-Language Integration, Context-Aware Design, Visual Communication Design, Self-Attention Based Generative Adversarial-tuned Robustly Optimized BERT Pretraining Approach (SAGA-ROBERTa).<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_53180f64958e6cc2\" class=\"markdown markdown-main-panel md-content enable-luminous-fast-follows enable-updated-hr-color stronger\" dir=\"ltr\" aria-busy=\"false\" aria-live=\"polite\">\n<ol>\n<li data-path-to-node=\"0\">[1] J. Zhu, (2025) &#8220;Visual contextual perception and user emotional feedback in visual communication design&#8221; BMC Psychology 13: 1-13. DOI: https:\/\/doi.org\/10.1186\/s40359-025-02615-1.<\/li>\n<li data-path-to-node=\"0\">[2] C. Zhang, W. Zeng, and L. Liu, (2021) &#8220;Urban VR: An immersive analytics system for context-aware urban design&#8221; Computers &amp; Graphics 99: 128-138. DOI: https:\/\/doi.org\/10.1016\/j.cag.2021.07.006.<\/li>\n<li data-path-to-node=\"0\">[3] H. Elfaik, (2021) &#8220;Combining context-aware embeddings and an attentional deep learning model for Arabic affect analysis on Twitter&#8221; IEEE Access 9: 111214-111230. DOI: https:\/\/doi.org\/10.1109\/ACCESS.2021.3102087.<\/li>\n<li data-path-to-node=\"0\">[4] A. Manolova, K. Tonchev, V. Poulkov, S. Dixit, and P. Lindgren, (2021) &#8220;Context-aware holographic communication based on semantic knowledge extraction&#8221; Wireless Personal Communications 120: 2307-2319. DOI: https:\/\/doi.org\/10.1007\/s11277-021-08560-7.<\/li>\n<li data-path-to-node=\"0\">[5] L. Lu and L. Huang, (2022) &#8220;Exploration and application of graphic design language based on artificial intelligence visual communication&#8221; Wireless Communications and Mobile Computing 2022: 9907303. DOI: https:\/\/doi.org\/10.1155\/2022\/9907303.<\/li>\n<li data-path-to-node=\"0\">[6] Z. Xie, B. Zhou, X. Cheng, E. Schoenfeld, and F. Ye, (2022) &#8220;Passive and context-aware in-home vital signs monitoring using co-located UWB-depth sensor fusion&#8221; ACM Transactions on Computing for Healthcare 3: 1-31. DOI: https:\/\/doi.org\/10.1145\/3549941.<\/li>\n<li data-path-to-node=\"0\">[7] A. Omolaja, A. Otebolaku, and A. Alfoudi, (2022) &#8220;Context-aware complex human activity recognition using hybrid deep learning models&#8221; Applied Sciences 12: 9305. DOI: https:\/\/doi.org\/10.3390\/app12189305.<\/li>\n<li data-path-to-node=\"0\">[8] A. Montanha, A. M. Oprescu, and M. Romero-Ternero, (2022) &#8220;A context-aware artificial intelligence-based system to support street crossings for pedestrians with visual impairments&#8221; Applied Artificial Intelligence 36: 2062818. DOI: https:\/\/doi.org\/10.1080\/08839514.2022.2062818.<\/li>\n<li data-path-to-node=\"0\">[9] Z. Ren, L. He, and J. Lu, (2023) &#8220;Context-aware edge-enhanced GAN for remote sensing image super-resolution&#8221; IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 17: 1363-1376. DOI: https:\/\/doi.org\/10.1109\/JSTARS.2023.3333271.<\/li>\n<li data-path-to-node=\"0\">[10] X. Chen, J. Li, T. Gao, Y. Piao, H. Ji, B. Yang, and W. Xu, (2024) &#8220;Dynamic context-aware pyramid network for infrared small target detection&#8221; IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing: DOI: https:\/\/doi.org\/10.1109\/JSTARS.2024.3434330.<\/li>\n<li data-path-to-node=\"0\">[11] A. Papastratis, K. Dimitropoulos, and P. Daras, (2021) &#8220;Continuous sign language recognition through a context-aware generative adversarial network&#8221; Sensors 21: 2437. DOI: https:\/\/doi.org\/10.3390\/s21072437.<\/li>\n<li data-path-to-node=\"0\">[12] T. Abirami, A. K. Dutta, and S. Alsubai, (2024) &#8220;Deep learning-powered visual place recognition for enhanced mobile multimedia communication in autonomous transport systems&#8221; Alexandria Engineering Journal 109: 950-962. DOI: https:\/\/doi.org\/10.1016\/j.aej.2024.09.060.<\/li>\n<li data-path-to-node=\"0\">[13] M. Rahman, M. A. M. Provath, K. Deb, P. K. Dhar, and T. Shimamura, (2025) &#8220;CAMFusion: Context-aware multi-modal fusion framework for detecting sarcasm and humor integrating video and textual cues&#8221; IEEE Access: DOI: https:\/\/doi.org\/10.1109\/ACCESS.2025.3535694.<\/li>\n<li data-path-to-node=\"0\">[14] A. Yin and K. Yin, (2025) &#8220;Robust image watermarking using bidirectional-interactive and context-aware networks&#8221; IEEE Transactions on Circuits and Systems for Video Technology: DOI: https:\/\/doi.org\/10.1109\/TCSVT.2025.3543969.<\/li>\n<li data-path-to-node=\"0\">[15] L. Liu and J. Yang, (2025) &#8220;A meta-learning-based multi-scene student posture detection method for enhancing learning motivation and engagement in smart classrooms&#8221; Journal of Applied Science and Engineering 29(3): 707-724. DOI: https:\/\/doi.org\/10.6180\/jase.202603_29(3).0021.<\/li>\n<li data-path-to-node=\"0\">[16] L. Lu, (2020) &#8220;Design of visual communication based on deep learning approaches&#8221; Soft Computing 24: 7861-7872. DOI: https:\/\/doi.org\/10.1007\/s00500-019-03954-z.<\/li>\n<\/ol>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1956,6],"tags":[2114],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation: BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202612_35.034 Download PDF Multimodal information integration is essential for developing context-aware visual communication designs,&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/11725"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=11725"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=11725"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=11725"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}