Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Cross Modal Fusion Network: Integrating Visual Language Features for Context Aware Visual Communication Design

Ying Zhang

Xinxiang Vocational and Technical College Henan, 453000, China

Received: May 09, 2026
Accepted: August 05, 2026
Publication Date: September 13, 2026

上傳圖片

Overall architecture of the proposed model

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202612_35.034

Download PDF

Multimodal information integration is essential for developing context-aware visual communication designs, especially as digital media increasingly relies on flexible and efficient interaction. Traditional approaches often treat text and images as separate elements, limiting contextual coherence and weakening message clarity. This study proposes a cross-modal fusion network that integrates visual and linguistic information to enhance context-aware visual communication. Using the Scene Parse 150 dataset, preprocessing involved tokenization, image resizing, and feature extraction through a convolutional neural network. The model introduces a Self-Attention Based Generative Adversarial-tuned Robustly Optimized BERT Pretraining Approach (SAGA-ROBERTa), where RoBERTa encodes textual descriptions to capture semantic richness and guide a generative adversarial network in producing or refining visual design layouts. A multi-channel fusion module with self-attention further captures intrinsic relationships between modalities, ensuring strong semantic alignment. The discriminator evaluates visual realism and coherence with textual intent. Experimental results demonstrate that SAGA-ROBERTa achieves notable improvements, including an Average IoU of 68.6% and pixel accuracy of 94.5%, outperforming conventional unimodal methods. The findings highlight the potential of cross-modal deep learning frameworks to support more adaptive, emotionally resonant, and semantically precise visual communication across diverse media contexts.

Keywords: Cross-Modal Fusion, Visual-Language Integration, Context-Aware Design, Visual Communication Design, Self-Attention Based Generative Adversarial-tuned Robustly Optimized BERT Pretraining Approach (SAGA-ROBERTa).

  1. [1] J. Zhu, (2025) “Visual contextual perception and user emotional feedback in visual communication design” BMC Psychology 13: 1-13. DOI: https://doi.org/10.1186/s40359-025-02615-1.
  2. [2] C. Zhang, W. Zeng, and L. Liu, (2021) “Urban VR: An immersive analytics system for context-aware urban design” Computers & Graphics 99: 128-138. DOI: https://doi.org/10.1016/j.cag.2021.07.006.
  3. [3] H. Elfaik, (2021) “Combining context-aware embeddings and an attentional deep learning model for Arabic affect analysis on Twitter” IEEE Access 9: 111214-111230. DOI: https://doi.org/10.1109/ACCESS.2021.3102087.
  4. [4] A. Manolova, K. Tonchev, V. Poulkov, S. Dixit, and P. Lindgren, (2021) “Context-aware holographic communication based on semantic knowledge extraction” Wireless Personal Communications 120: 2307-2319. DOI: https://doi.org/10.1007/s11277-021-08560-7.
  5. [5] L. Lu and L. Huang, (2022) “Exploration and application of graphic design language based on artificial intelligence visual communication” Wireless Communications and Mobile Computing 2022: 9907303. DOI: https://doi.org/10.1155/2022/9907303.
  6. [6] Z. Xie, B. Zhou, X. Cheng, E. Schoenfeld, and F. Ye, (2022) “Passive and context-aware in-home vital signs monitoring using co-located UWB-depth sensor fusion” ACM Transactions on Computing for Healthcare 3: 1-31. DOI: https://doi.org/10.1145/3549941.
  7. [7] A. Omolaja, A. Otebolaku, and A. Alfoudi, (2022) “Context-aware complex human activity recognition using hybrid deep learning models” Applied Sciences 12: 9305. DOI: https://doi.org/10.3390/app12189305.
  8. [8] A. Montanha, A. M. Oprescu, and M. Romero-Ternero, (2022) “A context-aware artificial intelligence-based system to support street crossings for pedestrians with visual impairments” Applied Artificial Intelligence 36: 2062818. DOI: https://doi.org/10.1080/08839514.2022.2062818.
  9. [9] Z. Ren, L. He, and J. Lu, (2023) “Context-aware edge-enhanced GAN for remote sensing image super-resolution” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 17: 1363-1376. DOI: https://doi.org/10.1109/JSTARS.2023.3333271.
  10. [10] X. Chen, J. Li, T. Gao, Y. Piao, H. Ji, B. Yang, and W. Xu, (2024) “Dynamic context-aware pyramid network for infrared small target detection” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing: DOI: https://doi.org/10.1109/JSTARS.2024.3434330.
  11. [11] A. Papastratis, K. Dimitropoulos, and P. Daras, (2021) “Continuous sign language recognition through a context-aware generative adversarial network” Sensors 21: 2437. DOI: https://doi.org/10.3390/s21072437.
  12. [12] T. Abirami, A. K. Dutta, and S. Alsubai, (2024) “Deep learning-powered visual place recognition for enhanced mobile multimedia communication in autonomous transport systems” Alexandria Engineering Journal 109: 950-962. DOI: https://doi.org/10.1016/j.aej.2024.09.060.
  13. [13] M. Rahman, M. A. M. Provath, K. Deb, P. K. Dhar, and T. Shimamura, (2025) “CAMFusion: Context-aware multi-modal fusion framework for detecting sarcasm and humor integrating video and textual cues” IEEE Access: DOI: https://doi.org/10.1109/ACCESS.2025.3535694.
  14. [14] A. Yin and K. Yin, (2025) “Robust image watermarking using bidirectional-interactive and context-aware networks” IEEE Transactions on Circuits and Systems for Video Technology: DOI: https://doi.org/10.1109/TCSVT.2025.3543969.
  15. [15] L. Liu and J. Yang, (2025) “A meta-learning-based multi-scene student posture detection method for enhancing learning motivation and engagement in smart classrooms” Journal of Applied Science and Engineering 29(3): 707-724. DOI: https://doi.org/10.6180/jase.202603_29(3).0021.
  16. [16] L. Lu, (2020) “Design of visual communication based on deep learning approaches” Soft Computing 24: 7861-7872. DOI: https://doi.org/10.1007/s00500-019-03954-z.