School of International Communication, Communication University of China, Nanjing 211172, Jiangsu, China
Received: May 20, 2026
Accepted: June 27, 2026
Publication Date: August 05, 2026
Architecture of the Multimodal Transformer for Cross-Cultural Emotion Recognition
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202611_34.019
Cross-cultural awareness and communication training for bilingual broadcasters should place the emphasis on high-engagement, tailored and transactive learning environments that support multilingual communication, socio-contextual processing, and cultural awareness. The objective is to establish an artificial intelligence (AI)-based virtual reality (VR) solution for immersive bilingual communication by utilizing multimodal emotion
analysis and real-time dynamic feedback. The proposed framework combines two publicly available multimodal datasets CAMEO and MELD with a systematically controlled human study dataset at the core collected from 110 bilingual speakers. In general, the procedures for this study are multimodal data acquisition, MFCC based speech feature extraction, multilingual BERT (mBERT) text embedding, bilingual code-switch processing, emotion encoding, and stratified data splitting. A multimodal transformer based on attention to jointly represent text and speech data for contextual emotion recognition and bilingual interaction assessment. The framework includes cross-cultural adaptation and AI-based real-time feedback in the immersive VR communication settings. Simulation results indicate that the proposed model achieves the best performance with an Accuracy of 98.63%,
a Precision of 98.62%, a Recall of 98.62%, and an F1-Score of 98.62%. The results suggest that the proposed framework surpass both the CNN, the BiLSTM, the transformer encoder, as well as other existing multimodal methods in the task of emotion recognition. In this work we report on real-time inference with an average response latency of 8.37ms that facilitates natural immersive interaction.
Keywords: Virtual Reality, Multimodal Transformer, Mel-Frequency Cepstral Coefficients, Bilingual Broadcasting, Bidirectional Encoder Representations from Transformers
- [1] Y.-J. Wu, T.-J. Ding, J.-C. Hsu, K.-L. Ou, and W. Tarng, (2025) “Exploratory Learning of Amis Indigenous Culture and Local Environments Using Virtual Reality and Drone Technology” ISPRS International Journal of Geo-Information 14(11): 441. DOI: 10.3390/ijgi14110441.
- [2] Z. Yang, (2025) “Design of a Visual Communication System for Animated Characters in Virtual Reality Using Sobel Edge Detection and Motion Capture” Journal of Applied Science and Engineering 29(1): 235–243. DOI: 10.6180/jase.202601_29(1).0023.
- [3] A. K. Jumani, K. Kumar, and M. A. Chhajro, (2021) “Systematic Analysis of Virtual Reality & Augmented Reality” International Journal of Information Engineering & Electronic Business 13(1): 1–12. DOI: 10.5815/ijieeb.2021.01.04.
- [4] A. A. Laghari, V. V. Estrela, H. Li, Y. Shoulin, A. A. Khan, M. S. Anwar, A. Wahab, and K. Bouraqia, (2024) “Quality of Experience Assessment in Virtual/Augmented Reality Serious Games for Healthcare: A Systematic Literature Review” Technology and Disability 36(1–2): 17–28. DOI: 10.3233/TAD-230035.
- [5] A. K. Jumani, J. Shi, A. A. Laghari, V. V. Estrela, G. A. Sampedro, A. Almadhor, N. Kryvinska, and A. U. Nabi, (2024) “Quality of Experience That Matters in Gaming Graphics: How to Blend Image Processing and Virtual Reality” Electronics 13(15): 2998. DOI: 10.3390/electronics13152998.
- [6] A. K. Jumani, J. Shi, A. A. Laghari, M. A. Amin, A. U. Nabi, K. Narwani, and Y. Zhang, (2025) “Quality of Experience (QoE) in Cloud Gaming: A Comparative Analysis of Deep Learning Techniques via Facial Emotions in a Virtual Reality Environment” Sensors 25(5): 1594. DOI: 10.3390/s25051594.
- [7] A. A. Laghari, S. Shahid, R. Yadav, S. Karim, A. Khan, H. Li, and Y. Shoulin, (2023) “The State of Art and Review on Video Streaming” Journal of High Speed Networks 29(3): 211–236. DOI: 10.3233/JHS-222087.
- [8] M. R. Sareddy and V. Kumar. “Virtual Reality Meets AI: Revolutionizing Physical Education with LiDAR and RL”. In: 2025 8th International Conference on Electronics, Materials Engineering & Nano-Technology (IEMEnTech). IEEE, 2025, 1–6. DOI: 10.1109/IEMEnTech65115.2025.10959637.
- [9] M. Maciejewski, J. Kocoń, W. Oleksy, M. Kopeć, and W. Kazienko. amu-cai/CAMEO Dataset. Accessed: May 9, 2026. 2026.
- [10] S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea. “MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations”. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. DOI: 10.18653/v1/P19-1050.
- [11] B. Xie, M. Sidulova, and C. H. Park, (2021) “Robust Multimodal Emotion Recognition from Conversation with Transformer-Based Crossmodality Fusion” Sensors 21(14): 4913. DOI: 10.3390/s21144913.
- [12] C. Qu et al., (2025) “Enhancing Emotion Recognition in Virtual Reality: A Multimodal Dataset and a Temporal Emotion Detector” Frontiers in Psychology 16(2): 1709943. DOI: 10.3389/fpsyg.2025.1709943.
- [13] H. Wang and M. Wang, (2026) “Subject-Independent Multimodal Interaction Modeling for Joint Emotion and Immersion Estimation in Virtual Reality” Symmetry 18(3): 451. DOI: 10.3390/sym18030451.
