{"id":7025,"date":"2026-05-21T21:43:56","date_gmt":"2026-05-21T13:43:56","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=7025"},"modified":"2026-05-21T22:58:14","modified_gmt":"2026-05-21T14:58:14","slug":"jase-202609-32-052","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202609-32-052","title":{"rendered":"Multimodal Fusion for Text-to-Image Synthesis: A GAN Framework Driven by CLIP and CAM"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=3671\" data-type=\"page\" data-id=\"1055\">Volume 32<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-05-21T21:43:56+08:00\">2026-05-21<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Qiuyong Huang and Ailong Tang<a href=\"mailto:110hqy@163.com\"><i class=\"fa fa-envelope\"><\/i><\/a> <\/p>\n\n\n\n<p style=\"font-size:14px\">College of Information Science and Engineering, Liuzhou Institute of Technology, Liuzhou 545616, Guangxi, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: October 25, 2025<br>Accepted:&nbsp;May 8, 2026<br>Publication Date:&nbsp;May 21, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/05\/32_052.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Structure of CLIP-CA-GAN<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a href=\"\/jase\/wp-content\/uploads\/2026\/05\/V32.0052.txt\" data-type=\"attachment\" data-id=\"6681\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202609_32.052\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202609_32.052<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/05\/052_2025_1622_V32.pdf\" data-type=\"attachment\" data-id=\"7035\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Targeting the issues of weak fine-grained alignment capability and insufficient semantic controllability in existing generative adversarial network approaches, this paper presents a multimodal fusion-based model, namely Contrastive Language\u2013Image Pretraining-Cross-Attention-Generative Adversarial Networks (CLIP CA-GAN). With GAN as the basic architecture, this model incorporates the Contrastive Language-Image Pretraining (CLIP) model to establish multimodal semantic constraints. It dynamically fuses the local features of text and images via the Cross-Attention Mechanism (CAM), and optimizes generation quality through a designed Feature Fusion Module and a comprehensive loss function (LF). Experimental results demonstrate that the performance of CLIP-CA-GAN outperforms mainstream methods. On MS-COCO, the Fr\u00e9chet Inception Distance (FID) decreases to 16.09, and the Inception Score (IS) rises to 4.91. On CUB, the FID stands at 14.06, the IS at 5.33, and the R\u2013precision (RP) reaches 79.24. Additionally, the model has a relatively small number of parameters and high training efficiency, thus providing a high-quality and low-complexity solution for image generation.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;CLIP-CA-GAN, multimodal, CLIP, fine-grained alignment, CAM<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<ol>\n<li>[1] F. Bie, Y. Yang, Z. Zhou, A. Ghanem, M. Zhang, Z. Yao, X. Wu, C. Holmes, P. Golnari, D. A. Clifton, et al., (2024) \u201cRenaissance: A survey into ai text-to-image generation in the era of large model\u201d IEEE transactions on pattern analysis and machine intelligence 47(3): 2212\u20132231. DOI: 10.1109\/TPAMI.2024.3522305.<\/li>\n<li>[2] V. Paananen, J. Oppenlaender, and A. Visuri, (2024) \u201cUsing text-to-image generation for architectural design ideation\u201d International Journal of Architectural Computing 22(3): 458\u2013474. DOI: 10.1177\/14780771231222783.<\/li>\n<li>[3] J. Oppenlaender. \u201cThe cultivated practices of text-to-image generation\u201d. In: Humane Autonomous Technology: Re-thinking Experience with and in Intelligent Systems. Springer, 2024, 325\u2013349. DOI: 10.1007\/978-3-031-66528-8_14.<\/li>\n<li>[4] L. H\u00f6llein, A. Bo\u017ei\u010d, N. M\u00fcller, D. Novotny, H.-Y. Tseng, C. Richardt, M. Zollh\u00f6fer, and M. Nie\u00dfner. \u201cViewdiff: 3d-consis<span class=\"citation-339\">tent image generation with text-to-image models\u201d. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 2024, 5043\u20135052. URL: <\/span><a class=\"ng-star-inserted\" href=\"https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Hollein_ViewDiff_3D-Consistent_Image_Generation_with_Text-to-Image_Models_CVPR_2024_paper.html\" target=\"_blank\" rel=\"noopener\"><span class=\"citation-339 citation-end-339\">https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Hollein_ViewDiff_3D-C<\/span>onsistent_Image_Generation_with_Text-to-Image_Models_CVPR_2024_paper.html<\/a>.<\/li>\n<li>[5] J. Gartner and M. Romanov, (2024) \u201cThe advantages of ai text to image generation\u201d International Journal of Art, Design, and Metaverse 2(1): 1\u20138. DOI: <a class=\"ng-star-inserted\" href=\"https:\/\/topazart.info\/e-journals\/index.php\/ijam\/article\/view\/65\" target=\"_blank\" rel=\"noopener\">https:\/\/topazart.info\/e-journals\/index.php\/ijam\/article\/view\/65<\/a>.<\/li>\n<li>[6] A. Dunkel, D. Burghardt, and M. Gugulica, (2024) \u201cGenerative text-to-image diffusion for automated map production based on geosocial media data\u201d KN-Journal of Cartography and Geographic Information 74(1): 3\u201315. DOI: 10.1007\/s42489-024-00159-9.<\/li>\n<li>[7] A. A. Laghari, V. V. Estrela, and S. Yin, (2024) \u201cHow to collect and interpret medical pictures captured in highly challenging environments that range from nanoscale to hyperspectral imaging\u201d Current Medical Imaging 20(1): e28122212228. DOI: 10.2174\/1573405619666221228094228.<\/li>\n<li>[8] S. Narasimhaswamy, U. Bhattacharya, X. Chen, I. Dasgupta, S. Mitra, and M. Hoai. \u201cHandiffuser: Text-to-image generation with realistic hand appearances\u201d. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2024, 2468\u20132479. URL: <a class=\"ng-star-inserted\" href=\"https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Narasimhaswamy_HanDiffuser_Text-to-Image_Generation_With_Realistic_Hand_Appearances_CVPR_2024_paper.html\" target=\"_blank\" rel=\"noopener\">https:\/\/openaccess.thecvf.com\/content\/CVPR2024\/html\/Narasimhaswamy_HanDiffuser_Text-to-Image_Generation_With_Realistic_Hand_Appearances_CVPR_2024_paper.html<\/a>.<\/li>\n<li>[9] A. A. Laghari, Y. Sun, M. Alhussein, K. Aurangzeb, M. S. Anwar, and M. Rashid, (2023) \u201cDeep residual-dense network based on bidirectional recurrent neural network for atrial fibrillation detection\u201d Scientific reports 13(1): 15109. DOI: 10.1038\/s41598-023-40343-x.<\/li>\n<li>[10] A. A. Laghari, S. Shahid, R. Yadav, S. Karim, A. Khan, H. Li, and Y. Shoulin, (2023) \u201cThe state of art and review on video streaming\u201d Journal of High Speed Networks 29(3): 211\u2013236. DOI: 10.3233\/JHS-222087.<\/li>\n<li>[11] M. A. Munir, R. A. Shah, M. Ali, A. A. Laghari, A. Almadhor, and T. R. Gadekallu, (2024) \u201cEnhancing gene mutation prediction with sparse regularized autoencoders in lung cancer radiomics analysis\u201d IEEE Access 13: 7407\u20137425. DOI: 10.1109\/ACCESS.2024.3523330.<\/li>\n<li>[12] S. Karim, Y. Zhang, A. A. Laghari, and M. R. Asif. \u201cImage processing based proposed drone for detecting and controlling street crimes\u201d. In: 2017 IEEE 17th International Conference on Communication Technology (ICCT). IEEE. 2017, 1725\u20131730. DOI: 10.1109\/ICCT.2017.8359925.<\/li>\n<li>[13] U. Saeed, K. Kumar, M. A. Khuhro, A. A. Laghari, A. A. Shaikh, and A. Rai, (2024) \u201cDeepLeukNet\u2014A CNN based microscopy adaptation model for acute lymphoblastic leukemia classification\u201d Multimedia Tools and Applications 83(7): 21019\u201321043. DOI: 10.1007\/s11042-023-16191-2.<\/li>\n<li>[14] G. Marcus, E. Davis, and S. Aaronson, (2022) \u201cA very preliminary analysis of DALL-E 2\u201d arXiv preprint arXiv:2204.13807: DOI: 10.48550\/arXiv.2204.13807.<\/li>\n<li>[15] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. \u201cHigh-resolution image synthesis with latent diffusion models\u201d. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 2022, 10684\u201310695. URL: <a class=\"ng-star-inserted\" href=\"https:\/\/www.google.com\/search?q=https:\/\/openaccess.thecvf.com\/content\/CVPR2022\/html\/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.html%3Futm_source%3Drns.dwaiat.de\" target=\"_blank\" rel=\"noopener\">https:\/\/openaccess.thecvf.com\/content\/CVPR2022\/html\/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.html?utm_source=rns.dwaiat.de<\/a>.<\/li>\n<li>[16] H. Li and X.-J. Wu, (2024) \u201cCrossFuse: A novel cross attention mechanism based infrared and visible image fusion approach\u201d Information Fusion 103: 102147. DOI: 10.1016\/j.inffus.2023.102147.<\/li>\n<li>[17] R. Tamilkodi, K. Suryakala, N. Vamsi, M. Arvind Reddy, M. Nithin Kumar, and M. Venkat. \u201cTransforming text into art: Exploring dall-e\u2019s text-to-image generation capabilities\u201d. In: International Conference on Smart Data Intelligence. Springer. 2024, 413\u2013421. DOI: 10.1007\/978-981-97-3191-6_31.<\/li>\n<li>[18] Y. Shi, M. Shang, and Z. Qi, (2023) \u201cIntelligent layout generation based on deep generative models: A comprehensive survey\u201d Information Fusion 100: 101940. DOI: 10.1016\/j.inffus.2023.101940.<\/li>\n<li>[19] S. Naveen, M. S. R. Kiran, M. Indupriya, T. Manikanta, and P. Sudeep, (2021) \u201cTransformer models for enhancing AttnGAN based text to image generation\u201d Image and Vision Computing 115: 104284. DOI: 10.1016\/j.imavis.2021.104284.<\/li>\n<li>[20] L. Yan, R. Yan, B. Chai, G. Ceng, P. Zhou, and J. Gao, (2024) \u201cDM-GAN: CNN hybrid vits for training GANs under limited data\u201d Pattern Recognition 156: 110810. DOI: 10.1016\/j.patcog.2024.110810.<\/li>\n<li>[21] R. Mehmood, R. Bashir, and K. J. Giri. \u201cComparative Analysis of AttnGAN, DF-GAN and SSA-GAN\u201d. In: 2021 3rd International Conference on Advances in Computing, Communication Control and Networking (ICAC3N). IEEE. 2021, 370\u2013375. DOI: 10.1109\/ICAC3N53548.2021.9725424.<\/li>\n<li>[22] D. S. Patra and S. Padhee. \u201cComparative Analysis of ControlGAN and ControlGAN-GP Models based Text-to-Image Synthesis\u201d. In: 2022 OITS International Conference on Information Technology (OCIT). IEEE. 2022, 564\u2013568. DOI: 10.1109\/OCIT56763.2022.00110.<\/li>\n<li>[23] W. Liao, K. Hu, M. Y. Yang, and B. Rosenhahn. \u201cText to image generation with semantic-spatial aware gan\u201d. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 2022, 18187\u201318196. URL: <a class=\"ng-star-inserted\" href=\"https:\/\/openaccess.thecvf.com\/content\/CVPR2022\/html\/Liao_Text_to_Image_Generation_With_Semantic-Spatial_Aware_GAN_CVPR_2022_paper.html\" target=\"_blank\" rel=\"noopener\">https:\/\/openaccess.thecvf.com\/content\/CVPR2022\/html\/Liao_Text_to_Image_Generation_With_Semantic-Spatial_Aware_GAN_CVPR_2022_paper.html<\/a>.<\/li>\n<li>[24] S. Hou, Z. Li, K. Wu, Y. Zhao, and H. Li, (2024) \u201cMasked cross-attention and multi-channel attention guiding single-stage generative adversarial networks for text-to-image generation\u201d The Visual Computer 40(12): 8639\u20138651. DOI: 10.1007\/s00371-024-03260-2.<\/li>\n<li>[25] S. A. Baumann, F. Krause, M. Neumayr, N. Stracke, M. Sevi, V. T. Hu, and B. Ommer. \u201cContinuous, subject-specific attribute control in t2i models by identifying semantic directions\u201d. In: Proceedings of the Computer Vision and Pattern Recognition Conference. 2025, 13231\u201313241. DOI: <a class=\"ng-star-inserted\" href=\"https:\/\/openaccess.thecvf.com\/content\/CVPR2025\/html\/Baumann_Continuous_Subject-Specific_Attribute_Control_in_T2I_Models_by_Identifying_Semantic_CVPR_2025_paper.html\" target=\"_blank\" rel=\"noopener\">https:\/\/openaccess.thecvf.com\/content\/CVPR2025\/html\/Baumann_Continuous_Subject-Specific_Attribute_Control_in_T2I_Models_by_Identifying_Semantic_CVPR_2025_paper.html<\/a>.<\/li>\n<li>[26] K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu, (2023) \u201cT2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation\u201d Advances in Neural Information Processing Systems 36: 78723\u201378747. URL: <a class=\"ng-star-inserted\" href=\"https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/f8ad010cdd9143dbb0e9308c093aff24-Paper-Datasets_and_Benchmarks.pdf\" target=\"_blank\" rel=\"noopener\">https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/f8ad010cdd9143dbb0e9308c093aff24-Paper-Datasets_and_Benchmarks.pdf<\/a>.<\/li>\n<li>[27] T. Hu, L. Li, J. Van de Weijer, H. Gao, F. S. Khan, J. Yang, M.-M. Cheng, K. Wang, and Y. Wang, (2024) \u201cToken merging for training-free semantic binding in text-to-image synthesis\u201d Advances in Neural Information Processing Systems 37: 137646\u2013137672. DOI: 10.52202\/079017-4372.<\/li>\n<li>[28] N. S. Mudiraj and S. Singh, (2025) \u201cSemantic mapping of Hindi text-to-image generation using CUB dataset\u201d Scientific Reports 15(1): 36632. DOI: 10.1038\/s41598-025-20537-1.<\/li>\n<li>[29] O. Durusoy et al., (2025) \u201cOpen-source datasets for image processing and artificial intelligence research: A comparison of imagenet and ms coco datasets\u201d Int. J. Sci. Innov. Eng 2: 639\u2013653. DOI: 10.70849\/IJSCI20250202575.<\/li>\n<li>[30] J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, (2022) \u201cCascaded diffusion models for high fidelity image generation\u201d Journal of Machine Learning Research 23(47): 1\u201333. URL: <a class=\"ng-star-inserted\" href=\"https:\/\/www.google.com\/search?q=http:\/\/jmlr.org\/papers\/v23\/21-0635.html\" target=\"_blank\" rel=\"noopener\">http:\/\/jmlr.org\/papers\/v23\/21-0635.html<\/a>.<\/li>\n<li>[31] S. Ramzan, M. M. Iqbal, and T. Kalsum, (2022) \u201cText-to-image generation using deep learning\u201d Engineering Proceedings 20(1): 16. DOI: 10.3390\/engproc2022020016.<\/li>\n<\/ol>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,720,6],"tags":[1461],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202609_32.052\u00a0\u00a0 Download PDF Targeting the issues of weak fine-grained alignment capability and insufficient semantic&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/7025"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=7025"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=7025"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=7025"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}