{"id":6883,"date":"2026-05-17T22:52:19","date_gmt":"2026-05-17T14:52:19","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=6883"},"modified":"2026-05-19T11:36:43","modified_gmt":"2026-05-19T03:36:43","slug":"jase-202609-32-043","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202609-32-043","title":{"rendered":"Quantitative Aesthetic Evaluation of Visual Artworks Using Vision Transformer with Multi-Dimensional Artistic Feature Fusion"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=3671\" data-type=\"page\" data-id=\"1055\">Volume 32<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-05-17T22:52:19+08:00\">2026-05-17<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Zeyu Gao<a href=\"mailto:396333836@qq.com\"><i class=\"fa fa-envelope\"><\/i><\/a> <\/p>\n\n\n\n<p style=\"font-size:14px\">School of Art, Bangkok Thonburi University, Bangkok, 10170 Thailand<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: April 8, 2026<br>Accepted:&nbsp;May 1, 2026<br>Publication Date:&nbsp;May 17, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/05\/32_043.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Qualitative&nbsp;Aesthetic&nbsp;Score&nbsp;Prediction&nbsp;on Sample Artworks&nbsp;<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a href=\"\/jase\/wp-content\/uploads\/2026\/05\/V32.0043.txt\" data-type=\"attachment\" data-id=\"6681\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202609_32.043\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202609_32.043<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/05\/043_2026_0784_V32.pdf\" data-type=\"attachment\" data-id=\"6911\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Automated quantitative aesthetic evaluation of visual artworks is a challenging cross-disciplinary task involving computer vision and art history. Traditional aesthetic assessment methods rely on handcrafted features or single-branch deep learning models, which fail to comprehensively capture the multi-faceted artistic attributes (e.g., color harmony, composition balance, texture, and semantic style) and long-range global dependencies critical to artistic appreciation. To address these limitations, this paper proposes a novel framework: Vision Transformer with Multi-Dimensional Artistic Feature Fusion (MDAF-ViT). Our model integrates a hierarchical Vision Transformer (ViT) backbone for global context modeling with multi-branch feature extractors to capture low-level visual attributes, mid-level compositional rules, and high-level semantic style features. A key innovation is the Dynamic Multi-Dimensional Attention Fusion (MDAF) module, which adaptively weights and fuses heterogeneous artistic features. Extensive experiments on standard art aesthetic datasets (BAID, APDDv2, JenAesthetics) demonstrate that MDAF-ViT significantly outperforms state-of-the-art CNN and ViT based methods, achieving superior performance in terms of Pearson Linear Correlation Coefficient (PLCC), Spearman Rank Correlation Coefficient (SRCC), and Mean Squared Error (MSE). This work provides a robust, interpretable foundation for large-scale digital art analysis and curation.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;Computational Aesthetics; Artwork Evaluation; Vision Transformer<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<ol>\n<li>[1] D. Jonauskaite, N. Dael, L. Baboulaz, L. Ch\u00e8vre, I. Cierny, N. Ducimeti\u00e8re, A. Fekete, P. Gabioud, H. Leder, M. Vetterli, et al., (2024) \u201cInteractive digital engagement with visual artworks and cultural artefacts enhances user aesthetic experiences in the laboratory and museum\u201d International Journal of Human\u2013Computer Interaction 40(6): 1369\u20131382. DOI: 10.1080\/10447318.2022.2143767.<\/li>\n<li>[2] E. Stamkou, D. Keltner, R. Corona, E. Aksoy, and A. S. Cowen, (2024) \u201cEmotional palette: A computational mapping of aesthetic experiences evoked by visual art\u201d Scientific Reports 14(1): 19932. DOI: 10.1038\/s41598-024-69686-9.<\/li>\n<li>[3] E. A. Vessel and H. Ovadia, (2025) \u201cThe role of the default mode network in aesthetic appeal\u201d Current Opinion in Behavioral Sciences 66: 101608. DOI: 10.1016\/j.cobeha.2025.101608.<\/li>\n<li>[4] T. Shi, C. Chen, X. Li, and A. Hao, (2024) \u201cSemantic and style based multiple reference learning for artistic and general image aesthetic assessment\u201d Neurocomputing 582: 127434. DOI: 10.1016\/j.neucom.2024.127434.<\/li>\n<li>[5] X. Zhang, Y. Xiao, J. Peng, X. Gao, and B. Hu, (2024) \u201cConfidence-based dynamic cross-modal memory network for image aesthetic assessment\u201d Pattern Recognition 149: 110227. DOI: 10.1016\/j.patcog.2023.110227.<\/li>\n<li>[6] X. Lu, Z. Lin, H. Jin, J. Yang, and J. Z. Wang, (2015) \u201cRating image aesthetics using deep learning\u201d IEEE Transactions on Multimedia 17(11): 2021\u20132034. DOI: 10.1109\/TMM.2015.2477040.<\/li>\n<li>[7] H. Jang and J.-S. Lee, (2021) \u201cAnalysis of deep features for image aesthetic assessment\u201d IEEE Access 9: 29850\u201329861. DOI: 10.1109\/ACCESS.2021.3060171.<\/li>\n<li>[8] J. McCormack and A. Lomas. \u201cUnderstanding aesthetic evaluation using deep learning\u201d. In: International conference on computational intelligence in music, sound, art and design (part of EvoStar). Springer. 2020, 118\u2013133. DOI: 10.1007\/978-3-030-43859-3_9.<\/li>\n<li>[9] X. Tian, Z. Dong, K. Yang, and T. Mei, (2015) \u201cQuery-dependent aesthetic model with deep learning for photo quality assessment\u201d IEEE Transactions on Multimedia 17(11): 2035\u20132048. DOI: 10.1109\/TMM.2015.2479916.<\/li>\n<li>[10] Y. Deng, C. C. Loy, and X. Tang, (2017) \u201cImage aesthetic assessment: An experimental survey\u201d IEEE Signal Processing Magazine 34(4): 80\u2013106. DOI: 10.1109\/MSP.2017.2696576.<\/li>\n<li>[11] Y. Ke, Y. Wang, K. Wang, F. Qin, J. Guo, and S. Yang, (2023) \u201cImage aesthetics assessment using composite features from transformer and CNN\u201d Multimedia Systems 29(5): 2483\u20132494. DOI: 10.1007\/s00530-023-01141-7.<\/li>\n<li>[12] M. Carrasco, C. Gonz\u00e1lez-Mart\u00edn, J. Aranda, and L. Oliveros, (2026) \u201cVision Transformer attention alignment with human visual perception in aesthetic object evaluation\u201d Plos one 21(4): e0344006. DOI: 10.1371\/journal.pone.0344006.<\/li>\n<li>[13] S. Li, H. Liang, M. Xie, and X. He. \u201cMulti-scale and multi-patch aggregation network based on dual-column vision fusion for image aesthetics assessment\u201d. In: 2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE. 2024, 1\u20136. DOI: 10.1109\/ICME57554.2024.10687850.<\/li>\n<li>[14] H. Takimoto, F. Omori, and A. Kanagawa, (2021) \u201cImage aesthetics assessment based on multi-stream CNN architecture and saliency features\u201d Applied Artificial Intelligence 35(1): 25\u201340. DOI: 10.1080\/08839514.2020.1839197.<\/li>\n<li>[15] D. Soydaner and J. Wagemans, (2024) \u201cMulti-task convolutional neural network for image aesthetic assessment\u201d Ieee Access 12: 4716\u20134729. DOI: 10.1109\/ACCESS.2024.3349961.<\/li>\n<li>[16] C.-H. Lee, J.-L. Shih, W.-L. Su, C.-C. Lien, and C.-C. Han. \u201cCombination of Global Maximum Pooling and Local Average Pooling for Unsupervised Fine-Grained Image Retrieval\u201d. In: Proceedings of the 2024 7th Artificial Intelligence and Cloud Computing Conference. 2024, 299\u2013307. DOI: 10.1145\/3719384.3719427.<\/li>\n<li>[17] J. Yao, J. Yao, Y. Yang, and C. Huang. \u201cHierarchical Adaptive Position Encoding-Based Transformer for Point Cloud Analysis\u201d. In: International Conference on Neural Information Processing. Springer. 2024, 197\u2013210. DOI: 10.1007\/978-981-96-6576-1_14.<\/li>\n<li>[18] R. Egele, J. Junior, J. CS, J. N. van Rijn, I. Guyon, X. Bar\u00f3, A. Clap\u00e9s, P. Balaprakash, S. Escalera, T. Moeslund, et al., (2024) \u201cAi competitions and benchmarks: Dataset development\u201d arXiv preprint arXiv:2404.09703: DOI: 10.48550\/arXiv.2404.09703.<\/li>\n<li>[19] U. Lee, Y. Son, J. Shin, G. Byun, Y. Lee, J. Koh, M. Jeon, and H. Kim. \u201cLLaVA-Docent-V2: Improving Data Quality and Pedagogical Data Generation to Train Large Multimodal Models for Art Appreciation Education\u201d. In: International Conference on Intelligent Tutoring Systems. Springer. 2025, 213\u2013228. DOI: 10.1007\/978-3-031-98284-2_17.<\/li>\n<li>[20] S. A. Amirshahi, G. U. Hayn-Leichsenring, J. Denzler, and C. Redies. \u201cJenaesthetics subjective dataset: analyzing paintings by subjective scores\u201d. In: European Conference on Computer Vision. Springer. 2014, 3\u201319. DOI: 10.1007\/978-3-319-16178-5_1.<\/li>\n<\/ol>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,720,6],"tags":[1452],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202609_32.043\u00a0\u00a0 Download PDF Automated quantitative aesthetic evaluation of visual artworks is a challenging cross-disciplinary&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/6883"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=6883"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=6883"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=6883"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}