School of Information Engineering, Zhengzhou University of Science and Technology, Zhengzhou 450064, China
Received: March 18, 2026
Accepted: May 3, 2026
Publication Date: July 18, 2026
Overall architecture of the proposed Lightweight Adaptive Transformer (LAT)
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202610_33.043
Image restoration is a fundamental task in computer vision, aiming to recover high quality images from degraded inputs. While Transformer-based methods have achieved state-of-the-art (SOTA) performance in this field, their heavy computational complexity and high memory consumption severely limit the deployment in real-time scenarios. To address the trade-off between restoration quality and inference efficiency, this paper proposes a lightweight adaptive Transformer (LAT) for real-time and high-quality image restoration. Specifically, we design a novel adaptive depthwise separable attention (ADSA) mechanism, which replaces the full self attention in traditional Transformers with depthwise separable convolutions and adaptive gating, reducing computational complexity fromO (N2) to O(N) (where N is the number of image tokens). Additionally, a dynamic feature fusion (DFF) module is introduced to adaptively integrate multi-scale features, enhancing the restoration of fine-grained details while maintaining model lightness. Furthermore, we propose a hybrid loss function (HLF) that combines perceptual loss, L1 loss, and adversarial loss, balancing objective accuracy and subjective visual quality. Extensive experiments are conducted on five benchmark datasets (DIV2K, Set14, BSD100, Urban100, and RealSR) for super-resolution, denoising, and deblurring tasks. Results demonstrate that the proposed LAT achieves SOTA restoration quality (e.g., PSNR of 38.21 dB and SSIM of 0.961 on DIV2K for 4× super-resolution) while reducing the model parameters by 72.3% and inference time by 68.5% compared to existing Transformer-based methods. Ablation studies verify the effectiveness of each proposed module, and real-world tests confirm its applicability in real-time scenarios.
Keywords: Image Restoration; Lightweight Transformer; Adaptive Attention; Real-Time Inference; Feature Fusion
- [1] M. V. Conde, G. Geigle, and R. Timofte. “Instructir: High-quality image restoration following human instructions”. In: European Conference on Computer Vision. Springer. 2024, 1–21. DOI: 10.1007/978-3-031-72764-1_1.
- [2] Y. Cui, W. Ren, X. Cao, and A. Knoll, (2024) “Revitalizing convolutional network for image restoration” IEEE Transactions on Pattern Analysis and Machine Intelligence 46(12): 9423–9438. DOI: 10.1109/TPAMI.2024.3419007.
- [3] C. Zhao, H. Li, and M. Jiang, “Layer Priors And Encoding-decoding Network For Image Dehazing” Journal of Applied Science and Engineering 29(6): 1391–1398. DOI: 10.6180/jase.202606_29(6).0007.
- [4] S. Yin, L. Wang, A. A. Laghari, L. Teng, G. Srivastava, A. Almadhor, and T. R. Gadekallu, (2026) “FGM-MLSD: A Fuzzy Region Competition and Gaussian Mixture Segment Model Via Modified Line Segment Detector Model for Airport Object Saliency Detection in Remote Sensing Images” IEEE Transactions on Fuzzy Systems 34(4): 1175–1186. DOI: 10.1109/TFUZZ.2026.3650858.
- [5] Y. Cui, M. Liu, W. Ren, and A. Knoll, (2025) “Modumer: Modulating transformer for image restoration” IEEE Transactions on Neural Networks and Learning Systems 36(9): 7099–17113. DOI: 10.1109/TNNLS.2025.3561924.
- [6] H. Liu and L. Li. “Image restoration employing cross-ViT combined generative adversarial networks”. In: Fourth International Conference on Image Processing and Intelligent Control (IPIC 2024). 13250. SPIE. 2024, 156–161. DOI: 10.1117/12.3038514.
- [7] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte. “Swinir: Image restoration using swin transformer”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2021, 1833–1844. DOI: 10.1109/ICCVW54120.2021.00210.
- [8] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang. “Restormer: Efficient transformer for high-resolution image restoration”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022, 5728–5739. DOI: 10.1109/CVPR52688.2022.00564.
- [9] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li. “Uformer: A general u-shaped transformer for image restoration”. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022, 17683–17693. DOI: 10.1109/CVPR52688.2022.01716.
- [10] W. Peebles and S. Xie. “Scalable diffusion models with transformers”. In: Proceedings of the IEEE/CVF international conference on computer vision. 2023, 4195–4205. DOI: 10.1109/ICCV51070.2023.00387.
- [11] K. Wu, J. Zhang, H. Peng, M. Liu, B. Xiao, J. Fu, and L. Yuan. “Tinyvit: Fast pretraining distillation for small vision transformers”. In: European conference on computer vision. Springer. 2022, 68–85. DOI: 10.1007/978-3-031-19803-8_5.
- [12] H. Choi, C. Na, J. Oh, S. Lee, J. Kim, S. Choe, J. Lee, T. Kim, and J. Yang. “Reciprocal attention mixing transformer for lightweight image restoration”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 5992–6002. DOI: 10.1109/CVPRW63382.2024.00606.
- [13] S. Yin, L. Wang, T. Chen, H. Huang, J. Gao, J. Zhang, M. Liu, P. Li, and C. Xu, (2026) “LKAFormer: A lightweight kolmogorov-arnold transformer model for image semantic segmentation” ACM Transactions on Intelligent Systems and Technology 17(3): 1–24. DOI: 10.1145/3759254.
- [14] L. Wang, H. Wang, S. Yin, and L. Wang, (2025) “Masked vision transformer for fast hyperspectral image classification” IEEE Transactions on Geoscience and Remote Sensing 63: DOI: 10.1109/TGRS.2025.3572242.
- [15] X. Huo, G. Sun, S. Tian, Y. Wang, L. Yu, J. Long, W. Zhang, and A. Li, (2024) “HiFuse: Hierarchical multiscale feature fusion network for medical image classification” Biomedical signal processing and control 87: 105534. DOI: 10.1016/j.bspc.2023.105534.
- [16] Q. Yang, P. Yan, Y. Zhang, H. Yu, Y. Shi, X. Mou, M. K. Kalra, Y. Zhang, L. Sun, and G. Wang, (2018) “Low-dose CT image denoising using a generative adversarial network with Wasserstein distance and perceptual loss” IEEE transactions on medical imaging 37(6): 1348–1357. DOI: 10.1109/TMI.2018.2827462.
- [17] J. Qiao, M. Cai, W. Li, Y. Liu, X. Huang, G. He, J. Xie, J. Hu, X. Chen, and S. Lin, (2025) “RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought” arXiv preprint arXiv:2506.16796: DOI: 10.48550/arXiv.2506.16796.
- [18] Y. Liu, S. Li, L. Zhou, H. Liu, and Z. Li, (2025) “Dark-Yolo: A low-light object detection algorithm integrating multiple attention mechanisms” Applied Sciences 15(9): 5170. DOI: 10.3390/app15095170.
