Lili Li, Zhaoyang Liu, and Qi Yang
School of Mechanical Engineering, Shenyang Ligong University, Shenyang 110159, China.
Received: April 23, 2025
Accepted: May 16, 2025
Publication Date: June 15, 2026
A lightweight feature extraction network
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202610_33.002
This paper presents LiteFusion, a lightweight 6D object pose estimation algorithm that addresses the parameter redundancy and computational inefficiency issues of the classic DenseFusion framework in edge computing environments. For image feature extraction, a NAS-optimized EfficientNetV2-B1 is used as the backbone network, and an artifact-free upsampling strategy combining bilinear interpolation and 1 × 1 convolutions is adopted. This architecture reduces model parameters substantially while maintaining the ability to extract high frequency geometric features. To reduce redundancy in high-dimensional fused features, a shared bottleneck layer is designed for efficient feature compression, and a decoupled multi-task lightweight prediction head is implemented to decrease computational overhead and mitigate overfitting risks. Furthermore, an energy efficient training engine is proposed, which addresses VRAM and training efficiency bottlenecks through delayed BatchNorm freezing with gradient accumulation, automatic mixed precision (AMP), and asynchronous data flow optimization. Experiments are conducted on the LineMOD and YCB-Video datasets. Results show that LiteFusion achieves higher pose estimation accuracy than the original DenseFusion with a parameter count of 9.35M, a 14.7% increase in inference speed and a 35% improvement in training efficiency. It also demonstrates significantly enhanced accuracy for weakly textured objects, achieving a favorable balance between precision and edge deployment efficiency.
Keywords: 6Dobject pose estimation; DenseFusion; lightweight network; energy-efficient training
- [1] B. Wan and C. Zhang, (2024) “Fundamental Coordinate Space for Object 6D Pose Estimation” IEEE Access 12: 146430–146440. DOI: 10.1109/ACCESS.2024.3473936.
- [2] J. Tremblay, T. To, B. Sundaralingam, Y. Xiang, D. Fox, and S. Birchfield, (2018) “Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects” arXiv preprint arXiv:1809.10790: DOI: 10.48550/arXiv.1809.10790. arXiv: 1809.10790 [cs.RO].
- [3] E. Marchand, H. Uchiyama, and F. Spindler, (2015) “Pose estimation for augmented reality: a hands-on survey” IEEE Transactions on Visualization and Computer Graphics 22(12): 2633–2651. DOI: 10.1109/TVCG.2015.2513408.
- [4] A. Gouda, S. Awasthi, C. Blesing, L. Manohar, F. Hoffmann, and A. Kirchheim. “MR6D: Benchmarking 6D Pose Estimation for Mobile Robots”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. IEEE/CVF. IEEE, 2025, 2426–2434. DOI: 10.48550/arXiv.2508.13775. arXiv: 2508.13775 [cs.CV].
- [5] R. Mur-Artal and J. D. Tardós, (2017) “ORB-SLAM2: An open-source slam system for monocular, stereo, and rgb-d cameras” IEEE Transactions on Robotics 33(5): 1255–1262. DOI: 10.1109/TRO.2017.2705103.
- [6] C. Wang, D. Xu, Y. Zhu, R. Martín-Martín, C. Lu, L. Fei-Fei, and S. Savarese. “DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2019, 3343–3352. DOI: 10.1109/CVPR.2019.00346.
- [7] Y. He, W. Sun, H. Huang, J. Liu, H. Fan, and J. Sun. “PVN3D: A deep point-wise 3d keypoints voting network for 6dof pose estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2020, 11632–11641. DOI: 10.48550/arXiv.1911.04231. arXiv: 1911.04231 [cs.CV].
- [8] L. Zuo, L. Xie, H. Pan, and Z. Wang, (2022) “A Lightweight Two-End Feature Fusion Network for Object 6D Pose Estimation” Machines 10(4): 254. DOI: 10.3390/machines10040254.
- [9] B. Wen, W. Yang, J. Kautz, and S. Birchfield. “FoundationPose: Unified 6d pose estimation and tracking of novel objects”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2024, 17868–17879. DOI: 10.48550/arXiv.2312.08344. arXiv: 2312.08344 [cs.CV].
- [10] J. Lin, L. Liu, D. Lu, and K. Jia. “SAM-6D: Segment anything model meets zero-shot 6d object pose estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2024, 27906–27916. DOI: 10.48550/arXiv.2311.15707. arXiv: 2311.15707 [cs.CV].
- [11] J. Huang, H. Yu, K.-T. Yu, N. Navab, S. Ilic, and B. Busam. “MatchU: Matching unseen objects for 6d pose estimation from rgb-d images”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2024, 10095–10105. DOI: 10.1109/CVPR52733.2024.00962.
- [12] J. Corsetti, D. Boscaini, C. Oh, A. Cavallaro, and F. Poiesi. “Open-vocabulary object 6d pose estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2024, 18071–18080. DOI: 10.1109/CVPR52733.2024.01711.
- [13] M. Tan and Q. Le. “EfficientNetV2: Smaller models and faster training”. In: International Conference on Machine Learning. PMLR. PMLR, 2021, 10096–10106. DOI: 10.48550/arXiv.2104.00298. arXiv: 2104.00298 [cs.CV].
- [14] C. R. Qi, H. Su, K. Mo, and L. J. Guibas. “PointNet: Deep learning on point sets for 3d classification and segmentation”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE. IEEE, 2017, 652–660. DOI: 10.1109/CVPR.2017.16.
- [15] Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox, (2017) “PoseCNN: A convolutional neural network for 6d object pose estimation in cluttered scenes” arXiv preprint arXiv:1711.00199: DOI: 10.48550/arXiv.1711.00199. arXiv: 1711.00199 [cs.CV].
- [16] P. Besl and N. D. McKay, (1992) “A method for registration of 3-D shapes” IEEE Transactions on Pattern Analysis and Machine Intelligence 14(2): 239–256. DOI: 10.1109/34.121791.
- [17] Y. Xu, K.-Y. Lin, G. Zhang, X. Wang, and H. Li. “RNNPose: Recurrent 6-DoF Object Pose Refinement with Robust Correspondence Field Estimation and Pose Optimization”. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE/CVF. IEEE, 2022, 14860–14870. DOI: 10.1109/CVPR52688.2022.01446.
- [18] Y. He, Y. Wang, H. Fan, J. Sun, and Q. Chen. “FS6D: Few-shot 6d pose estimation of novel objects”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2022, 6814–6824. DOI: 10.1109/CVPR52688.2022.01446.
- [19] N. Pereira and L. A. Alexandre. “MaskedFusion: Mask-based 6D object pose estimation”. In: 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA). IEEE. IEEE, 2020, 71–78. DOI: 10.1109/ICMLA51294.2020.00021.
- [20] X. Jiang, D. Li, H. Chen, Y. Zheng, R. Zhao, and L. Wu. “Uni6D: A unified cnn framework without projection breakdown for 6d pose estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE/CVF. IEEE, 2022, 11174–11184. DOI: 10.48550/arXiv.2203.14531. arXiv: 2203.14531 [cs.CV].
- [21] D. Xu, D. Anguelov, and A. Jain. “PointFusion: Deep sensor fusion for 3d bounding box estimation”. In: Proceedings of the IEEE conference on computer vision and pattern recognition. IEEE. IEEE, 2018, 244–253. DOI: 10.1109/CVPR.2018.00033.
