STEMM Institute Press
Science, Technology, Engineering, Management and Medicine
Research on Robust Reinforcement Learning Methods for Embodied Perception of Robotic Arms in Unstructured Environments
DOI: https://doi.org/10.62517/jes.202602307
Author(s)
Zhiyao Wang
Affiliation(s)
Hangzhou Dianzi University ITMO Joint Institute, Hangzhou Dianzi University, Hangzhou, China *Corresponding Author
Abstract
Embodied perception is a core capability for robotic arms to operate autonomously in unstructured environments. However, traditional methods struggle to effectively distinguish between lighting changes, sensor noise accumulation, and object occlusions due to complex disturbances, potentially leading to operational errors or task failure. To address these challenges, this paper proposes a robust optimization framework for embodied perception in robotic arms based on deep reinforcement learning, systematically evaluating the contributions of each module through ablation studies. The problem of perceptual robustness is modeled as a Markov Decision Process (MDP), with a multimodal attention fusion module designed to incorporate signal-to-noise ratio-aware perception, and contrastive learning employed to enhance feature discriminability. Perceptual confidence is introduced as a state variable, enabling a dynamic perception gain modulation mechanism that achieves end-to-end co-optimization of perception weights and action policies. Additionally, an intrinsic motivation exploration strategy based on prediction error is integrated to alleviate the challenge of sparse rewards in unstructured scenarios. A large-scale parallel training environment is built on the NVIDIA IsaacGym platform, using the UR5e robotic arm to generate diverse task scenarios involving extreme lighting conditions, texture variations, and random occlusions for grasping, placing, and assembly tasks. Ablation results demonstrate that the dynamic perception gain mechanism significantly improves success rates under both normal and strong interference conditions; the intrinsic motivation exploration strategy effectively enhances exploration efficiency in sparse-reward settings; and the multimodal attention fusion module provides a solid foundation for gain adjustment by estimating weight distributions. The complete framework outperforms standard baselines such as DDPG across all three task categories, maintaining low perceptual error levels even under extreme occlusion conditions.
Keywords
Unstructured Environment; Robotic Arm; Embodied Perception; Reinforcement Learning; Robustness Optimization
References
[1] Sun, C., Yuan, X., Wang, Y., & Liu, W. (2025). Embodied Intelligence-based Autonomous Unmanned System Technology. Journal of Automatica Sinica, 51(4), 762-777. [2] Zheng, N., Yang, M., Jiang, W., Sun, H., & Ding, N. (2025). Development Trends and Prospects of Embodied Intelligence. Engineering Sciences, 1-13. [3] Li, H., Chen, Y., Cui, W., Liu, W., Liu, K., Zhou, M., Zhang, Z., & Zhao, D. (2026). A Survey of Vision-Language-Action Models for Embodied Manipulation. Journal of Automatica Sinica, 52(1), 18-51. [4] Wang, N. (2023). Embodied Intelligent Visual Navigation Based on Entropy-Estimated Energy Models Daliann Maritime University]. [5] Mao, J., Wang, Z., Zhou, X., Xia, F., & Zhang, C. (2025). Design and Experimental Verification of a Dynamic Obstacle Avoidance Algorithm for Robotic Manipulators Based on Deep Reinforcement Learning. Experimental Technology and Management(4). [6] Shao, Y., Zhou, H., Fan, X., Xu, Q., & Liu, N. (2024). An Intelligent Control Algorithm for Single Robotic Manipulators Based on Deep Reinforcement Learning. Journal of Tianjin University of Technology, 40(6), 70-77. [7] Yuan, Q., Qi, J., & Yu, H. (2025). Intelligent Visual Servo Control of Robotic Manipulators Based on Deep Reinforcement Learning. Computer Integrated Manufacturing System(3). [8] You, X., Zhang, Y., Qiang, H., Xiang, G., Liao, Y., Guo, B., Zhong, Y., Fang, H., Zhao, T., & Dian, S. (2026). An Adaptive Prescribed Performance Control Method for Robotic Manipulators Based on Reinforcement Learning. [9] Hu, C., Du, B., Wang, Z., Zhang, Q., Zhang, R., & Gao, H. (2025). Perceptual Uncertainty-Aware Motion Planning for Autonomous Driving Based on Adaptive Heuristic Reinforcement Learning. Intelligent Transportation Systems, IEEE Transactions on, 26(8), 12676-12687. [10] Dong, J., Zhou, L., Sun, Q., Xu, J., & Tang, Y. (2026). Bio-Inspired Intelligent Robotics Technology: From Morphological Structure to Embodied Intelligence. Information and Control, 1-19. [11] Zou, Q., Tang, Y., Gao, B., Zhao, X., & Zhang, Z. (2025). Multi-agent Thinking Semi-multi-turn Communication Network Based on Deep Reinforcement Learning. Journal of Control Theory and Applications, 42(3), 553-562. [12] Xie, X., & Zhang, D. (2025). An Industrial Embodied Agent Demonstration Learning System Based on the Internet of Things. Robot Technique and Application(6), 53-56.
Copyright @ 2020-2035 STEMM Institute Press All Rights Reserved