基于动态演化的机器人安全强化学习运动规划方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP183

基金项目:


Dynamic evolution-based safe reinforcement learning for robotic motion planning
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    基于安全强化学习的机器人运动规划方法在真实物理环境中展现出巨大潜力, 但在复杂风险敏感场景下, 现有方法的稳定性和适应性有限, 难以满足安全要求. 鉴于此, 研究一种具有时序动态演化特性的安全强化学习方法具有重要意义. 首先, 构建基于时间连续常微分方程架构的策略网络, 并引入融合欧拉离散化方法, 使策略由静态映射转变为可微分的动态系统, 从而提升智能体对非均匀时间序列的建模能力, 增强其在复杂环境中的自适应性; 其次, 通过自适应约束近端策略优化算法解决安全约束求解问题, 实现灵活且安全的机器人运动控制; 最后, 在机器人安全基准测试平台Safety-Gymnasium和自主搭建的Gazebo物理场景中进行实验测试. 测试结果表明: 基于动态演化的安全强化学习运动规划方法在满足约束的前提下, 能有效提升机器人在规划任务上的表现; 在物理场景下成功率较基线模型PPO-Lag提升了53.7%; 同时, 在不同难度的物理场景中, 所提方法的平均奖励和成功率最高, 实现了泛化场景的最优规划效果.

    Abstract:

    Safe reinforcement learning-based robot motion planning demonstrates significant potential in real-world physical environments. However, existing methods exhibit limited stability and adaptability in complex, risk-sensitive scenarios, often failing to meet safety requirements. To address this, we propose a safe reinforcement learning method with temporally dynamic evolution. First, we develop a policy network within a continuous-time ordinary differential equation framework, integrating Euler discretization to transform the policy from a static mapping into a differentiable dynamical system. This approach enhances the modeling of non-uniform temporal sequences, thereby improving adaptability in complex environments. Second, we introduce an adaptive constrained proximal policy optimization algorithm to solve safety-constrained optimization, enabling flexible and safe motion control. Finally, we evaluate the method on the Safety-Gymnasium benchmark and in custom Gazebo physical environments. Experimental results show that the proposed dynamically evolving safe reinforcement learning motion planning method outperforms existing methods, with a 53.7% higher success rate than the baseline PPO-Lag. It achieves the highest average return and success rate across environments of varying difficulty, demonstrating superior planning performance in generalized scenarios.

    参考文献
    相似文献
    引证文献
引用本文

刘京奇,韩奇松,胡春鹤,等.基于动态演化的机器人安全强化学习运动规划方法[J].控制与决策,2026,41(8):2332-2344

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-08-21
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-13
  • 出版日期:
文章二维码