基于自适应控制障碍函数的安全强化学习路径规划框架
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP24

基金项目:

国家自然科学基金项目(62273028).


An adaptive control barrier function-based safe reinforcement learning framework for path planning
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对强化学习路径规划中控制性能与安全性的权衡问题, 提出一种融合控制障碍函数的安全强化学习算法 RL + adpCBF. 首先, 在约束马尔可夫决策过程框架下, 引入控制障碍函数以提供安全保证, 并基于双延迟深度确定性策略梯度构建端到端的安全强化学习框架; 随后, 设计 CBFNet 网络结构, 在约束强化学习范式下加入基于控制障碍函数的安全约束, 并通过拉格朗日乘子法自适应平衡安全性与控制性能; 同时, 在满足安全约束的前提下, 将基于先验动力学与安全集构建的 CBF 约束实现为可微分的优化层, 使安全约束以显式凸优化的形式嵌入策略网络, 并通过端到端反向传播实现梯度更新; 最后, 在仿真环境中对所提出方法进行验证. 实验结果表明, RL + adpCBF 算法能够在保持较高控制性能的同时有效规避不安全动作, 对环境变化具备快速响应与实时策略调整能力, 能够显著提升移动机器人的安全性与运行效率.

    Abstract:

    This paper addresses the trade-off between control performance and safety in reinforcement learning (RL)-based path planning by proposing RL+adpCBF, a safe RL algorithm augmented with control barrier functions (CBFs). First, within the framework of constrained Markov decision processes (CMDPs), CBFs are incorporated to provide formal safety guarantees, and an end-to-end safe RL framework is constructed based on twin delayed deep deterministic policy gradient (TD3). Second, we design the CBFNet architecture, which introduces CBF-based safety constraints under the constrained RL paradigm and achieves adaptive safety-performance balancing via the Lagrangian multiplier method. Meanwhile, the CBF constraints derived from prior dynamics and safety sets are formulated as a differentiable optimization layer, embedding explicit convex safety constraints into the policy network and enabling end-to-end gradient backpropagation. Finally, the proposed method is validated in simulation. Experimental results demonstrate that the RL+adpCBF effectively avoids unsafe actions while maintaining high control performance, responds rapidly to environmental changes with real-time policy adjustments, and significantly improves the safety and operational efficiency of mobile robots.

    参考文献
    相似文献
    引证文献
引用本文

刘永康,张严心,柳向斌,等.基于自适应控制障碍函数的安全强化学习路径规划框架[J].控制与决策,2026,41(8):2345-2352

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-10-21
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-13
  • 出版日期:
文章二维码