基于双源不确定性的有模型离线强化学习
DOI:
CSTR:
作者:
作者单位:

中国矿业大学信息与控制工程学院

作者简介:

通讯作者:

中图分类号:

TP273

基金项目:

国家自然科学基金


Model-Based Offline Reinforcement Learning with Dual Sources of Uncertainty
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    有模型离线强化学习利用静态数据集学习动力学模型, 并通过模型生成样本辅助策略优化, 为真实世界中高风险、高成本的决策问题提供了可行的解决方案. 然而, 现有有模型离线强化学习方法面临双重挑战:静态数据集覆盖不足导致经验贝尔曼算子与真实贝尔曼算子之间存在认知误差; 动力学模型在多步预测中存在误差累积, 所生成的虚拟样本引入模型幻觉, 进一步加剧价值估计偏差. 针对上述问题, 本文提出一种基于双源不确定性的有模型离线Actor-Critic方法(Model-Based Offline Actor-Critic with Dual-Source Uncertainty, MDSU). 分别针对静态数据集和动力学模型构建不确定性量化器, 用以刻画两类数据源下的估计风险. 将不确定性量化器以惩罚形式引入静态贝尔曼目标与虚拟贝尔曼目标中, 并通过加权融合机制构造双源不确定性融合算子, 从而在策略评估阶段同时抑制两类误差的负面影响. 理论分析证明, 当静态数据不确定性量化器与动力学模型不确定性量化器能够有效约束相应认知误差时, MDSU的次优后悔具有显式上界. 在MuJoCo连续控制任务上的实验结果表明, MDSU在大多数任务上获得最高的归一化得分, 验证了该方法在性能提升与训练稳定性方面的有效性.

    Abstract:

    Model-based offline reinforcement learning learns a dynamics model from a static dataset and uses model-generated samples to assist policy optimization, providing a feasible solution for high-risk and high-cost decision-making problems in the real world. However, existing model-based offline reinforcement learning methods face two major challenges: insufficient coverage of the static dataset leads to cognitive error between the empirical Bellman operator and the true Bellman operator; meanwhile, error accumulation in multi-step predictions of the dynamics model causes generated virtual samples to introduce model hallucinations, further exacerbating bias in value estimation. To address these issues, this paper proposes a Model-Based Offline Deterministic Actor-Critic with Dual-Source Uncertainty(MDSU) method. Uncertainty quantifiers are constructed for the static dataset and the dynamics model, respectively, to characterize the estimation risks associated with the two data sources. The uncertainty quantifiers are incorporated as penalty terms into the static Bellman target and the virtual Bellman target, and a dual-source uncertainty fusion operator is constructed through a weighted fusion mechanism, thereby suppressing the negative effects of both types of errors during policy evaluation. Theoretical analysis proves that when the static-data uncertainty quantifier and the dynamics-model uncertainty quantifier can effectively constrain the corresponding cognitive error, MDSU has an explicit suboptimality regret bound. Experimental results on MuJoCo continuous-control tasks show that MDSU achieves the highest normalized scores on most tasks, verifying its effectiveness in improving performance and training stability.

    参考文献
    相似文献
    引证文献
引用本文
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-24
  • 最后修改日期:2026-07-23
  • 录用日期:2026-07-24
  • 在线发布日期:
  • 出版日期:
文章二维码