基于多智能体强化学习的综合能源微网供需协同优化调度
CSTR:
作者:
作者单位:

山东大学控制科学与工程学院

作者简介:

通讯作者:

中图分类号:

TK01;TP18

基金项目:

国家重点研发项目;国家自然科学基金项目


Multi-Agent Reinforcement Learning-Based Supply-Demand Coordinated Optimization for Integrated Energy Microgrid
Author:
Affiliation:

Fund Project:

National Key Research and Development Program of China; National Natural Science Foundation of China

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    激活需求侧温控负荷、电动汽车等柔性资源并推动供需双向互动,是提升综合能源微网经济性与灵活性的关键. 然而,传统集中式优化难以兼顾隐私保护与决策自主性,多智能体强化学习在复杂多能流耦合场景中又面临信息隔离导致的协同困难及物理硬约束保障不足等问题. 为此,本文提出一种改进的多智能体深度强化学习供需协同调度方法. 首先,构建供能侧与需求侧双智能体协同决策模型,在保持各主体自主决策的同时,实现系统经济性与用户舒适性的协同优化;其次,在集中训练、分散执行框架下,引入交叉注意力机制构建集中式价值网络,增强跨主体耦合信息感知与协同学习能力;最后,提出基于物理边界的安全动作投影机制,将策略输出约束于可行域,提高策略安全性与可实施性. 算例结果表明,所提方法能够降低运行成本、改善用户舒适度,并满足物理约束,在收敛性、经济性、舒适度和约束满足度方面优于基线方法.

    Abstract:

    Activating demand-side flexible resources, such as building thermal loads and electric vehicles, and promoting bidirectional supply-demand interaction are key to improving the economic performance and operational flexibility of integrated energy microgrids (IEMs). However, traditional centralized optimization methods struggle to balance privacy protection and decision-making autonomy, while existing multi-agent deep reinforcement learning (MADRL) methods face coordination difficulties caused by information isolation and insufficient guarantees of physical constraint satisfaction in complex multi-energy coupling scenarios. To address these challenges, an improved MADRL-based coordinated supply-demand dispatch method is proposed. First, a dual-agent collaborative decision-making model comprising supply-side and demand-side agents is established to coordinate system economy and user comfort while preserving the decision-making autonomy of each agent. Second, under the centralized training and decentralized execution framework, a cross-attention-based centralized critic network is developed to enhance the perception of cross-agent coupling features under information isolation and improve cooperative learning efficiency. Finally, a safe action projection mechanism based on system physical boundaries is introduced to constrain policy outputs within the feasible action set, thereby improving policy safety and practical implementability. Case studies demonstrate that the proposed method reduces system operating costs, improves user comfort, and satisfies physical constraints during both training and operation while maintaining decentralized decision-making autonomy. Compared with benchmark methods, it achieves better performance in terms of convergence, economic efficiency, user comfort, and constraint satisfaction.

    参考文献
    相似文献
    引证文献
引用本文
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-25
  • 最后修改日期:2026-06-17
  • 录用日期:2026-06-19
  • 在线发布日期: 2026-07-09
  • 出版日期:
文章二维码