基于分层强化学习的有无人协同自适应决策权限分配算法
CSTR:
作者:
作者单位:

1.中国人民解放军国防科技大学;2.信息支援部队工程大学

作者简介:

通讯作者:

中图分类号:

TP181

基金项目:

国家社会科学基金军事学青年项目(2025-SKJJ-D-048)


Hierarchical Reinforcement Learning-based Adaptive Decision Authority Allocation Algorithm for Manned-Unmanned Teaming
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对有无人协同作战中决策权限静态分配难以适应动态战场态势的问题,提出一种基于分层强化学习的自适应决策权限分配算法(HAR). 算法采用双层决策架构:上层权限分配网络根据态势因子输出三级权限等级,下层动作策略网络在权限约束下执行任务. 通过危险评分偏置与指数移动平均平滑机制保障权限切换的稳定性,引入认知负荷指数量化人机协同效果,并采用三阶段渐进式课程学习提升复杂任务泛化能力. 仿真实验结果表明:HAR在困难任务阶段成功率达61%,优于MAPPO(58%)等基线算法;高权限使用比例从简单任务的14.87%提升至困难任务的69.82%,变化幅度达54.95个百分点;认知负荷变化幅度为0.45,显著优于对比算法,验证了所提方法能够有效实现决策权限与任务难度的动态匹配.

    Abstract:

    To address the challenge that static decision authority allocation in manned-unmanned teaming (MUM-T) cannot adapt to dynamic battlefield situations, a Hierarchical Adaptive authority allocation algorithm based on Reinforcement learning (HAR) is proposed. The algorithm adopts a two-layer architecture: the upper permission allocation network outputs three-level permission grades based on situational factors, while the lower action policy network executes tasks under permission constraints. Permission switching stability is ensured through danger score bias and exponential moving average smoothing mechanisms. A cognitive load index is introduced to quantify human-machine collaboration effectiveness, and a three-stage progressive curriculum learning framework is employed to enhance generalization in complex tasks. Simulation results demonstrate that HAR achieves a 61% task success rate in difficult tasks, outperforming MAPPO (58%) and other baseline algorithms. The high-permission usage ratio adaptively increases from 14.87% in simple tasks to 69.82% in difficult tasks, with a variation of 54.95 percentage points. The cognitive load variation reaches 0.45, significantly outperforming comparison algorithms, validating that the proposed method effectively achieves dynamic matching between decision authority and task difficulty.

    参考文献
    相似文献
    引证文献
引用本文
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-12-22
  • 最后修改日期:2026-06-26
  • 录用日期:2026-06-30
  • 在线发布日期: 2026-07-17
  • 出版日期:
文章二维码