Abstract:Frequent switching between task execution and recharging leads to control conflicts and local congestion in persistent-operation scenarios for laser-charged unmanned aerial vehicle (UAV) swarms. This paper proposes an improved multi-agent proximal policy optimization (MAPPO) method, termed MAPPO-RCY, to address these issues. Specifically, the method designed a role-split policy structure that explicitly separated the continuous control behaviors associated with charging and task execution, thereby alleviating low-level control conflicts. Furthermore, a dual-value-head critic mechanism evaluated returns for charging and task execution separately, which mitigated value-estimation coupling and improved the consistency between advantage estimation and role semantics under heterogeneous objectives. In addition, the method incorporated an active-yield bias at the execution stage to improve local passage conditions in high-density core areas. Simulation results showed that in a 15-UAV scenario, MAPPO-RCY achieved a task completion rate of 98.7% and outperformed MAPPO, HAPPO, HATRPO, IPPO, and Greedy. MAPPO-RCY reduced the average task response time to 2.5 s and task execution duration to 8.2 s, while it maintained a high average state of charge for the system. In the scalability test with 25 UAVs, the task completion rate remained at 97.9%. These results demonstrate that the proposed method improves the cooperative scheduling performance of laser-charged UAV swarms in high-density persistent-operation scenarios.