Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment

📄 arXiv: 2608.20834v1 📥 PDF

作者: Siqi Ding, Xuanhe Wang, Pei Guo, Guoyang Shi, Changquan Yu, Yiting Wang, Xianming Song, Xiang Gu, Zhengyuan Chen, Lei Xing, Yapeng Zhang, Jianguo Chen, Tianyuan Liu

分类: physics.plasm-ph, cs.AI

发布日期: 2026-08-21


💡 一句话要点

提出Advantage级聚合强化学习以解决X点目标磁配置控制问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 强化学习 多目标优化 等离子体控制 核聚变 X点目标 实验验证 控制系统

📋 核心要点

  1. 现有方法依赖于预计算的前馈波形和PID控制,缺乏针对次零点的专用闭环反馈,导致XPT操作的可重复性不足。
  2. 论文提出将XPT反馈视为多目标强化学习控制问题,设计Advantage聚合方法以保持目标特定的时间信用。
  3. 实验结果表明,AdvA-PPO在500毫秒的回放中将最差通道得分从0.23提升至0.81,X点通量均方根误差减少约20倍。

📝 摘要(中文)

管理紧凑高功率托卡马克的偏滤器热负荷是一个核心挑战。为提高局部通量扩展并将耗散体积与核心解耦,EHL-2采用了X点目标(XPT)偏滤器。该研究将XPT反馈建模为多目标强化学习控制问题,提出Advantage聚合(AdvA)方法,以应对等离子体电流、形状和零点约束之间的强耦合。通过在EXL-50U实验校准的自由边界环境中进行评估,AdvA-PPO在多个实验条件下显著提高了控制性能,为未来实时XPT验证提供了基础。

🔬 方法详解

问题定义:本研究旨在解决X点目标磁配置控制中的反馈闭环问题,现有方法在处理等离子体电流、形状和零点约束时存在强耦合,导致控制效果不理想。

核心思路:论文提出Advantage聚合(AdvA)方法,通过保持目标特定的时间信用,解决了奖励标量化在时间上的崩溃问题,从而提升了控制性能。

技术框架:整体架构包括多目标强化学习框架,AdvA模块用于奖励处理,PPO算法用于策略优化,实验环境则基于EXL-50U的自由边界校准。

关键创新:AdvA的核心创新在于引入了对最差目标的非线性标量化和残差修正,显著改善了策略更新的效果,与传统方法相比具有本质区别。

关键设计:在设计中,采用了特定的损失函数以平衡各目标的影响,网络结构经过优化以适应复杂的等离子体行为,确保在不同初始平衡状态下的有效性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,AdvA-PPO在500毫秒的回放中将最差通道得分从0.23提升至0.81,X点通量均方根误差减少约20倍,且在测量不确定性下,AdvA-PPO是唯一能够完成全时域操作的学习控制器。

🎯 应用场景

该研究的成果可广泛应用于高功率托卡马克的实时控制系统,特别是在偏滤器热负荷管理和等离子体稳定性控制方面。通过优化XPT操作,未来可能推动核聚变研究的进展,提升能源的可持续性和安全性。

📄 摘要(原文)

Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.