Multi-Objective Deep Reinforcement Learning for Secure and Stable Power System Operation

📄 arXiv: 2608.20914v1 📥 PDF

作者: Ioannis Papadopoulos, Georgios Tsaousoglou, Johanna Vorwerk

分类: eess.SY

发布日期: 2026-08-21


💡 一句话要点

提出多目标深度强化学习以实现安全稳定的电力系统运行

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 深度强化学习 电力系统 多目标优化 热安全 阻尼特性 智能电网 实时控制

📋 核心要点

  1. 现有的电力系统控制方法通常只关注单一目标,无法有效应对多目标之间的权衡问题。
  2. 本文提出了一种多目标深度强化学习代理,旨在同时维持热安全和提高系统的阻尼特性。
  3. 实验结果表明,所提代理在多个操作目标之间取得了更好的平衡,显著提升了系统的稳定性。

📝 摘要(中文)

随着能源转型的推进,电力系统的稳定运行面临挑战,急需在不确定性下快速决策。尽管强化学习在电力系统控制中展现出潜力,但现有应用通常仅关注单一操作标准,如热安全或小信号稳定性。本文开发了一种统一控制的深度强化学习代理,能够在随机负载变化下维持热安全,同时引导系统朝向具有更好阻尼特性的操作点。与仅关注热安全的代理和常规政策相比,所提代理在多个操作目标之间实现了更好的平衡,显著提高了阻尼效果且几乎没有热安全违规。最后,研究展示了在小信号和大信号扰动下,提高关键阻尼的操作价值,表明更高的阻尼可加快振荡衰减并改善关键清除时间。

🔬 方法详解

问题定义:本文旨在解决电力系统在多目标操作中的稳定性问题,现有方法往往只关注单一目标,导致在面对复杂负载变化时的决策不足。

核心思路:通过设计一个统一控制的深度强化学习代理,本文实现了在维持热安全的同时,优化系统的阻尼特性,以应对多目标的挑战。

技术框架:该方法包括环境建模、代理设计、训练过程和评估阶段。代理通过与环境交互学习,优化决策策略以实现多目标平衡。

关键创新:本研究的创新在于提出了一个能够同时考虑热安全和阻尼特性的深度强化学习框架,突破了传统方法的单一目标限制。

关键设计:在网络结构上,采用了深度神经网络来处理复杂的输入状态,并设计了适应多目标的损失函数,以平衡不同目标的权重。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果显示,所提代理在热安全和阻尼特性方面均优于传统方法,热安全违规几乎为零,同时在小信号和大信号扰动下,系统的振荡衰减速度显著提高,关键清除时间也得到了改善。

🎯 应用场景

该研究的潜在应用领域包括电力系统的实时控制与优化,尤其是在可再生能源比例逐渐提高的背景下,能够有效提升电力系统的稳定性和安全性。未来,该方法有望在智能电网和微电网等新兴领域发挥重要作用。

📄 摘要(原文)

The ongoing energy transition challenges the stable operation of power systems and increases the need for rapid decision-making under uncertainty. While reinforcement learning has emerged as a promising framework for power system control and operation, existing applications typically focus on a single operational criterion, such as thermal security or small-signal stability. However, power system operation is inherently multi-objective and may involve trade-offs between objectives. This paper develops a unified-control deep reinforcement learning agent that maintains thermal security under stochastic load variations while steering the system toward operating points with improved damping of the most critical mode. Compared to a thermal-security-only agent and a business-as-usual policy, the proposed agent achieves a better balance among the operational objectives considered, with notably improved damping and negligible thermal-security violations. Finally, the operational value of increased critical damping is demonstrated under small- and large-signal disturbances, where operating points with higher damping lead to faster oscillation decay and improved critical clearing times.