Process-Knowledge-Embedded Safe DRL for Real-Time Dispatch of Process Loads in Industrial Microgrids

📄 arXiv: 2608.03149v1 📥 PDF

作者: Daniyaer Paizulamu, Lin Cheng, Fashun Shi, Yuchi Zhang, Zhaoyang Dong

分类: eess.SY

发布日期: 2026-08-04

备注: 10 pages


💡 一句话要点

提出过程知识嵌入的安全深度强化学习以优化工业微电网调度

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 深度强化学习 工业微电网 过程知识 调度优化 可再生能源 安全性 成本降低

📋 核心要点

  1. 现有方法在多阶段耦合下难以平衡成本降低与生产过程可行性,造成决策的局限性。
  2. 提出了一种过程知识嵌入的安全深度强化学习框架,利用过程距离引导动作处理,确保调度的安全性与可行性。
  3. 通过实际数据的案例研究,验证了该方法在电力成本和过程损失方面的显著提升,具有良好的实用性。

📝 摘要(中文)

钢铁制造过程负载(SPLs)是灵活的资源,能够提升本地可再生能源的利用率并降低工业微电网的电力采购成本。然而,现有决策的多阶段耦合性使得传统深度强化学习在降低成本的同时难以保持生产过程的可行性。本文提出了一种过程知识嵌入的安全深度强化学习框架,用于实时调度工业微电网中的SPLs。通过构建无损的主动边界动作空间和基于过程距离的动作处理机制,确保了可接受的执行和可行的持续性。案例研究表明,相较于基于规则的调度和滚动MILP,该方法在可接受的计算时间内实现了零过程损失和电力成本分别降低49.2%和25.9%。

🔬 方法详解

问题定义:本文旨在解决工业微电网中钢铁制造过程负载调度的多阶段耦合问题,现有方法难以在降低成本的同时保持生产过程的可行性。

核心思路:提出的框架通过嵌入过程知识,构建无损的主动边界动作空间,并利用过程距离引导动作处理机制,确保调度的安全性和可行性。

技术框架:整体架构包括动作空间构建、过程距离引导的动作处理、递归过程可行性建立和基于PPO的过程修正距离内嵌等主要模块,形成一个闭环的调度优化系统。

关键创新:最重要的创新在于通过过程知识的嵌入和安全动作偏好调整,显著提升了调度的安全性和可行性,区别于传统的深度强化学习方法。

关键设计:设计中采用了过程距离作为动作处理的依据,结合修正预算和原始-对偶更新策略,确保了过程知识的有效内化,同时量化了原始策略对安全处理的依赖性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,所提方法在实际数据上实现了零过程损失,并在电力成本上相较于基于规则的调度和滚动MILP分别降低了49.2%和25.9%。这些结果展示了该方法在工业微电网调度中的显著优势。

🎯 应用场景

该研究的潜在应用领域包括工业微电网的实时调度、智能制造和能源管理系统。通过优化钢铁制造过程的调度策略,能够有效提升可再生能源的利用率,降低生产成本,具有重要的实际价值和广泛的应用前景。

📄 摘要(原文)

Steelmaking process loads (SPLs) are flexible resources that enhance local renewable-energy utilization and reduce electricity procurement costs in industrial microgrids. However, strong multistage coupling makes current decisions affect subsequent feasibility, challenging conventional deep reinforcement learning to reduce costs while maintaining process feasibility throughout production. This paper proposes a process-knowledge-embedded safe deep reinforcement learning framework for the real-time dispatch of SPLs in industrial microgrids. Specifically, a lossless active-frontier action space is constructed, and a process-distance-guided action-processing mechanism reallocates excluded-action probabilities according to process distance and the actor's safe-action preference. Recursive process feasibility is established to guarantee admissible execution and feasible continuation. Furthermore, the expected process-correction distance is incorporated into PPO through a correction budget and a primal-dual update to internalize process knowledge into the raw policy, while a derived bound quantifies the raw policy's dependence on safety processing. Case studies using real-world data demonstrate zero process losses, electricity-cost reductions of 49.2% and 25.9% relative to rule-based scheduling and rolling MILP, respectively, within an acceptable computation time.