PDF(2174 KB)
多智能体强化学习的混合型移相变压器协同运行控制
辛锋, 刘漫雨, 崔高辰, 靳晓强, 刘若溪, 戴罕奇, 桑晓魁, 赵千川
清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (8) : 1715-1725.
PDF(2174 KB)
PDF(2174 KB)
多智能体强化学习的混合型移相变压器协同运行控制
Multiagent reinforcement learning-based coordinated control for hybrid transformers
晶闸管控制的混合型移相变压器(thyristor controlled hybrid transformer,TCHT)可实现配电网柔性互联,其兼备经济性和可靠性。然而在实际工程中,用户下达的功率传输指令需要进一步映射至TCHT装置补偿的电压矢量档位,而配电网络元件参数的缺失会导致映射出现偏差。采用已有的路径寻优法虽然针对单装置可以搜索到最优的补偿档位,但是在多装置的场景下由于电网潮流的耦合特性难以实现协同性。为此,该文提出了基于多智能体强化学习的协同运行控制方法,将多TCHT装置的协同控制建模为Markov决策过程,通过在线学习训练出具有协同性的多装置调控路径。根据数值实验结果,该文所提出的方法相比于独立的路径寻优法,平均能够降低36.6%的所需调节次数和58.9%的平均调节偏差。
Objective: Flexible interconnection technology can realize asynchronous closed-loop operation among multiple alternating current distribution feeders. It can support dedicated power flow transfer, balance feeder loading, reduce network losses, and enable optimal allocation of controllable resources. The thyristor-controlled hybrid transformer (TCHT) can be used to realize flexible interconnection in distribution networks, offering good economic performance and high reliability. In device operation and control, the power-transfer command—whether obtained from an optimal power flow algorithm or specified manually based on engineering experience—must be converted into voltage compensation tap positions that are executable by the TCHT. However, in practical engineering, the accurate determination of the equivalent circuit parameters of lines, transformers, and other components is challenging. Existing studies have shown that a feedback mechanism can be introduced for a single TCHT, whereby the compensation voltage that minimizes power transfer error can be determined through a shortest-path search method. However, this method is difficult to extend to coordinate control when multiple TCHTs are coupled through the network topology. Methods: This study proposed a coordinated operation and control method based on multiagent reinforcement learning (RL). First, the coordinated control of multiple TCHTs was modeled as a Markov decision process (MDP). The system state space was defined as the local power regulation error of each TCHT. The action space was defined as the neighboring points of the current three-phase voltage compensation tap position of the TCHT, which avoids the convergence complexities of the training and control processes caused by an excessively large action space. The reward function was defined as the maximum regulation error. Consequently, the original problem was transformed into an equivalent MDP. Second, an online solution method based on multiagent RL was developed. A parameterized policy function was adopted to handle the continuous state space. Generalized advantage estimation was applied to reduce the approximation variance of the policy gradient, after which policy parameters were updated via the proximal policy optimization algorithm. Based on this, a coordinated policy training algorithm for multiple TCHTs was developed. Guided by the global reward function, this method enabled coordination among different TCHT devices. Results: Numerical analyses were performed on typical public network topologies. A flexible simulation platform was constructed for the interconnection distribution network based on Python and pandapower, with three homogeneous TCHT devices installed in the process. The results showed that the policy iteration process of the proposed method was stable, with only small fluctuations. The reward value increased significantly from iterations 0 to 50. From iterations 50 to 350, the value converged gradually to the optimum. From iterations 350 to 400, it remained basically stable around the optimum. Additionally, each decision step required only one forward pass of the neural network. On an Intel i7-13700 processor, each computation required only 10 ms on average, meeting real-time requirements. Compared with the independent shortest-path search method, the proposed method reduced the average regulation steps and error by 36.6% and 58.9%, respectively. Conclusions: These results show that although existing methods cannot coordinate multiple TCHTs within a common distribution network in the absence of power grid parameters, our algorithm significantly improves the accuracy and efficiency of coordinated control. Thus, this study fills the methodological gap in the collaborative control problem of multiple TCHTs.
柔性互联装置 / 多智能体强化学习 / Markov决策过程 / 运行控制
flexible interconnective device / multiagent reinforcement learning / Markov decision process / operational control
| 1 |
|
| 2 |
贾焦心, 李豪迈, 邵晨, 等. 面向有源配电网的电磁式柔性互联装置研究进展[J/OL]. 电工技术学报, 1–22 (2025-06-11). [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381.
JIA J X, LI H M, SHAO C, et al. Research progress on electromagnetic flexible interconnection devices for active distribution networks[J/OL]. Transactions of China Electrotechnical Society, 1–22 (2025-06-11) [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381. (in Chinese)
|
| 3 |
颜湘武, 卢俊达, 贾焦心, 等. 面向配电台区经济性及综合承载力提升的旋转潮流控制器双层规划模型[J]. 电工技术学报, 2025, 40(7): 2078- 2094.
|
| 4 |
葛少云, 杨赞, 刘洪, 等. 含四端SOP有源配电网可靠性和供电能力评估[J]. 高电压技术, 2020, 46(4): 1124- 1132.
|
| 5 |
|
| 6 |
谭振龙, 张春朋, 姜齐荣, 等. 旋转潮流控制器与统一潮流控制器和Sen Transformer的对比[J]. 电网技术, 2016, 40(3): 868- 874.
|
| 7 |
|
| 8 |
余梦泽, 李俭, 刘雷, 等. 电磁混合式统一潮流控制器的拓扑结构与控制策略优化[J]. 电工技术学报, 2015, 30(S2): 169- 175.
|
| 9 |
陈柏超, 刘雷, 余梦泽, 等. 电磁混合式潮流控制器本体优化及控制[J]. 高电压技术, 2017, 43(4): 1086- 1094.
|
| 10 |
桑晓魁. 基于晶闸管多级调压的中压柔性合环控制技术[D]. 北京: 北京交通大学, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.
SANG X K. Medium voltage flexible loop closing control technology based on multi-stage voltage regulation with thyristor [D]. Beijing: Beijing Jiaotong University, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.(inChinese)
|
| 11 |
徐波. 基于SOP选址定容的柔性互联配电网调度方法研究[D]. 贵阳: 贵州大学, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.
XU B. Scheduling method for flexibly interconnected distribution networks based on soft open point siting and sizing[D]. Guiyang: Guizhou University, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.(inChinese)
|
| 12 |
|
| 13 |
|
| 14 |
魏炜, 季文文, 黄旭, 等. 含高比例分布式电源配电系统的参数辨识方法[J]. 电力建设, 2025, 46(05): 96- 111.
|
| 15 |
YU C, VELU A, VINITSKY E, et al. The surprising effectiveness of PPO in cooperative multi-agent games[C]// Proceedings of the 36th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2022: 1787.
|
| 16 |
梁志峰, 康重庆, 隋凌峰, 等. 含高比例分布式光伏的主配网运行风险评估与防控策略研究[J]. 清华大学学报(自然科学版), 2024, 64(11): 1964- 1978.
|
| 17 |
|
| 18 |
|
| 19 |
|
| 20 |
SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation[C]// Proceedings of the 4th International Conference on Learning Representations. San Juan, USA: ICLR, 2016.
|
| 21 |
KINGMA D P, BA J. Adam: A method for stochastic optimization[C]// Proceedings of the 3rd International Conference on Learning Representations. San Diego, USA: ICLR, 2015.
|
| 22 |
RUDION K, ORTHS A, STYCZYNSKI Z A, et al. Design of benchmark of medium voltage distribution network for investigation of DG integration[C]// Proceedings of 2006 IEEE Power Engineering Society General Meeting. Montreal, Canada: IEEE, 2006: 1–6.
|
| 23 |
|
| 24 |
|
| 25 |
PASZKE A, GROSS S, MASSA F, et al. PyTorch: An imperative style, high-performance deep learning library[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc., 2019: 721.
|
| 26 |
PJM. Data miner 2: Hourly load: Metered [DB/OL]. PJM, 2019 [2025-11-10]. https://dataminer2.pjm.com.
|
/
| 〈 |
|
〉 |