多智能体强化学习的混合型移相变压器协同运行控制

辛锋, 刘漫雨, 崔高辰, 靳晓强, 刘若溪, 戴罕奇, 桑晓魁, 赵千川

清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (8) : 1715-1725.

PDF(2174 KB)
PDF(2174 KB)
清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (8) : 1715-1725. DOI: 10.16511/j.cnki.qhdxxb.2026.28.014
自动化

多智能体强化学习的混合型移相变压器协同运行控制

作者信息 +

Multiagent reinforcement learning-based coordinated control for hybrid transformers

Author information +
文章历史 +

摘要

晶闸管控制的混合型移相变压器(thyristor controlled hybrid transformer,TCHT)可实现配电网柔性互联,其兼备经济性和可靠性。然而在实际工程中,用户下达的功率传输指令需要进一步映射至TCHT装置补偿的电压矢量档位,而配电网络元件参数的缺失会导致映射出现偏差。采用已有的路径寻优法虽然针对单装置可以搜索到最优的补偿档位,但是在多装置的场景下由于电网潮流的耦合特性难以实现协同性。为此,该文提出了基于多智能体强化学习的协同运行控制方法,将多TCHT装置的协同控制建模为Markov决策过程,通过在线学习训练出具有协同性的多装置调控路径。根据数值实验结果,该文所提出的方法相比于独立的路径寻优法,平均能够降低36.6%的所需调节次数和58.9%的平均调节偏差。

Abstract

Objective: Flexible interconnection technology can realize asynchronous closed-loop operation among multiple alternating current distribution feeders. It can support dedicated power flow transfer, balance feeder loading, reduce network losses, and enable optimal allocation of controllable resources. The thyristor-controlled hybrid transformer (TCHT) can be used to realize flexible interconnection in distribution networks, offering good economic performance and high reliability. In device operation and control, the power-transfer command—whether obtained from an optimal power flow algorithm or specified manually based on engineering experience—must be converted into voltage compensation tap positions that are executable by the TCHT. However, in practical engineering, the accurate determination of the equivalent circuit parameters of lines, transformers, and other components is challenging. Existing studies have shown that a feedback mechanism can be introduced for a single TCHT, whereby the compensation voltage that minimizes power transfer error can be determined through a shortest-path search method. However, this method is difficult to extend to coordinate control when multiple TCHTs are coupled through the network topology. Methods: This study proposed a coordinated operation and control method based on multiagent reinforcement learning (RL). First, the coordinated control of multiple TCHTs was modeled as a Markov decision process (MDP). The system state space was defined as the local power regulation error of each TCHT. The action space was defined as the neighboring points of the current three-phase voltage compensation tap position of the TCHT, which avoids the convergence complexities of the training and control processes caused by an excessively large action space. The reward function was defined as the maximum regulation error. Consequently, the original problem was transformed into an equivalent MDP. Second, an online solution method based on multiagent RL was developed. A parameterized policy function was adopted to handle the continuous state space. Generalized advantage estimation was applied to reduce the approximation variance of the policy gradient, after which policy parameters were updated via the proximal policy optimization algorithm. Based on this, a coordinated policy training algorithm for multiple TCHTs was developed. Guided by the global reward function, this method enabled coordination among different TCHT devices. Results: Numerical analyses were performed on typical public network topologies. A flexible simulation platform was constructed for the interconnection distribution network based on Python and pandapower, with three homogeneous TCHT devices installed in the process. The results showed that the policy iteration process of the proposed method was stable, with only small fluctuations. The reward value increased significantly from iterations 0 to 50. From iterations 50 to 350, the value converged gradually to the optimum. From iterations 350 to 400, it remained basically stable around the optimum. Additionally, each decision step required only one forward pass of the neural network. On an Intel i7-13700 processor, each computation required only 10 ms on average, meeting real-time requirements. Compared with the independent shortest-path search method, the proposed method reduced the average regulation steps and error by 36.6% and 58.9%, respectively. Conclusions: These results show that although existing methods cannot coordinate multiple TCHTs within a common distribution network in the absence of power grid parameters, our algorithm significantly improves the accuracy and efficiency of coordinated control. Thus, this study fills the methodological gap in the collaborative control problem of multiple TCHTs.

关键词

柔性互联装置 / 多智能体强化学习 / Markov决策过程 / 运行控制

Key words

flexible interconnective device / multiagent reinforcement learning / Markov decision process / operational control

引用本文

导出引用
辛锋, 刘漫雨, 崔高辰, . 多智能体强化学习的混合型移相变压器协同运行控制[J]. 清华大学学报(自然科学版). 2026, 66(8): 1715-1725 https://doi.org/10.16511/j.cnki.qhdxxb.2026.28.014
Feng XIN, Manyu LIU, Gaochen Cui, et al. Multiagent reinforcement learning-based coordinated control for hybrid transformers[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(8): 1715-1725 https://doi.org/10.16511/j.cnki.qhdxxb.2026.28.014
中图分类号: TP272   

参考文献

1
WANG J, ZHOU N C, CHUNG C Y, et al. Coordinated planning of converter-based DG units and soft open points incorporating active management in unbalanced distribution networks[J]. IEEE Transactions on Sustainable Energy, 2020, 11(3): 2015- 2027.
2
贾焦心, 李豪迈, 邵晨, 等. 面向有源配电网的电磁式柔性互联装置研究进展[J/OL]. 电工技术学报, 1–22 (2025-06-11). [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381.
JIA J X, LI H M, SHAO C, et al. Research progress on electromagnetic flexible interconnection devices for active distribution networks[J/OL]. Transactions of China Electrotechnical Society, 1–22 (2025-06-11) [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381. (in Chinese)
3
颜湘武, 卢俊达, 贾焦心, 等. 面向配电台区经济性及综合承载力提升的旋转潮流控制器双层规划模型[J]. 电工技术学报, 2025, 40(7): 2078- 2094.
Yan X W, Lu J D, Jia J X, et al. A bi-levelprogramming model of rotary power flow controller for improving the economy and comprehensive carrying capacity of distribution station area[J]. Transactions of China Electrotechnical Society, 2025, 40(7): 2078- 2094.
4
葛少云, 杨赞, 刘洪, 等. 含四端SOP有源配电网可靠性和供电能力评估[J]. 高电压技术, 2020, 46(4): 1124- 1132.
GE S Y, YANG Z, LIU H, et al. Reliability and power supply capability evaluation of active distribution networks with four-terminal soft open points[J]. High Voltage Engineering, 2020, 46(4): 1124- 1132.
5
BA A O, PENG T, LEFEBVRE S. Rotary power-flow controller for dynamic performance evaluation—part Ⅱ: RPFC application in a transmission corridor[J]. IEEE Transactions on Power Delivery, 2009, 24(3): 1417- 1425.
6
谭振龙, 张春朋, 姜齐荣, 等. 旋转潮流控制器与统一潮流控制器和Sen Transformer的对比[J]. 电网技术, 2016, 40(3): 868- 874.
TAN Z L, ZHANG C P, JIANG Q R, et al. Comparative research on rotary power flow controller, unified power flow controller and Sen transformer[J]. Power System Technology, 2016, 40(3): 868- 874.
7
HUO Y D, LI P, JI H R, et al. Data-driven adaptive operation of soft open points in active distribution networks[J]. IEEE Transactions on Industrial Informatics, 2021, 17(12): 8230- 8242.
8
余梦泽, 李俭, 刘雷, 等. 电磁混合式统一潮流控制器的拓扑结构与控制策略优化[J]. 电工技术学报, 2015, 30(S2): 169- 175.
YU M Z, LI J, LIU L, et al. Topology structure and control strategy optimization of electromagnetic hybrid unified power flow controller[J]. Transactions of China Electrotechnical Society, 2015, 30(S2): 169- 175.
9
陈柏超, 刘雷, 余梦泽, 等. 电磁混合式潮流控制器本体优化及控制[J]. 高电压技术, 2017, 43(4): 1086- 1094.
CHEN B C, LIU L, YU M Z, et al. Ontology optimization and control strategy of electromagnetic hybrid power flow controller[J]. High Voltage Engineering, 2017, 43(4): 1086- 1094.
10
桑晓魁. 基于晶闸管多级调压的中压柔性合环控制技术[D]. 北京: 北京交通大学, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.
SANG X K. Medium voltage flexible loop closing control technology based on multi-stage voltage regulation with thyristor [D]. Beijing: Beijing Jiaotong University, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.(inChinese)
11
徐波. 基于SOP选址定容的柔性互联配电网调度方法研究[D]. 贵阳: 贵州大学, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.
XU B. Scheduling method for flexibly interconnected distribution networks based on soft open point siting and sizing[D]. Guiyang: Guizhou University, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.(inChinese)
12
FARIVAR M, LOW S H. Branch flow model: relaxations and convexification—Part Ⅰ[J]. IEEE Transactions on Power Systems, 2013, 28(3): 2554- 2564.
13
FARIVAR M, LOW S H. Branch flow model: relaxations and convexification—Part Ⅱ[J]. IEEE Transactions on Power Systems, 2013, 28(3): 2565- 2572.
14
魏炜, 季文文, 黄旭, 等. 含高比例分布式电源配电系统的参数辨识方法[J]. 电力建设, 2025, 46(05): 96- 111.
WEI W, JI W W, HUANG X, et al. Parameter Identification Method for Distribution Systems with a High Proportion of Distributed Renewable Energy[J]. Electric Power Construction, 2025, 46(05): 96- 111.
15
YU C, VELU A, VINITSKY E, et al. The surprising effectiveness of PPO in cooperative multi-agent games[C]// Proceedings of the 36th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2022: 1787.
16
梁志峰, 康重庆, 隋凌峰, 等. 含高比例分布式光伏的主配网运行风险评估与防控策略研究[J]. 清华大学学报(自然科学版), 2024, 64(11): 1964- 1978.
LIANG Z F, KANG C Q, SUI L F, et al. Research on risk warning and prevention strategies for main distribution networks with high proportion distributed photovoltaics[J]. Journal of Tsinghua University (Science and Technology), 2024, 64(11): 1964- 1978.
17
HU Q Y, YUE W Y. Markov decision processes with their applications[M]. New York: Springer, 2008.
18
LECUN Y, BENGIO Y, HINTON G. Deep learning[J]. Nature, 2015, 521(7553): 436- 444.
19
WILLIAMS R J. Simple statistical gradient-following algorithms for connectionist reinforcement learning[J]. Machine Learning, 1992, 8(3-4): 229- 256.
20
SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation[C]// Proceedings of the 4th International Conference on Learning Representations. San Juan, USA: ICLR, 2016.
21
KINGMA D P, BA J. Adam: A method for stochastic optimization[C]// Proceedings of the 3rd International Conference on Learning Representations. San Diego, USA: ICLR, 2015.
22
RUDION K, ORTHS A, STYCZYNSKI Z A, et al. Design of benchmark of medium voltage distribution network for investigation of DG integration[C]// Proceedings of 2006 IEEE Power Engineering Society General Meeting. Montreal, Canada: IEEE, 2006: 1–6.
23
VAN ROSSUM G, DRAKE F L. Python 3 reference manual[M]. Scotts Valley, USA: CreateSpace, 2009.
24
THURNER L, SCHEIDLER A, SCHÄFER F, et al. pandapower—an open-source Python tool for convenient modeling, analysis, and optimization of electric power systems[J]. IEEE Transactions on Power Systems, 2018, 33(6): 6510- 6521.
25
PASZKE A, GROSS S, MASSA F, et al. PyTorch: An imperative style, high-performance deep learning library[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc., 2019: 721.
26
PJM. Data miner 2: Hourly load: Metered [DB/OL]. PJM, 2019 [2025-11-10]. https://dataminer2.pjm.com.

基金

国网北京市电力公司项目“配电网多制式组网及柔性互联关键技术研究”(520223250003)

版权

版权所有,未经授权,不得转载。
PDF(2174 KB)

Accesses

Citation

Detail

段落导航
相关文章

/