PDF(2174 KB)
Multiagent reinforcement learning-based coordinated control for hybrid transformers
Feng XIN, Manyu LIU, Gaochen Cui, Xiaoqiang JIN, Ruoxi LIU, Hanqi DAI, Xiaokui SANG, Qianchuan ZHAO
Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (8) : 1715-1725.
PDF(2174 KB)
PDF(2174 KB)
Multiagent reinforcement learning-based coordinated control for hybrid transformers
Objective: Flexible interconnection technology can realize asynchronous closed-loop operation among multiple alternating current distribution feeders. It can support dedicated power flow transfer, balance feeder loading, reduce network losses, and enable optimal allocation of controllable resources. The thyristor-controlled hybrid transformer (TCHT) can be used to realize flexible interconnection in distribution networks, offering good economic performance and high reliability. In device operation and control, the power-transfer command—whether obtained from an optimal power flow algorithm or specified manually based on engineering experience—must be converted into voltage compensation tap positions that are executable by the TCHT. However, in practical engineering, the accurate determination of the equivalent circuit parameters of lines, transformers, and other components is challenging. Existing studies have shown that a feedback mechanism can be introduced for a single TCHT, whereby the compensation voltage that minimizes power transfer error can be determined through a shortest-path search method. However, this method is difficult to extend to coordinate control when multiple TCHTs are coupled through the network topology. Methods: This study proposed a coordinated operation and control method based on multiagent reinforcement learning (RL). First, the coordinated control of multiple TCHTs was modeled as a Markov decision process (MDP). The system state space was defined as the local power regulation error of each TCHT. The action space was defined as the neighboring points of the current three-phase voltage compensation tap position of the TCHT, which avoids the convergence complexities of the training and control processes caused by an excessively large action space. The reward function was defined as the maximum regulation error. Consequently, the original problem was transformed into an equivalent MDP. Second, an online solution method based on multiagent RL was developed. A parameterized policy function was adopted to handle the continuous state space. Generalized advantage estimation was applied to reduce the approximation variance of the policy gradient, after which policy parameters were updated via the proximal policy optimization algorithm. Based on this, a coordinated policy training algorithm for multiple TCHTs was developed. Guided by the global reward function, this method enabled coordination among different TCHT devices. Results: Numerical analyses were performed on typical public network topologies. A flexible simulation platform was constructed for the interconnection distribution network based on Python and pandapower, with three homogeneous TCHT devices installed in the process. The results showed that the policy iteration process of the proposed method was stable, with only small fluctuations. The reward value increased significantly from iterations 0 to 50. From iterations 50 to 350, the value converged gradually to the optimum. From iterations 350 to 400, it remained basically stable around the optimum. Additionally, each decision step required only one forward pass of the neural network. On an Intel i7-13700 processor, each computation required only 10 ms on average, meeting real-time requirements. Compared with the independent shortest-path search method, the proposed method reduced the average regulation steps and error by 36.6% and 58.9%, respectively. Conclusions: These results show that although existing methods cannot coordinate multiple TCHTs within a common distribution network in the absence of power grid parameters, our algorithm significantly improves the accuracy and efficiency of coordinated control. Thus, this study fills the methodological gap in the collaborative control problem of multiple TCHTs.
flexible interconnective device / multiagent reinforcement learning / Markov decision process / operational control
| 1 |
|
| 2 |
贾焦心, 李豪迈, 邵晨, 等. 面向有源配电网的电磁式柔性互联装置研究进展[J/OL]. 电工技术学报, 1–22 (2025-06-11). [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381.
JIA J X, LI H M, SHAO C, et al. Research progress on electromagnetic flexible interconnection devices for active distribution networks[J/OL]. Transactions of China Electrotechnical Society, 1–22 (2025-06-11) [2025-11-03]. https://doi.org/10.19595/j.cnki.1000-6753.tces.250381. (in Chinese)
|
| 3 |
颜湘武, 卢俊达, 贾焦心, 等. 面向配电台区经济性及综合承载力提升的旋转潮流控制器双层规划模型[J]. 电工技术学报, 2025, 40(7): 2078- 2094.
|
| 4 |
葛少云, 杨赞, 刘洪, 等. 含四端SOP有源配电网可靠性和供电能力评估[J]. 高电压技术, 2020, 46(4): 1124- 1132.
|
| 5 |
|
| 6 |
谭振龙, 张春朋, 姜齐荣, 等. 旋转潮流控制器与统一潮流控制器和Sen Transformer的对比[J]. 电网技术, 2016, 40(3): 868- 874.
|
| 7 |
|
| 8 |
余梦泽, 李俭, 刘雷, 等. 电磁混合式统一潮流控制器的拓扑结构与控制策略优化[J]. 电工技术学报, 2015, 30(S2): 169- 175.
|
| 9 |
陈柏超, 刘雷, 余梦泽, 等. 电磁混合式潮流控制器本体优化及控制[J]. 高电压技术, 2017, 43(4): 1086- 1094.
|
| 10 |
桑晓魁. 基于晶闸管多级调压的中压柔性合环控制技术[D]. 北京: 北京交通大学, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.
SANG X K. Medium voltage flexible loop closing control technology based on multi-stage voltage regulation with thyristor [D]. Beijing: Beijing Jiaotong University, 2024. DOI: 10.26944/d.cnki.gbfju.2024.001451.(inChinese)
|
| 11 |
徐波. 基于SOP选址定容的柔性互联配电网调度方法研究[D]. 贵阳: 贵州大学, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.
XU B. Scheduling method for flexibly interconnected distribution networks based on soft open point siting and sizing[D]. Guiyang: Guizhou University, 2024. DOI:10.27047/d.cnki.ggudu.2024.000477.(inChinese)
|
| 12 |
|
| 13 |
|
| 14 |
魏炜, 季文文, 黄旭, 等. 含高比例分布式电源配电系统的参数辨识方法[J]. 电力建设, 2025, 46(05): 96- 111.
|
| 15 |
YU C, VELU A, VINITSKY E, et al. The surprising effectiveness of PPO in cooperative multi-agent games[C]// Proceedings of the 36th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2022: 1787.
|
| 16 |
梁志峰, 康重庆, 隋凌峰, 等. 含高比例分布式光伏的主配网运行风险评估与防控策略研究[J]. 清华大学学报(自然科学版), 2024, 64(11): 1964- 1978.
|
| 17 |
|
| 18 |
|
| 19 |
|
| 20 |
SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation[C]// Proceedings of the 4th International Conference on Learning Representations. San Juan, USA: ICLR, 2016.
|
| 21 |
KINGMA D P, BA J. Adam: A method for stochastic optimization[C]// Proceedings of the 3rd International Conference on Learning Representations. San Diego, USA: ICLR, 2015.
|
| 22 |
RUDION K, ORTHS A, STYCZYNSKI Z A, et al. Design of benchmark of medium voltage distribution network for investigation of DG integration[C]// Proceedings of 2006 IEEE Power Engineering Society General Meeting. Montreal, Canada: IEEE, 2006: 1–6.
|
| 23 |
|
| 24 |
|
| 25 |
PASZKE A, GROSS S, MASSA F, et al. PyTorch: An imperative style, high-performance deep learning library[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems. Red Hook, USA: Curran Associates Inc., 2019: 721.
|
| 26 |
PJM. Data miner 2: Hourly load: Metered [DB/OL]. PJM, 2019 [2025-11-10]. https://dataminer2.pjm.com.
|
/
| 〈 |
|
〉 |