A TD3-based Deep Reinforcement Learning algorithm for Voltage Regulation of Dual Active Bridge (DAB) DC-DC Converter under Triple Phase-Shift Modulation

K. Girinath Babu J. Sivavara Prasad V. Vasudevan

Журнал: International Journal of Engineering and Manufacturing @ijem

Статья в выпуске: 5 vol.16, 2026 года.

Бесплатный доступ

Owing to the bidirectional power transfer, galvanic isolation and high conversion efficiency, the Dual Active Bridge (DAB) converter has become a popular topology for bidirectional DC-DC power conversion. It applies to electric vehicles, battery energy storage systems and DC microgrids due to its properties. The non-linear input-output characteristics and parameter variations with operating conditions create a difficulty in controlling the output for accurate voltage regulation. A Deep Reinforcement Learning (DRL) control method is proposed to overcome the nonlinearity and voltage regulation issue of the DAB converter. Initially, a state space model and generalized averaging model are used to accurately model the converter dynamics. The control performance is explored in multiple phase-shift modulation techniques like Single Phase Shift (SPS), Extended Phase Shift (EPS), Dual Phase Shift (DPS) and Triple Phase Shift (TPS). A DDPG agent is first implemented for continuous phase-shift control; then, to improve the learning stability, convergence speed and voltage regulation performance, a TD3 agent is introduced. The performance of the proposed DRL is tested and compared to the Grey Wolf Optimizer tuned Proportional–Integral (GWO-PI) control and Model Predictive Control (MPC) for various input voltage and loading conditions. The simulation results show that the settling time of the proposed TD3 controller is 4.3ms, which is 35% lower than that of the MPC and GWO-PI controllers. The steady state voltage ripple of the proposed TD3 controller is reduced by 35% compared to the MPC and GWO-PI controllers. The voltage regulation accuracy of the proposed TD3 controller is better than that of the MPC and GWO-PI controllers.

Dual Active Bridge (DAB) converter \ Deep Reinforcement Learning (DRL) \ Adaptive Voltage Regulation \ Twin Delayed Deep Deterministic Policy Gradient (TD3) \ Deep Deterministic Policy Gradient (DDPG) \ Phase-Shift Modulation Techniques \ Model Predictive Contr

Короткий адрес: https://sciup.org/15020733

IDS: 15020733   |   DOI: 10.5815/ijem.2026.05.23