Comparative Study of Exploration Strategies in Q- Learning for Multi-Controller SDN Load Balancing

Binod Sapkota Utkarsha Shukla Babu R. Dawadi Shashidhar R. Joshi

Журнал: International Journal of Wireless and Microwave Technologies @ijwmt

Статья в выпуске: 5 Vol.16, 2026 года.

Бесплатный доступ

Software Defined Networking enables centralized control and dynamic programmability, but achieving efficient load balancing across distributed controllers remains challenging because of traffic variability and scalability constraints. Despite the significant potential of reinforcement learning for multicontroller load balancing, existing studies primarily focus on developing new learning architectures or switch migration mechanisms, with limited attention given to systematically comparing exploration strategies under identical network conditions. The experimental setup consists of twelve switches for data forwarding and three distributed controllers following east-west communication. This study compares five exploration strategies of Q-learning using throughput, latency, controller response time, load balancing efficiency, fairness, scalability, and congestion-related metrics. Existing studies also often neglect fault tolerance, energy efficiency, and real-world validation. Under congested conditions, SOFT Q-learning records the lowest throughput reduction of 0.92, followed by upper confidence bound with 1.56, epsilon-greedy strategy with 2.34, prioritized experience replay with 3.29, and Boltzmann exploration with 6.06. The epsilon-greedy strategy achieves the highest throughput of 7.88 and fairness of 0.9999, soft reinforcement learning records the lowest latency of 0.0115 seconds, upper confidence bound achieves the fastest controller response time of 0.0285 seconds, and Boltzmann exploration attains the highest load balance ratio of 0.9865.

Boltzmann Exploration \ Multicontroller \ Load balancing \ Reinforcement learning \ Software Defined Networking

Короткий адрес: https://sciup.org/15020790

IDS: 15020790   |   DOI: 10.5815/ijwmt.2026.05.01