Comparing Load Balancing Algorithms in AI Inference Services
DOI: https://doi.org/10.62517/jbdc.202601307
Author(s)
Yibo Li
Affiliation(s)
Singapore Institute of Management, Singapore
Abstract
The study focuses on load balancing in AI inference systems. The rapid growth of Artificial Intelligence (AI), in large language models (LLMs), has significantly impacted information systems. Classical algorithms (Round Robin, Least Outstanding Requests (LOR), and Power of Two Choices (P2C)) and AI-specific algorithms (Token-Aware and Prefix-Cache routing) are compared to determine their performance by applying a simulation-based approach, SimPy, under 25%, 60%, and 90% load conditions using response time, tail latency, throughput, and server utilization. Under low load conditions, all algorithms displayed similar results, while under moderate and high load conditions, AI-specific algorithms established remarkable results than classical algorithms.
Keywords
Load Balancing; AI Inference; LLMs; Token-Aware; Prefix-Cache
References
[1] Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Jin, H., Chen, T., & Jia, Z. (2025). Towards efficient generative large language model serving: A survey from algorithms to systems. ACM Computing Surveys, 58(1), 1-37.
[2] Dam, S. K., Hong, C. S., Qiao, Y., & Zhang, C. (2024). A complete survey on llm-based ai chatbots. arXiv preprint arXiv:2406.16937.
[3] Gupta, A., Joshi, N., Reddy, R., & Nair, R. (2021). Enhancing Conversational Commerce through AI: Leveraging Natural Language Processing and Reinforcement Learning in Chatbot Development. International Journal of AI Advancements, 10(1).
[4] Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., ... & Stoica, I. (2023). Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles (pp. 611-626).
[5] Al Nuaimi, K., Mohamed, N., Al Nuaimi, M., & Al-Jaroodi, J. (2012). A survey of load balancing in cloud computing: Challenges and algorithms. In 2012 second symposium on network cloud computing and applications (pp. 137-142). IEEE.
[6] Mitzenmacher, M. (2002). The power of two choices in randomized load balancing. IEEE transactions on parallel and distributed systems, 12(10), 1094-1104.
[7] Li, B., Jiang, Y., Gadepally, V., & Tiwari, D. (2024). Llm inference serving: Survey of recent advances and opportunities. In 2024 IEEE High Performance Extreme Computing Conference (HPEC) (pp. 1-8). IEEE.
[8] Sharma, P., & Bhattarai, S. (2026). A Review on Retrieval-Augmented Generation: Architectures, Research Challenges, and Emerging Frontiers. Journal of Future Artificial Intelligence and Technologies, 2(4), 616-628.
[9] Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., ... & Zhang, H. (2024). {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) (pp. 193-210).
[10] Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., ... & Sheng, Y. (2024). Sglang: Efficient execution of structured language model programs. Advances in neural information processing systems, 37, 62557-62583.
[11] SimPy Development Team (2024) SimPy (Version 4.1.1)
[12] Kumar, P., & Kumar, R. (2019). Issues and challenges of load balancing techniques in cloud computing: A survey. ACM computing surveys (CSUR), 51(6), 1-35.
[13] Amazon Web Services (2019) Application Load Balancer now supports Least Outstanding Requests algorithm for load balancing requests.
[14] Yu, G. I., Jeong, J. S., Kim, G. W., Kim, S., & Chun, B. G. (2022). Orca: A distributed serving system for {Transformer-Based} generative models. In 16th USENIX symposium on operating systems design and implementation (OSDI 22) (pp. 521-538).
[15] Srivatsa, V., He, Z., Abhyankar, R., Li, D., & Zhang, Y. (2024). Preble: Efficient distributed prompt scheduling for llm serving. arXiv preprint arXiv:2407.00023.
[16] Sun, B., Huang, Z., Zhao, H., Xiao, W., Zhang, X., Li, Y., & Lin, W. (2024). Llumnix: Dynamic scheduling for large language model serving. In 18th USENIX symposium on operating systems design and implementation (OSDI 24) (pp. 173-191).
[17] Kleinrock, L. (1975). Queuing systems, volume i: Theory.
[18] Matloff, N. (2008). Introduction to discrete-event simulation and the simpy language. Davis, CA. Dept of Computer Science. University of California at Davis. Retrieved on August, 2(2009), 1-33.
[19] Banks, J., Carson, J.S., Nelson, B.L. & Nicol, D.M. (2014) Discrete-Event System Simulation. 5th edition. Harlow: Pearson.