source: arXiv Artificial Intelligence: A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing

level: research

Researchers analyzed a one-year production trace from Chutes, a platform serving many large language models. The study provides a global characterization and longitudinal view of real-world LLM serving workloads. Unlike prior work limited to short periods or few models, this trace captures full production behavior across popular and long-tail models. The analysis covers aggregate, temporal, model-level, and user-level perspectives, revealing how workloads evolve over time and how user-model interactions shape traffic.

The trace includes detailed request logs, model metadata, and user interactions. Key findings show significant workload evolution, with shifts in model popularity and request patterns over the year. Long-tail models receive sparse but persistent traffic, complicating caching strategies. Load-balancing must account for heterogeneous model sizes and varying request rates. The study quantifies cache hit rates and load distribution, providing concrete evidence for system design. Limitations include a single platform and potential biases in user behavior.

This work offers a valuable benchmark for LLM serving systems. The longitudinal data enables testing of caching and load-balancing algorithms under realistic conditions. Practitioners can use the trace to evaluate new serving architectures and improve resource utilization. The findings highlight the need for adaptive systems that handle both popular and niche models efficiently. The dataset is publicly available, supporting further research in AI infrastructure.

why it matters: Realistic workload traces help engineers design more efficient LLM serving systems, reducing costs and improving user experience.


source: arXiv Artificial Intelligence: A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing