A Survey of Workloads for Large Language Model Inference Serving: Description, Public Datasets, and Evaluation Practice
A five-aspect framework for describing LLM inference-serving workloads, with a survey of public traces, load generators, and evaluation practice.
