Research
Source review in progress
AI Inference Optimization Runtimes & Serving Engines
vLLM, TensorRT-LLM, SGLang, PagedAttention, continuous batching, chunked prefill, and speculative decoding
This brief is being checked against individual publications, benchmarks and methods. Its earlier generated source list and confidence labels did not meet the publication standard. The topic remains available here while the evidence is reviewed.