Find a term
Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.
Terms beginning with L
4 termsLatency breakdown
Latency breakdown is the serving concept concerned with latency breakdown during AI inference.
Inference performanceLatency budget
Latency budget is the serving concept concerned with latency budget during AI inference.
Inference performanceLLM latency
LLM latency is the time associated with receiving a language model response. Common measures include time to first token, time between generated tokens, and time to the final token, each describing a different user experience.
Inference performanceLong context window
Long context window is the serving concept concerned with long context window during AI inference.
Inference performance