Find a term
Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.
Terms beginning with I
10 termsInference admission
Inference admission is the serving concept concerned with inference admission during AI inference.
Inference performanceInference concurrency
Inference concurrency is the serving concept concerned with inference concurrency during AI inference.
Inference performanceInference load shedding
Inference load shedding is the serving concept concerned with inference load shedding during AI inference.
Inference performanceInference priority queue
Inference priority queue is the serving concept concerned with inference priority queue during AI inference.
Inference performanceInference queue time
Inference queue time is the serving concept concerned with inference queue time during AI inference.
Inference performanceInference throughput
Inference throughput is the amount of model inference work completed in a period of time. It may be expressed as requests per second, input tokens per second, output tokens per second, or another workload-specific measure.
Inference performanceInference token budget
Inference token budget is the serving concept concerned with inference token budget during AI inference.
Inference performanceINT4 inference
INT4 inference is the serving concept concerned with int4 inference during AI inference.
Inference performanceINT8 inference
INT8 inference is the serving concept concerned with int8 inference during AI inference.
Inference performanceINTer token latency
INTer token latency is the serving concept concerned with inter token latency during AI inference.
Inference performance