Find a term
Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.
Terms beginning with M
11 termsMax concurrent requests
Max concurrent requests is the serving concept concerned with max concurrent requests during AI inference.
Inference performanceMax tokens
Max tokens is the serving concept concerned with max tokens during AI inference.
Inference performanceMemory bandwidth
Memory bandwidth is the serving concept concerned with memory bandwidth during AI inference.
Inference performanceMemory bound inference
Memory bound inference is the serving concept concerned with memory bound inference during AI inference.
Inference performanceMicrobatching
Microbatching is the serving concept concerned with microbatching during AI inference.
Inference performanceMixed precision inference
Mixed precision inference is the serving concept concerned with mixed precision inference during AI inference.
Inference performanceModel loading
Model loading is the serving concept concerned with model loading during AI inference.
Inference performanceModel parallelism
Model parallelism is the serving concept concerned with model parallelism during AI inference.
Inference performanceModel replica
Model replica is the serving concept concerned with model replica during AI inference.
Inference performanceModel replication
Model replication is the serving concept concerned with model replication during AI inference.
Inference performanceModel warmup
Model warmup is the serving concept concerned with model warmup during AI inference.
Inference performance