Find a term
Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.
Terms beginning with D
4 termsData parallelism
Data parallelism is the serving concept concerned with data parallelism during AI inference.
Inference performanceDecode phase
Decode phase is the serving concept concerned with decode phase during AI inference.
Inference performanceDisaggregated serving
Disaggregated serving is the serving concept concerned with disaggregated serving during AI inference.
Inference performanceDynamic batching
Dynamic batching is the serving concept concerned with dynamic batching during AI inference.
Inference performance