A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with D

4 terms
  • Data parallelism

    Data parallelism is the serving concept concerned with data parallelism during AI inference.

    Inference performance
  • Decode phase

    Decode phase is the serving concept concerned with decode phase during AI inference.

    Inference performance
  • Disaggregated serving

    Disaggregated serving is the serving concept concerned with disaggregated serving during AI inference.

    Inference performance
  • Dynamic batching

    Dynamic batching is the serving concept concerned with dynamic batching during AI inference.

    Inference performance