A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with E

2 terms
  • End to end inference latency

    End to end inference latency is the serving concept concerned with end to end inference latency during AI inference.

    Inference performance
  • Expert parallelism

    Expert parallelism is the serving concept concerned with expert parallelism during AI inference.

    Inference performance