A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with W

3 terms
  • Warm start

    Warm start is the serving concept concerned with warm start during AI inference.

    Inference performance
  • Websocket streaming

    Websocket streaming is the serving concept concerned with websocket streaming during AI inference.

    Inference performance
  • Weight only quantization

    Weight only quantization is the serving concept concerned with weight only quantization during AI inference.

    Inference performance