A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with R

5 terms
  • Request coalescing

    Request coalescing is the serving concept concerned with request coalescing during AI inference.

    Inference performance
  • Request rate

    Request rate is the serving concept concerned with request rate during AI inference.

    Inference performance
  • Request scheduler

    Request scheduler is the serving concept concerned with request scheduler during AI inference.

    Inference performance
  • Response cache

    Response cache is the serving concept concerned with response cache during AI inference.

    Inference performance
  • Retry amplification

    Retry amplification is the serving concept concerned with retry amplification during AI inference.

    Inference performance