A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with L

4 terms
  • Latency breakdown

    Latency breakdown is the serving concept concerned with latency breakdown during AI inference.

    Inference performance
  • Latency budget

    Latency budget is the serving concept concerned with latency budget during AI inference.

    Inference performance
  • LLM latency

    LLM latency is the time associated with receiving a language model response. Common measures include time to first token, time between generated tokens, and time to the final token, each describing a different user experience.

    Inference performance
  • Long context window

    Long context window is the serving concept concerned with long context window during AI inference.

    Inference performance