A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with I

10 terms
  • Inference admission

    Inference admission is the serving concept concerned with inference admission during AI inference.

    Inference performance
  • Inference concurrency

    Inference concurrency is the serving concept concerned with inference concurrency during AI inference.

    Inference performance
  • Inference load shedding

    Inference load shedding is the serving concept concerned with inference load shedding during AI inference.

    Inference performance
  • Inference priority queue

    Inference priority queue is the serving concept concerned with inference priority queue during AI inference.

    Inference performance
  • Inference queue time

    Inference queue time is the serving concept concerned with inference queue time during AI inference.

    Inference performance
  • Inference throughput

    Inference throughput is the amount of model inference work completed in a period of time. It may be expressed as requests per second, input tokens per second, output tokens per second, or another workload-specific measure.

    Inference performance
  • Inference token budget

    Inference token budget is the serving concept concerned with inference token budget during AI inference.

    Inference performance
  • INT4 inference

    INT4 inference is the serving concept concerned with int4 inference during AI inference.

    Inference performance
  • INT8 inference

    INT8 inference is the serving concept concerned with int8 inference during AI inference.

    Inference performance
  • INTer token latency

    INTer token latency is the serving concept concerned with inter token latency during AI inference.

    Inference performance