A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with F

3 terms
  • Fair queuing

    Fair queuing is the serving concept concerned with fair queuing during AI inference.

    Inference performance
  • Flash attention

    Flash attention is the serving concept concerned with flash attention during AI inference.

    Inference performance
  • FP16 inference

    FP16 inference is the serving concept concerned with fp16 inference during AI inference.

    Inference performance