A reference from Weave

Find a term

Explore latency, throughput, caching, and serving techniques. Understand which parts of a model request consume time and compute.

Terms beginning with M

11 terms
  • Max concurrent requests

    Max concurrent requests is the serving concept concerned with max concurrent requests during AI inference.

    Inference performance
  • Max tokens

    Max tokens is the serving concept concerned with max tokens during AI inference.

    Inference performance
  • Memory bandwidth

    Memory bandwidth is the serving concept concerned with memory bandwidth during AI inference.

    Inference performance
  • Memory bound inference

    Memory bound inference is the serving concept concerned with memory bound inference during AI inference.

    Inference performance
  • Microbatching

    Microbatching is the serving concept concerned with microbatching during AI inference.

    Inference performance
  • Mixed precision inference

    Mixed precision inference is the serving concept concerned with mixed precision inference during AI inference.

    Inference performance
  • Model loading

    Model loading is the serving concept concerned with model loading during AI inference.

    Inference performance
  • Model parallelism

    Model parallelism is the serving concept concerned with model parallelism during AI inference.

    Inference performance
  • Model replica

    Model replica is the serving concept concerned with model replica during AI inference.

    Inference performance
  • Model replication

    Model replication is the serving concept concerned with model replication during AI inference.

    Inference performance
  • Model warmup

    Model warmup is the serving concept concerned with model warmup during AI inference.

    Inference performance