A reference from Weave

Find a term

Understand how AI systems are tested, scored, and compared. Learn to distinguish a useful evaluation result from a number that does not transfer to your workload.

Terms beginning with E

15 terms
  • Edge-case set

    Edge-case set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Elo rating

    Elo rating is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation case

    Evaluation case is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation criterion

    Evaluation criterion is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation dataset

    An evaluation dataset is a collection of cases used to assess a system against defined criteria. For an AI application, it can include inputs, expected behavior, reference answers, grading information, and the context needed to reproduce each case.

    Evaluations and benchmarks
  • Evaluation harness

    An evaluation harness is the software and configuration that runs an evaluation consistently. It typically loads cases, invokes a system, applies grading rules, records metrics, and produces results that can be compared across versions.

    Evaluations and benchmarks
  • Evaluation metric

    An evaluation metric is a defined calculation used to summarize how a system performs against an evaluation criterion. The calculation can compare predictions with references, classify outcomes, measure latency or cost, or combine several signals.

    Evaluations and benchmarks
  • Evaluation objective

    Evaluation objective is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation plan

    Evaluation plan is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation question

    Evaluation question is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation reproducibility

    Evaluation reproducibility is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation scenario

    Evaluation scenario is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation scope

    Evaluation scope is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation versioning

    Evaluation versioning is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Expected calibration error

    Expected calibration error is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks