A reference from Weave

Find a term

Understand how AI systems are tested, scored, and compared. Learn to distinguish a useful evaluation result from a number that does not transfer to your workload.

Search results

98 terms
  • Accuracy

    Accuracy is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Adversarial set

    Adversarial set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Aggregate score

    Aggregate score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark contamination

    Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.

    Evaluations and benchmarks
  • Benchmark coverage

    Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark leakage

    Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark reliability

    Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark saturation

    Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark transfer

    Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark validity

    Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • BERTScore

    BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Bias evaluation

    Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • BLEU score

    BLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Bootstrap confidence interval

    Bootstrap confidence interval is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Calibration error

    Calibration error is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Calibration set

    Calibration set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Challenge set

    Challenge set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Citation evaluation

    Citation evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Code-generation evaluation

    Code-generation evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • CodeBLEU score

    CodeBLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Cohen's kappa

    Cohen's kappa is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Coherence evaluation

    Coherence evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Completeness evaluation

    Completeness evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Composite score

    Composite score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Confidence score

    Confidence score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Cost evaluation

    Cost evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Counterfactual set

    Counterfactual set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Criterion-based evaluation

    Criterion-based evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Data contamination

    Data contamination is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Development set

    Development set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Edge-case set

    Edge-case set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Elo rating

    Elo rating is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation case

    Evaluation case is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation criterion

    Evaluation criterion is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation dataset

    An evaluation dataset is a collection of cases used to assess a system against defined criteria. For an AI application, it can include inputs, expected behavior, reference answers, grading information, and the context needed to reproduce each case.

    Evaluations and benchmarks
  • Evaluation harness

    An evaluation harness is the software and configuration that runs an evaluation consistently. It typically loads cases, invokes a system, applies grading rules, records metrics, and produces results that can be compared across versions.

    Evaluations and benchmarks
  • Evaluation metric

    An evaluation metric is a defined calculation used to summarize how a system performs against an evaluation criterion. The calculation can compare predictions with references, classify outcomes, measure latency or cost, or combine several signals.

    Evaluations and benchmarks
  • Evaluation objective

    Evaluation objective is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation plan

    Evaluation plan is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation question

    Evaluation question is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation reproducibility

    Evaluation reproducibility is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation scenario

    Evaluation scenario is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation scope

    Evaluation scope is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Evaluation versioning

    Evaluation versioning is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Expected calibration error

    Expected calibration error is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • F1 score

    F1 score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Factuality evaluation

    Factuality evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Format-adherence evaluation

    Format-adherence evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Importance sampling

    Importance sampling is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Instruction-following evaluation

    Instruction-following evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks