Find a term
Understand how AI systems are tested, scored, and compared. Learn to distinguish a useful evaluation result from a number that does not transfer to your workload.
Search results
98 termsAccuracy
Accuracy is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksAdversarial set
Adversarial set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksAggregate score
Aggregate score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark contamination
Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.
Evaluations and benchmarksBenchmark coverage
Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark leakage
Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark reliability
Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark saturation
Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark transfer
Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark validity
Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBERTScore
BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBias evaluation
Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBLEU score
BLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBootstrap confidence interval
Bootstrap confidence interval is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCalibration error
Calibration error is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCalibration set
Calibration set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksChallenge set
Challenge set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCitation evaluation
Citation evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCode-generation evaluation
Code-generation evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCodeBLEU score
CodeBLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCohen's kappa
Cohen's kappa is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCoherence evaluation
Coherence evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCompleteness evaluation
Completeness evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksComposite score
Composite score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksConfidence score
Confidence score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCost evaluation
Cost evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCounterfactual set
Counterfactual set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksCriterion-based evaluation
Criterion-based evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksData contamination
Data contamination is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksDevelopment set
Development set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEdge-case set
Edge-case set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksElo rating
Elo rating is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation case
Evaluation case is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation criterion
Evaluation criterion is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation dataset
An evaluation dataset is a collection of cases used to assess a system against defined criteria. For an AI application, it can include inputs, expected behavior, reference answers, grading information, and the context needed to reproduce each case.
Evaluations and benchmarksEvaluation harness
An evaluation harness is the software and configuration that runs an evaluation consistently. It typically loads cases, invokes a system, applies grading rules, records metrics, and produces results that can be compared across versions.
Evaluations and benchmarksEvaluation metric
An evaluation metric is a defined calculation used to summarize how a system performs against an evaluation criterion. The calculation can compare predictions with references, classify outcomes, measure latency or cost, or combine several signals.
Evaluations and benchmarksEvaluation objective
Evaluation objective is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation plan
Evaluation plan is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation question
Evaluation question is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation reproducibility
Evaluation reproducibility is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation scenario
Evaluation scenario is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation scope
Evaluation scope is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksEvaluation versioning
Evaluation versioning is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksExpected calibration error
Expected calibration error is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksF1 score
F1 score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksFactuality evaluation
Factuality evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksFormat-adherence evaluation
Format-adherence evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksImportance sampling
Importance sampling is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksInstruction-following evaluation
Instruction-following evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarks