Find a term
Understand how AI systems are tested, scored, and compared. Learn to distinguish a useful evaluation result from a number that does not transfer to your workload.
Terms beginning with B
11 termsBenchmark contamination
Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.
Evaluations and benchmarksBenchmark coverage
Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark leakage
Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark reliability
Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark saturation
Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark transfer
Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark validity
Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBERTScore
BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBias evaluation
Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBLEU score
BLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBootstrap confidence interval
Bootstrap confidence interval is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarks