Find a term
Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.
Terms beginning with B
82 termsBackdoor criterion
Backdoor criterion is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.
Measurement and experimentationBackfill
Backfill loads or recomputes historical records after a repair.
Engineering analyticsBacklog aging
Backlog aging is the elapsed time since a work item entered a backlog or became ready for consideration. It shows how long demand has waited before entering active delivery.
Flow and capacity planningBacklog health
Backlog health is an assessment of whether queued work has enough clarity, relevance, and prioritization to support reliable replenishment. It is a judgment supported by measures rather than a single universal score.
Flow and capacity planningBacklog refinement
Backlog refinement is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.
Flow and capacity planningBackpressure
Backpressure is a model-routing or gateway concept used to manage reliable request governance for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysBackup integrity
Backup integrity is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityBackup retention
Backup retention is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityBackward compatibility
Backward compatibility is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.
Code quality and technical debtBackward-compatible change
Backward-compatible change is a release engineering and DevOps concept for controlling how software changes are prepared, introduced, or understood.
DORA and DevOpsBaggage
Baggage is request-scoped key-value context propagated across service boundaries.
Reliability and observabilityBaseline regression
Baseline regression is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.
Engineering analyticsBatch inference
Batch inference processes multiple model inputs together in one serving operation. Grouping requests can improve hardware utilization, but it may add waiting time while a batch fills and must account for different input and output lengths.
Inference performanceBatch padding
Batch padding is the serving concept concerned with batch padding during AI inference.
Inference performanceBatch size
Batch size is the amount of work grouped into one processing, review, release, or handoff unit. The unit may be a change, pull request, story, deployment, or set of requests.
Flow and capacity planningBatch wait
Batch wait is the elapsed delay caused by holding work until a batch threshold, calendar window, or group of related items is ready. It is a queue effect created by batching policy.
Flow and capacity planningBeam search decoding
Beam search decoding is a language-model concept about instruction design and control. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsBehavior-driven development
Behavior-driven development is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtBehavior-preserving change
Behavior-preserving change is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.
Code quality and technical debtBenchmark baseline
Benchmark baseline is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark cohort
Benchmark cohort is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark comparability
Benchmark comparability is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark confidence
Benchmark confidence is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark contamination
Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.
Evaluations and benchmarksBenchmark coverage
Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark distribution
Benchmark distribution is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark drift
Benchmark drift is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark gaming
Benchmark gaming is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark interpretation
Benchmark interpretation is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark leakage
Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark normalization
Benchmark normalization is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark percentile
Benchmark percentile is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark quality
Benchmark quality is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark range
Benchmark range is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark reliability
Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark sample size
Benchmark sample size is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark saturation
Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark target
Benchmark target is an analytical concept for using reference values to understand engineering performance, variation, or capability.
Engineering analyticsBenchmark transfer
Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBenchmark validity
Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBERTScore
BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBeta testing
Beta testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtBF16 inference
BF16 inference is the serving concept concerned with bf16 inference during AI inference.
Inference performanceBias evaluation
Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksBidirectional encoder
Bidirectional encoder is a language-model concept about serving behavior and operational tradeoffs. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.
LLM fundamentalsBlack-box monitoring
Black-box monitoring is external evaluation of service behavior without relying on internal implementation knowledge.
Reliability and observabilityBlack-box testing
Black-box testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtBlameless culture
Blameless culture is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.
Developer productivityBlameless postmortem
Blameless postmortem is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.
Developer productivityBlast radius
Blast radius is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observability