A reference from Weave

Find a term

Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.

Terms beginning with B

82 terms
  • Backdoor criterion

    Backdoor criterion is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • Backfill

    Backfill loads or recomputes historical records after a repair.

    Engineering analytics
  • Backlog aging

    Backlog aging is the elapsed time since a work item entered a backlog or became ready for consideration. It shows how long demand has waited before entering active delivery.

    Flow and capacity planning
  • Backlog health

    Backlog health is an assessment of whether queued work has enough clarity, relevance, and prioritization to support reliable replenishment. It is a judgment supported by measures rather than a single universal score.

    Flow and capacity planning
  • Backlog refinement

    Backlog refinement is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.

    Flow and capacity planning
  • Backpressure

    Backpressure is a model-routing or gateway concept used to manage reliable request governance for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.

    Model routing and gateways
  • Backup integrity

    Backup integrity is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Backup retention

    Backup retention is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Backward compatibility

    Backward compatibility is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt
  • Backward-compatible change

    Backward-compatible change is a release engineering and DevOps concept for controlling how software changes are prepared, introduced, or understood.

    DORA and DevOps
  • Baggage

    Baggage is request-scoped key-value context propagated across service boundaries.

    Reliability and observability
  • Baseline regression

    Baseline regression is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.

    Engineering analytics
  • Batch inference

    Batch inference processes multiple model inputs together in one serving operation. Grouping requests can improve hardware utilization, but it may add waiting time while a batch fills and must account for different input and output lengths.

    Inference performance
  • Batch padding

    Batch padding is the serving concept concerned with batch padding during AI inference.

    Inference performance
  • Batch size

    Batch size is the amount of work grouped into one processing, review, release, or handoff unit. The unit may be a change, pull request, story, deployment, or set of requests.

    Flow and capacity planning
  • Batch wait

    Batch wait is the elapsed delay caused by holding work until a batch threshold, calendar window, or group of related items is ready. It is a queue effect created by batching policy.

    Flow and capacity planning
  • Beam search decoding

    Beam search decoding is a language-model concept about instruction design and control. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Behavior-driven development

    Behavior-driven development is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Behavior-preserving change

    Behavior-preserving change is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt
  • Benchmark baseline

    Benchmark baseline is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark cohort

    Benchmark cohort is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark comparability

    Benchmark comparability is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark confidence

    Benchmark confidence is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark contamination

    Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.

    Evaluations and benchmarks
  • Benchmark coverage

    Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark distribution

    Benchmark distribution is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark drift

    Benchmark drift is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark gaming

    Benchmark gaming is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark interpretation

    Benchmark interpretation is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark leakage

    Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark normalization

    Benchmark normalization is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark percentile

    Benchmark percentile is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark quality

    Benchmark quality is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark range

    Benchmark range is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark reliability

    Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark sample size

    Benchmark sample size is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark saturation

    Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark target

    Benchmark target is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark transfer

    Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark validity

    Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • BERTScore

    BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Beta testing

    Beta testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • BF16 inference

    BF16 inference is the serving concept concerned with bf16 inference during AI inference.

    Inference performance
  • Bias evaluation

    Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Bidirectional encoder

    Bidirectional encoder is a language-model concept about serving behavior and operational tradeoffs. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Black-box monitoring

    Black-box monitoring is external evaluation of service behavior without relying on internal implementation knowledge.

    Reliability and observability
  • Black-box testing

    Black-box testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Blameless culture

    Blameless culture is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Blameless postmortem

    Blameless postmortem is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Blast radius

    Blast radius is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability