A reference from Weave

Engineering & AI glossary

Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.

All terms

2,008 terms
  • Behavior-driven development

    Behavior-driven development is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Behavior-preserving change

    Behavior-preserving change is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt
  • Benchmark baseline

    Benchmark baseline is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark cohort

    Benchmark cohort is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark comparability

    Benchmark comparability is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark confidence

    Benchmark confidence is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark contamination

    Benchmark contamination occurs when evaluation examples, answers, or close duplicates appear in a model's training data or development process. The model may then recall the benchmark instead of demonstrating transferable capability.

    Evaluations and benchmarks
  • Benchmark coverage

    Benchmark coverage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark distribution

    Benchmark distribution is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark drift

    Benchmark drift is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark gaming

    Benchmark gaming is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark interpretation

    Benchmark interpretation is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark leakage

    Benchmark leakage is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark normalization

    Benchmark normalization is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark percentile

    Benchmark percentile is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark quality

    Benchmark quality is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark range

    Benchmark range is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark reliability

    Benchmark reliability is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark sample size

    Benchmark sample size is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark saturation

    Benchmark saturation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark target

    Benchmark target is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Benchmark transfer

    Benchmark transfer is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Benchmark validity

    Benchmark validity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • BERTScore

    BERTScore is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Beta testing

    Beta testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • BF16 inference

    BF16 inference is the serving concept concerned with bf16 inference during AI inference.

    Inference performance
  • Bias evaluation

    Bias evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Bidirectional encoder

    Bidirectional encoder is a language-model concept about serving behavior and operational tradeoffs. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Black-box monitoring

    Black-box monitoring is external evaluation of service behavior without relying on internal implementation knowledge.

    Reliability and observability
  • Black-box testing

    Black-box testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Blameless culture

    Blameless culture is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Blameless postmortem

    Blameless postmortem is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Blast radius

    Blast radius is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Blended AI rate

    Blended AI rate is an average rate combining models, token classes, or pricing tiers. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.

    Token costs and AI ROI
  • BLEU score

    BLEU score is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Blocked time

    Blocked time is the elapsed duration during which a work item is unable to make its next meaningful transition because an explicit obstacle prevents progress. It is a part of waiting time, not active implementation time.

    Flow and capacity planning
  • Blocked work

    Blocked work is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.

    Flow and capacity planning
  • Blocker

    Blocker is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.

    Flow and capacity planning
  • Blocker aging

    Blocker aging is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.

    Flow and capacity planning
  • Blocking reason

    A blocking reason is the classified cause attached to work that cannot proceed. Common classes include dependency, decision, environment, capacity, defect, and approval, but the useful taxonomy depends on the workflow.

    Flow and capacity planning
  • Blue-green deployment

    Blue-green deployment is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.

    DORA and DevOps
  • Blue-green routing

    Blue-green routing is a model-routing or gateway concept used to manage traffic selection for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.

    Model routing and gateways
  • Boolean parameter

    A function argument whose true or false value selects behavior or a mode.

    Code quality and technical debt
  • Bootstrap confidence interval

    Bootstrap confidence interval is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Bootstrap resampling

    Bootstrap resampling is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • BOS token

    BOS token is a language-model concept about instruction design and control. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Bottleneck

    A bottleneck is a workflow stage or capability whose effective capacity limits the rate at which the whole system can complete work. It is identified through sustained evidence of constrained flow, not simply because a stage feels busy.

    Flow and capacity planning
  • Boundary value analysis

    Boundary value analysis is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Box plot

    Box plot is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • Branch by abstraction

    Branch by abstraction is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt