Engineering & AI glossary
Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.
All terms
2,008 termsTest set
Test set is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksTest strategy
Test strategy is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtTest suite
Test suite is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtTest suite optimization
Test suite optimization is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtTest-first development
Test-first development is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtTestability
The ease with which software behavior can be isolated, stimulated, observed, and checked.
Code quality and technical debtTestability risk
The likelihood that important behavior is difficult to verify reliably before release.
Code quality and technical debtTestability-driven development
Testability-driven development is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtTheory of constraints
The theory of constraints focuses improvement on the system's current limiting constraint. The familiar cycle is to identify the constraint, use it effectively, align other work, elevate capacity when needed, and repeat when the constraint moves.
Flow and capacity planningThreshold alert
Threshold alert is an alert evaluated when a signal crosses a numeric boundary for a defined duration.
Reliability and observabilityThreshold metric
Threshold metric is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksThroughput
Throughput is the number of work items completed during a defined period. In software delivery, the item might be a pull request, deployed change, or customer request, and the chosen item boundary determines what the result means.
Flow and capacity planningThroughput forecast
Throughput forecast is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.
Flow and capacity planningThroughput latency tradeoff
Throughput latency tradeoff is the serving concept concerned with throughput latency tradeoff during AI inference.
Inference performanceThroughput per gpu
Throughput per gpu is the serving concept concerned with throughput per gpu during AI inference.
Inference performanceThroughput rate
Throughput rate is completed work divided by the interval in which it was completed. It is a rate version of throughput and requires a stable definition of item, completion, and time period.
Flow and capacity planningThroughput variability
Throughput variability is the spread of completed-item counts across time periods for a defined work population. It affects capacity planning and forecasting because the average rate does not describe every interval.
Flow and capacity planningTime to approval
Time to approval measures the elapsed period from a proposed change to the point at which required approval is recorded.
Code reviewTime to first byte
Time to first byte is the serving concept concerned with time to first byte during AI inference.
Inference performanceTime to first change
Time to first change is the elapsed time from a defined starting event, such as joining a team or creating a service, to a developer's first accepted code change in that environment. The start and completion events must be defined for the comparison to be meaningful.
Developer productivityTime to first review
Time to first review is the elapsed time between a reviewable change being submitted and the first substantive reviewer response.
Code reviewTime to first token
Time to first token, or TTFT, is the elapsed time from an inference request being accepted until the first output token is delivered. It captures startup and queue delay before generation becomes visible to a user.
Inference performanceTime to last token
Time to last token is the serving concept concerned with time to last token during AI inference.
Inference performanceTime to merge
Time to merge is the elapsed interval from a proposed change entering the workflow to its integration into the target branch.
Code reviewTime-based segment
Time-based segment is an analytical concept for separating engineering observations into populations whose differences may matter to a decision.
Engineering analyticsTime-window metric
Time-window metric groups observations within a defined period.
Engineering analyticsTimeout budget
Timeout budget is the serving concept concerned with timeout budget during AI inference.
Inference performanceTimeout policy
Timeout policy is a model-routing or gateway concept used to manage reliable request governance for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysTimestamp
Timestamp anchors ordering, duration, windows, and freshness.
Engineering analyticsTimestamp skew
Timestamp skew is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.
Engineering analyticsTimezone bias
Timezone bias is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.
Engineering analyticsToil ratio
Toil ratio is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.
DORA and DevOpsToken budget
A token budget is a configured limit or allowance for the tokens an AI system may process or generate during a request, task, or accounting period. The exact scope can refer to output length, context capacity, spending, or a workflow's total usage.
Token costs and AI ROIToken budget policy
Token budget policy is a model-routing or gateway concept used to manage policy enforcement for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysToken budget utilization
Token budget utilization is the percentage of a token budget consumed in a period. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken budget variance
Token budget variance is the difference between planned token consumption and actual consumption. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken embedding
Token embedding is a language-model concept about generation behavior and sampling. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsToken forecast
Token forecast is an estimate of future token consumption from workload history and planned demand. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken ledger
Token ledger is a durable accounting record of token usage and explanatory dimensions. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken meter
Token meter is a mechanism that records token usage for requests or workflows. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken overrun
Token overrun is usage that exceeds a configured allowance or expected request envelope. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken quota
Token quota is a token allowance assigned to a user, team, application, or account. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken rate
Token rate is the number of tokens processed per unit of time. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken streaming
Token streaming is the serving concept concerned with token streaming during AI inference.
Inference performanceToken utilization
Token utilization is how much of an available token allowance a workload consumes. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROIToken volume
Token volume is the amount of input and output text processed by an AI workload. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.
Token costs and AI ROITokenization
Tokenization is the process of converting text into the token units a language model receives and generates. A token can represent a word, part of a word, punctuation, or another piece of text, depending on the tokenizer.
LLM fundamentalsTokens per second
Tokens per second is the serving concept concerned with tokens per second during AI inference.
Inference performanceTool call normalization
Tool call normalization is a model-routing or gateway concept used to manage interface consistency for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysTool calling
Tool calling is an interface in which a language model returns a structured request for an application-defined function or external action. The application validates and executes the tool, then supplies the result back to the model.
AI coding and agents