Engineering & AI glossary
Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.
All terms
2,008 termsGateway metrics
Gateway metrics is a model-routing or gateway concept used to manage interface consistency for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysGateway request lifecycle
The gateway request lifecycle is the ordered set of stages an AI request passes through before a response reaches the application. It commonly includes authentication, admission, policy evaluation, routing, provider execution, retries or fallback, response handling, and usage recording.
Model routing and gatewaysGauge metric
Gauge metric is a value that can rise or fall, such as queue depth, memory, or active connections.
Reliability and observabilityGeneralization
Generalization is a language-model concept about evaluation design and failure analysis. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.
LLM fundamentalsGitOps
GitOps is a release engineering and DevOps concept for controlling how software changes are prepared, introduced, or understood.
DORA and DevOpsGoal setting
Goal setting is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.
Developer productivityGolden path
Golden path is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.
Developer productivityGolden signals
The golden signals are latency, traffic, errors, and saturation. They are a monitoring framework that focuses attention on the most useful high-level indicators of service health from a user's and operator's perspective.
Reliability and observabilityGoodhart's law
Goodhart's law is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.
Engineering analyticsGPU memory utilization
GPU memory utilization is the serving concept concerned with gpu memory utilization during AI inference.
Inference performanceGraceful degradation
Graceful degradation is a model-routing or gateway concept used to manage traffic selection for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysGraceful shutdown
Graceful shutdown is a release engineering and DevOps concept for controlling how software changes are prepared, introduced, or understood.
DORA and DevOpsGradient accumulation
Gradient accumulation is a language-model concept about evaluation design and failure analysis. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.
LLM fundamentalsGradient clipping
Gradient clipping is a language-model concept about training behavior and measurement. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.
LLM fundamentalsGrammar-constrained generation
Grammar-constrained generation is a language-model concept about generation behavior and sampling. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsGraph of thoughts
Graph of thoughts is a language-model concept about representation and similarity. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsGray-box testing
Gray-box testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.
Code quality and technical debtGrouped-query attention
Grouped-query attention is a language-model concept about serving behavior and operational tradeoffs. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.
LLM fundamentalsHallucination
An AI hallucination is a generated statement or artifact that is unsupported, fabricated, or incorrect for the task and available evidence. A fluent answer can still contain hallucinations.
LLM fundamentalsHandoff count
Handoff count is a developer productivity concept that helps teams understand handoff count in the context of software delivery.
Developer productivityHandoff delay
Handoff delay is the elapsed time between one participant completing its part of work and the next participant beginning the corresponding action. It is a process delay, not necessarily a sign of individual inactivity.
Flow and capacity planningHawthorne effect
Hawthorne effect is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.
Measurement and experimentationHead-based sampling
Head-based sampling is a sampling decision made near the beginning of a trace before its outcome is known.
Reliability and observabilityHealth check
Health check is a request or probe reporting whether a service meets a defined operational condition.
Reliability and observabilityHealth signal
Health signal is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityHeterogeneous treatment effect
Heterogeneous treatment effect is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.
Measurement and experimentationHidden state
Hidden state is a language-model concept about representation and similarity. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsHistogram
Histogram is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.
Measurement and experimentationHistogram bucket
Histogram bucket is a count of observations at or below a defined boundary in a distribution.
Reliability and observabilityHistorical cycle time
Historical cycle time is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.
Flow and capacity planningHistorical throughput
Historical throughput is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.
Flow and capacity planningHotfix rate
Hotfix rate is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.
DORA and DevOpsHuman evaluation
Human evaluation is a structured review in which people rate or compare model outputs using defined criteria. It captures qualities such as usefulness, factuality, tone, and task fit that may be difficult to reduce to one automated score.
Measurement and experimentationHuman In The Loop Agent
Human In The Loop Agent is a software-engineering concept describing how an AI coding system, its tools, or human collaborators handle a defined task.
AI coding and agentsHypothesis multiplicity
Hypothesis multiplicity is an analytical risk or quality concern that can make an engineering analysis appear more certain, comparable, or causal than it is.
Engineering analyticsIdea to production time
Idea to production time is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.
DORA and DevOpsIdempotency key
Idempotency key is a model-routing or gateway concept used to manage reliable request governance for AI requests. It describes a distinct decision, control, interface, or observation point between an application and one or more model providers.
Model routing and gatewaysIdentity resolution
Identity resolution matches the same person, team, service, or item across sources.
Engineering analyticsIdle capacity
Idle capacity is unused capacity under a stated boundary and time window. It can be valuable slack that absorbs variation, or it can signal missing demand, a dependency, or a policy that prevents safe pulling.
Flow and capacity planningImmutable artifact
Immutable artifact is a release engineering and DevOps concept for controlling how software changes are prepared, introduced, or understood.
DORA and DevOpsImmutable infrastructure
Immutable infrastructure is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.
DORA and DevOpsImpact effort matrix
Impact effort matrix is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.
Flow and capacity planningImportance sampling
Importance sampling is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
Evaluations and benchmarksImprovement hypothesis
Improvement hypothesis is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.
Developer productivityImputation
Imputation is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.
Measurement and experimentationIn-context learning
In-context learning is a language-model concept about context selection and limits. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.
LLM fundamentalsInappropriate intimacy
A code smell in which two modules know too much about each other’s internal details.
Code quality and technical debtIncident alert
Incident alert is a notification that evidence suggests a disruption, degradation, or risk requiring coordinated attention.
Reliability and observabilityIncident bridge
Incident bridge is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident budget
Incident budget is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.
DORA and DevOps