Find a term
Understand service health, incidents, and the telemetry used to investigate production systems. Connect reliability outcomes to engineering decisions.
Search results
140 termsAdmission control
Admission control is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityAlert deduplication
Alert deduplication is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityAlert evaluation window
Alert evaluation window is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityAlert fatigue
Alert fatigue is the reduced ability or willingness to respond carefully to alerts after repeated exposure to notifications that are noisy, low priority, or rarely actionable.
Reliability and observabilityAlert inhibition
Alert inhibition is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityAlert policy
Alert policy is documented conditions, grouping, suppression, ownership, and response expectations for alerts.
Reliability and observabilityAlert routing
Alert routing is sending an alert to the team, channel, or escalation path responsible for response.
Reliability and observabilityAlert rule
Alert rule is the query, condition, window, grouping, and action that create an alert.
Reliability and observabilityAlert suppression
Alert suppression is the intentional prevention or grouping of notifications when known conditions make individual alerts redundant.
Reliability and observabilityApplication monitoring
Application monitoring is observation of service-level errors, latency, dependencies, and business operations.
Reliability and observabilityAuto-instrumentation
Auto-instrumentation is automatic addition of standard telemetry to supported libraries or runtimes.
Reliability and observabilityBackup integrity
Backup integrity is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityBackup retention
Backup retention is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityBaggage
Baggage is request-scoped key-value context propagated across service boundaries.
Reliability and observabilityBlack-box monitoring
Black-box monitoring is external evaluation of service behavior without relying on internal implementation knowledge.
Reliability and observabilityBlast radius
Blast radius is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityBurn-rate alert
Burn-rate alert is detection of how quickly a service consumes error budget relative to its SLO window.
Reliability and observabilityCapacity headroom
Capacity headroom is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityCardinality management
Cardinality management is control of distinct attribute combinations so telemetry remains queryable and affordable.
Reliability and observabilityCircuit breaker state
Circuit breaker state is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityCommon-mode failure
Common-mode failure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityConnection pool exhaustion
Connection pool exhaustion is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityConsistency check
Consistency check is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityContext propagation
Context propagation is the transport of correlation information across processes, threads, services, and asynchronous work.
Reliability and observabilityCounter metric
Counter metric is a value that increases as occurrences happen, such as requests, jobs, or errors.
Reliability and observabilityCPU throttling
CPU throttling is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityCritical path analysis
Critical path analysis is identification of dependent work that determines when an operation can complete.
Reliability and observabilityCustomer impact window
Customer impact window is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityData reconciliation
Data reconciliation is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDead letter queue
Dead letter queue is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDegraded mode
Degraded mode is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDependency contract
Dependency contract is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDependency failure
Dependency failure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDependency health
Dependency health is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDependency isolation
Dependency isolation is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDependency map
Dependency map is a record of components that rely on one another and the behavior of those relationships.
Reliability and observabilityDisk pressure
Disk pressure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityDistributed trace
Distributed trace is a record following one operation across multiple processes or services with linked spans.
Reliability and observabilityDistributed tracing
Distributed tracing is a method for recording the path and timing of a request as it moves through multiple services, processes, or other components. A trace groups related spans that describe individual operations along that path.
Reliability and observabilityDurable queue
Durable queue is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityEndpoint monitoring
Endpoint monitoring is measurement of availability, latency, correctness, and errors for a particular API endpoint.
Reliability and observabilityError budget
An error budget is the amount of unreliability permitted by a service level objective over its measurement window. If the objective is 99.9% availability, the budget is the remaining 0.1% of allowed unavailability under the defined measurement rules.
Reliability and observabilityError budget alert
Error budget alert is notification when the allowance for unsuccessful or slow service events reaches a threshold.
Reliability and observabilityError budget exhaustion
Error budget exhaustion is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityError budget reset
Error budget reset is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityEscalation policy
Escalation policy is the sequence by which an unresolved alert moves from an initial responder to backups.
Reliability and observabilityExemplar
Exemplar is a sample observation attached to an aggregated metric point, often with trace context.
Reliability and observabilityFailure domain
Failure domain is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityFailure injection
Failure injection is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityGame day
Game day is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observability