Find a term
Understand service health, incidents, and the telemetry used to investigate production systems. Connect reliability outcomes to engineering decisions.
Terms beginning with E
7 termsEndpoint monitoring
Endpoint monitoring is measurement of availability, latency, correctness, and errors for a particular API endpoint.
Reliability and observabilityError budget
An error budget is the amount of unreliability permitted by a service level objective over its measurement window. If the objective is 99.9% availability, the budget is the remaining 0.1% of allowed unavailability under the defined measurement rules.
Reliability and observabilityError budget alert
Error budget alert is notification when the allowance for unsuccessful or slow service events reaches a threshold.
Reliability and observabilityError budget exhaustion
Error budget exhaustion is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityError budget reset
Error budget reset is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityEscalation policy
Escalation policy is the sequence by which an unresolved alert moves from an initial responder to backups.
Reliability and observabilityExemplar
Exemplar is a sample observation attached to an aggregated metric point, often with trace context.
Reliability and observability