A reference from Weave

Reliability and observability glossary

Understand service health, incidents, and the telemetry used to investigate production systems. Connect reliability outcomes to engineering decisions.

All terms

140 terms
  • Admission control

    Admission control is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Alert deduplication

    Alert deduplication is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Alert evaluation window

    Alert evaluation window is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Alert fatigue

    Alert fatigue is the reduced ability or willingness to respond carefully to alerts after repeated exposure to notifications that are noisy, low priority, or rarely actionable.

    Reliability and observability
  • Alert inhibition

    Alert inhibition is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Alert policy

    Alert policy is documented conditions, grouping, suppression, ownership, and response expectations for alerts.

    Reliability and observability
  • Alert routing

    Alert routing is sending an alert to the team, channel, or escalation path responsible for response.

    Reliability and observability
  • Alert rule

    Alert rule is the query, condition, window, grouping, and action that create an alert.

    Reliability and observability
  • Alert suppression

    Alert suppression is the intentional prevention or grouping of notifications when known conditions make individual alerts redundant.

    Reliability and observability
  • Application monitoring

    Application monitoring is observation of service-level errors, latency, dependencies, and business operations.

    Reliability and observability
  • Auto-instrumentation

    Auto-instrumentation is automatic addition of standard telemetry to supported libraries or runtimes.

    Reliability and observability
  • Backup integrity

    Backup integrity is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Backup retention

    Backup retention is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Baggage

    Baggage is request-scoped key-value context propagated across service boundaries.

    Reliability and observability
  • Black-box monitoring

    Black-box monitoring is external evaluation of service behavior without relying on internal implementation knowledge.

    Reliability and observability
  • Blast radius

    Blast radius is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Burn-rate alert

    Burn-rate alert is detection of how quickly a service consumes error budget relative to its SLO window.

    Reliability and observability
  • Capacity headroom

    Capacity headroom is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Cardinality management

    Cardinality management is control of distinct attribute combinations so telemetry remains queryable and affordable.

    Reliability and observability
  • Circuit breaker state

    Circuit breaker state is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Common-mode failure

    Common-mode failure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Connection pool exhaustion

    Connection pool exhaustion is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Consistency check

    Consistency check is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Context propagation

    Context propagation is the transport of correlation information across processes, threads, services, and asynchronous work.

    Reliability and observability
  • Counter metric

    Counter metric is a value that increases as occurrences happen, such as requests, jobs, or errors.

    Reliability and observability
  • CPU throttling

    CPU throttling is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Critical path analysis

    Critical path analysis is identification of dependent work that determines when an operation can complete.

    Reliability and observability
  • Customer impact window

    Customer impact window is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Data reconciliation

    Data reconciliation is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dead letter queue

    Dead letter queue is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Degraded mode

    Degraded mode is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dependency contract

    Dependency contract is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dependency failure

    Dependency failure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dependency health

    Dependency health is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dependency isolation

    Dependency isolation is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Dependency map

    Dependency map is a record of components that rely on one another and the behavior of those relationships.

    Reliability and observability
  • Disk pressure

    Disk pressure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Distributed trace

    Distributed trace is a record following one operation across multiple processes or services with linked spans.

    Reliability and observability
  • Distributed tracing

    Distributed tracing is a method for recording the path and timing of a request as it moves through multiple services, processes, or other components. A trace groups related spans that describe individual operations along that path.

    Reliability and observability
  • Durable queue

    Durable queue is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Endpoint monitoring

    Endpoint monitoring is measurement of availability, latency, correctness, and errors for a particular API endpoint.

    Reliability and observability
  • Error budget

    An error budget is the amount of unreliability permitted by a service level objective over its measurement window. If the objective is 99.9% availability, the budget is the remaining 0.1% of allowed unavailability under the defined measurement rules.

    Reliability and observability
  • Error budget alert

    Error budget alert is notification when the allowance for unsuccessful or slow service events reaches a threshold.

    Reliability and observability
  • Error budget exhaustion

    Error budget exhaustion is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Error budget reset

    Error budget reset is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Escalation policy

    Escalation policy is the sequence by which an unresolved alert moves from an initial responder to backups.

    Reliability and observability
  • Exemplar

    Exemplar is a sample observation attached to an aggregated metric point, often with trace context.

    Reliability and observability
  • Failure domain

    Failure domain is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Failure injection

    Failure injection is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Game day

    Game day is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability