Find a term
Understand service health, incidents, and the telemetry used to investigate production systems. Connect reliability outcomes to engineering decisions.
Terms beginning with I
18 termsIncident alert
Incident alert is a notification that evidence suggests a disruption, degradation, or risk requiring coordinated attention.
Reliability and observabilityIncident bridge
Incident bridge is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident channel
Incident channel is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident closure
Incident closure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident commander handoff
Incident commander handoff is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident communications plan
Incident communications plan is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident detection latency
Incident detection latency is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident handoff
Incident handoff is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident impact assessment
Incident impact assessment is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident recovery time
Incident recovery time is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident reopen
Incident reopen is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident response
Incident response is the coordinated process of detecting, assessing, containing, communicating about, and recovering from an event that threatens a service or users. It includes the operational actions during the event and the learning work that follows.
Reliability and observabilityIncident response time
Incident response time is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident severity
Incident severity is a classification of the impact, urgency, and scope of a service incident. A severity level guides response priorities and communication; it is not a measure of personal fault.
Reliability and observabilityIncident severity matrix
Incident severity matrix is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident stakeholder
Incident stakeholder is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityIncident status update
Incident status update is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityInfrastructure monitoring
Infrastructure monitoring is observation of hosts, containers, networks, storage, and orchestration resources.
Reliability and observability