Find a term
Understand service health, incidents, and the telemetry used to investigate production systems. Connect reliability outcomes to engineering decisions.
Terms beginning with S
17 termsService level agreement
Service level agreement is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityService level credit
Service level credit is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityService level indicator
A service level indicator, or SLI, is a carefully defined quantitative measure of a service behavior that matters to users. Common examples include availability, request latency, error rate, and throughput.
Reliability and observabilityService level objective
A service level objective, or SLO, is a target value or range for a service level indicator. It states the level of service a team aims to provide over a defined period and measurement boundary.
Reliability and observabilityService map
Service map is a representation of services and observed communication paths between them.
Reliability and observabilityShared fate
Shared fate is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilitySingle point of failure
Single point of failure is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilitySLO alerting
SLO alerting is creation of notifications from measured service objectives and remaining error budget.
Reliability and observabilitySpan
Span is a timed unit of work within a trace with operation, timing, status, attributes, and events.
Reliability and observabilitySpan attribute
Span attribute is a key-value property attached to a span to provide searchable operation context.
Reliability and observabilitySpan event
Span event is a timestamped annotation attached to a span to record something during its lifetime.
Reliability and observabilitySpan ID
Span ID is the identifier for one span within a trace that distinguishes it from its parent and siblings.
Reliability and observabilitySplit brain
Split brain is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityStartup probe
Startup probe is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.
Reliability and observabilityStructured logging
Structured logging is writing log records as named fields instead of only free-form text.
Reliability and observabilitySummary metric
Summary metric is a client-side statistical summary of observations, commonly count, sum, and selected quantiles.
Reliability and observabilitySynthetic monitoring
Synthetic monitoring is scheduled scripted requests or workflows run from controlled conditions.
Reliability and observability