Instrument distributed applications with OpenTelemetry SDKs and propagate trace context across service and network boundaries.
Observability Engineering with Distributed Tracing and Metrics
Build observability pipelines that unify distributed traces, metrics and structured logs so engineering teams can pinpoint latency, errors and saturation across microservice architectures.
Course Overview
When a single user request crosses dozens of microservices, a slow page load or a failed checkout can no longer be traced through log files alone. Engineering teams need telemetry that shows the exact path a request took, where time was spent, and which dependency failed, without drowning on-call staff in noise. This course teaches participants to instrument applications with OpenTelemetry, propagate trace context across service boundaries, and correlate spans with metrics and structured logs inside a single query. Participants design cardinality-aware metric schemas, apply the RED and USE methods to choose what to measure, and configure head-based and tail-based sampling to control telemetry volume and cost. Sessions cover building dashboards on top of metrics and trace data, writing alert rules tied to service level objectives rather than raw thresholds, and using exemplars to jump from a metric spike straight to the trace that caused it. Practical labs work through a multi-service application with an injected latency fault, requiring participants to isolate the failing dependency using traces, metrics and logs together. By the end, participants can design an observability stack, define meaningful service level indicators and objectives, and reduce mean time to diagnosis for production incidents.
Expected Learning Outcomes
Design metric schemas using the RED and USE methods while controlling label cardinality to keep query performance stable.
Configure head-based and tail-based sampling strategies that balance trace completeness against storage and ingestion cost.
Build dashboards that correlate metrics, distributed traces and structured logs within a single triage workflow.
Define service level indicators, objectives and error budgets that drive alerting instead of arbitrary static thresholds.
Diagnose latency and failure incidents by tracing a request across microservices to the exact dependency at fault.
Evaluate observability tool choices and instrumentation coverage gaps across an existing microservice estate.
Who Should Attend
Backend and platform engineers instrumenting microservices for production monitoring.
Site reliability engineers responsible for latency and error budget tracking.
DevOps engineers building or maintaining metrics and tracing monitoring stacks.
Software architects designing telemetry standards for multi-service systems.
Engineering managers who need to interpret observability data during incident reviews.
Support and on-call engineers who triage production incidents using traces and dashboards.
Course Modules
Select any module to see its sessions and points.
01Instrumentation Foundations with OpenTelemetry
2 sessions · 8 points
Session 1Spans, Traces and Context Propagation
- Instrument service code with OpenTelemetry SDKs to emit spans that capture operation names, durations and attributes.
- Propagate trace context across HTTP, gRPC and message-queue boundaries using W3C Trace Context headers.
- Enrich spans with semantic conventions for HTTP, database and messaging attributes to keep traces comparable across teams.
- Configure the OpenTelemetry Collector to receive, batch and export telemetry over the OTLP protocol.
Session 2Sampling and Telemetry Volume Control
- Compare head-based and tail-based sampling strategies and their effect on trace completeness for rare errors.
- Set sampling rates per service and route to protect latency-sensitive paths from instrumentation overhead.
- Apply span and attribute limits to prevent high-cardinality data from overwhelming the tracing backend.
- Test instrumentation changes in a staging pipeline before promoting sampling configuration to production.
02Metrics Design and Dashboarding
2 sessions · 8 points
Session 1Metric Schemas with the RED and USE Methods
- Apply the RED method to define rate, error and duration metrics for request-driven services.
- Apply the USE method to define utilisation, saturation and error metrics for infrastructure resources.
- Design metric names and labels that avoid unbounded cardinality while remaining queryable at scale.
- Write queries that aggregate latency percentiles and error rates across service instances.
Session 2Building Correlated Dashboards
- Build dashboards that link metric panels to exemplars and the underlying trace for a spiking value.
- Design dashboard layouts around golden signals so on-call engineers can triage within the first minute.
- Correlate structured logs with trace identifiers to move from a dashboard alert to root cause evidence.
- Version dashboards and alert rules as code to keep observability configuration under change control.
03Service Level Objectives and Alerting
2 sessions · 8 points
Session 1Defining SLIs, SLOs and Error Budgets
- Select service level indicators that reflect what users actually experience rather than internal system health.
- Set service level objectives and error budgets that balance reliability targets against release velocity.
- Calculate burn rate for error budgets to distinguish a brief blip from a sustained reliability breach.
- Negotiate SLOs with product owners so reliability targets become shared commitments rather than engineering assumptions.
Session 2Alert Design and Noise Reduction
- Write multi-window, multi-burn-rate alert rules that catch fast and slow error budget depletion.
- Reduce alert fatigue by consolidating duplicate alerts and routing them through a single incident channel.
- Design runbooks linked to each alert so responders reach diagnostic dashboards without manual searching.
- Review alert precision and recall periodically to retire rules that no longer predict real incidents.
04Diagnosing Incidents Across Distributed Systems
2 sessions · 8 points
Session 1Root Cause Analysis with Correlated Telemetry
- Trace a slow user request end to end to identify which downstream service or database call caused the delay.
- Cross-reference trace spans with infrastructure metrics to distinguish application bugs from resource exhaustion.
- Use log correlation identifiers to pull the exact log lines associated with a failing trace span.
- Reconstruct a cascading failure timeline from traces and metrics during a multi-service outage.
Session 2Maturing the Observability Practice
- Audit instrumentation coverage across services to find blind spots before they cause undiagnosed outages.
- Estimate storage and ingestion cost for traces, metrics and logs to justify retention and sampling decisions.
- Establish an observability standard that keeps instrumentation consistent as teams and services grow.
- Plan a phased rollout of tracing and SLOs across a legacy estate that currently relies only on log files.
What the participant receives
4 course modules
A structured syllabus
8 training sessions
across 5 days
32 detailed points
Applied, detailed content
Accredited attendance certificate
On completing the programme
Complete your registration
We will contact you within one business day to confirm.
Ready to start?
Reserve your seat and start building the skill.
