Explain the role of each control plane and node component in a running Kubernetes cluster.
Kubernetes Cluster Operations and Workload Right-Sizing
Operate production Kubernetes clusters and right-size workloads with resource requests, limits and autoscaling, covering control plane health, upgrades, scheduling efficiency and per-namespace cost visibility.
Course Overview
A Kubernetes cluster that runs without incident on day one tends to drift over months into pods with no resource limits, nodes sitting mostly idle, and upgrades postponed until a supported version is no longer available. This course covers the operational discipline that keeps a cluster healthy and affordable long after initial rollout. It starts with the control plane and node components that make up a cluster, then moves to workload right-sizing: setting resource requests and limits from real usage data, choosing quality-of-service classes deliberately, and configuring horizontal, vertical and cluster autoscaling so capacity tracks demand instead of a fixed guess. Participants design namespaces, resource quotas and limit ranges for multi-team clusters, apply pod disruption budgets and node affinity rules to keep workloads available during maintenance, and plan control plane and node upgrades that respect version skew policy. The course also covers etcd backup and restore, role-based access control and network policy for tenant isolation, finishing with per-namespace cost visibility so a team can see the cost consequence of its own scheduling and sizing decisions.
Expected Learning Outcomes
Set resource requests and limits from measured usage rather than default or guessed values.
Configure horizontal, vertical and cluster autoscaling so capacity tracks real demand.
Design namespaces, resource quotas and limit ranges for a multi-team cluster.
Apply pod disruption budgets and affinity rules to protect availability during maintenance.
Plan control plane and node upgrades that respect version skew and workload continuity.
Report per-namespace resource cost so teams see the impact of their own sizing decisions.
Who Should Attend
Platform engineers operating Kubernetes clusters for multiple internal teams.
Site reliability engineers responsible for cluster availability and upgrades.
Application teams wanting to right-size workloads running on shared clusters.
Infrastructure engineers planning capacity for growing Kubernetes environments.
DevOps engineers troubleshooting scheduling and resource contention issues.
Engineering managers accountable for the cost of their team's Kubernetes workloads.
Course Modules
Select any module to see its sessions and points.
01Cluster Architecture and Health
2 sessions · 8 points
Session 1Control Plane and Node Components
- Explain the role of the API server, etcd, scheduler and controller manager in cluster operation.
- Describe how the kubelet and container runtime manage pods on each node.
- Diagnose a control plane component failure using logs and component status checks.
- Back up and restore etcd so cluster state can be recovered after a serious failure.
Session 2Node Pools and Cluster Topology
- Design node pools separated by workload type, instance size or hardware requirement.
- Apply taints and tolerations to keep specialised workloads on their intended nodes.
- Plan node lifecycle including draining, cordoning and safe replacement.
- Size a node pool for both steady-state load and expected burst demand.
02Resource Requests, Limits and Autoscaling
2 sessions · 8 points
Session 1Setting Requests, Limits and Quality of Service
- Set CPU and memory requests and limits based on measured pod usage over time.
- Choose Guaranteed, Burstable or BestEffort quality of service deliberately per workload.
- Diagnose throttling and out-of-memory terminations caused by incorrect limits.
- Configure liveness, readiness and startup probes that reflect real application health.
Session 2Horizontal, Vertical and Cluster Autoscaling
- Configure a horizontal pod autoscaler driven by CPU, memory or custom metrics.
- Apply a vertical pod autoscaler to recommend or adjust requests automatically.
- Configure cluster autoscaling so node count tracks pending pod demand.
- Test autoscaling behaviour under a simulated load increase and decrease.
03Multi-Team Governance and Availability
2 sessions · 8 points
Session 1Namespaces, Quotas and Multi-Tenancy
- Design namespace structure that separates teams, environments or applications.
- Apply resource quotas that cap total consumption per namespace.
- Set limit ranges that enforce sensible default and maximum resource values per pod.
- Apply role-based access control and network policy to isolate namespaces from each other.
Session 2Availability and Disruption Management
- Configure pod disruption budgets so voluntary disruptions do not remove too many replicas at once.
- Apply pod affinity and anti-affinity rules to spread or group workloads deliberately.
- Design workloads to tolerate node drains during planned maintenance without an outage.
- Test failover behaviour when a node is removed unexpectedly from the cluster.
04Upgrades, Cost Visibility and Ongoing Operations
2 sessions · 8 points
Session 1Cluster and Workload Upgrades
- Plan control plane upgrades that respect the supported version skew with nodes.
- Sequence node upgrades so workload availability is maintained throughout the process.
- Validate workload compatibility against a new Kubernetes version before upgrading.
- Roll back an upgrade safely if a post-upgrade issue is detected.
Session 2Cost Visibility and Continuous Right-Sizing
- Attribute compute cost to namespaces and teams based on actual resource consumption.
- Identify chronically over-provisioned workloads through ongoing usage analysis.
- Report scheduling efficiency and bin-packing quality across the cluster's nodes.
- Build a recurring right-sizing review that keeps requests and limits aligned with usage.
What the participant receives
4 course modules
A structured syllabus
8 training sessions
across 5 days
32 detailed points
Applied, detailed content
Accredited attendance certificate
On completing the programme
Complete your registration
We will contact you within one business day to confirm.
Ready to start?
Reserve your seat and start building the skill.
