Write a testable experiment hypothesis with primary, secondary and guardrail metrics.
Online Controlled Experiments and A/B Testing for Digital Services
Design, run and analyse online controlled experiments for digital services, from power analysis and sample size through to advanced adaptive designs.
Course Overview
An online experiment can clear every significance threshold and still be wrong, undone by a mismatched sample ratio, a fading novelty effect, or interference leaking between treatment and control groups. This course sets out how to design and analyse online controlled experiments so that a result can be trusted before it drives a shipping decision: choosing primary and guardrail metrics up front, calculating sample size through power analysis, and picking a randomisation unit that actually matches how the treatment operates on users. Variance reduction, sequential testing and Bayesian approaches are applied to reach conclusions without wasting traffic, alongside the diagnostic habits that catch sample ratio mismatch, novelty effects and marketplace interference before they cause a wrong call. A further module covers multi-armed bandits, factorial designs and long-term holdouts for situations a simple fixed-horizon test cannot handle. The final session turns individual experiments into an organisational habit, through a pre-launch review process and a shared repository of past results, wins and failures alike.
Expected Learning Outcomes
Calculate minimum detectable effect and sample size through power analysis before launch.
Select the randomisation unit and stratification approach that matches the treatment mechanism.
Apply variance reduction and sequential testing methods to reach conclusions efficiently.
Detect sample ratio mismatch, novelty effects and interference before trusting a result.
Design bandit, factorial or switchback experiments for adaptive or interference-prone contexts.
Embed experiment review and a shared results repository into the product decision process.
Who Should Attend
Data scientists and analysts running online controlled experiments on digital products.
Product managers deciding whether a tested change should ship, iterate or be abandoned.
Growth and marketing teams testing pricing, messaging or onboarding changes.
Engineers building or maintaining an internal experimentation platform.
Analytics leaders setting statistical standards for experiment design and analysis.
Two-sided marketplace teams managing interference between treatment and control groups.
Course Modules
Select any module to see its sessions and points.
01Designing Valid Online Experiments
2 sessions · 8 points
Session 1Hypotheses, Metrics and Sample Size
- Writing a testable hypothesis that ties a specific product change to an expected metric movement.
- Choosing primary, secondary and guardrail metrics before an experiment is allowed to launch.
- Calculating minimum detectable effect and required sample size through formal power analysis.
- Selecting the randomisation unit that matches how the treatment actually operates on users.
Session 2Randomisation, Stratification and Assignment Infrastructure
- Implementing consistent hashing so a user receives the same variant across sessions and devices.
- Stratifying randomisation on key segments to reduce variance in a heterogeneous user population.
- Ramping traffic allocation gradually to limit exposure while confirming instrumentation is correct.
- Checking for sample ratio mismatch as the first validity test run on any experiment's data.
02Statistical Analysis of Experiment Results
2 sessions · 8 points
Session 1Significance Testing and Variance Reduction
- Applying the significance test suited to a metric's distribution, such as a t-test or chi-square test.
- Using CUPED or regression adjustment with pre-experiment covariates to reduce variance and shorten run time.
- Correcting for multiple comparisons when an experiment reports results on many secondary metrics.
- Interpreting a confidence interval rather than a single point estimate when sizing a business decision.
Session 2Sequential and Bayesian Approaches
- Applying always-valid sequential testing so a team can monitor results without inflating false positives.
- Comparing frequentist and Bayesian A/B testing approaches for clarity in stakeholder communication.
- Avoiding the peeking problem where early stopping on a promising trend biases the final result.
- Deciding a pre-registered stopping rule before an experiment launches rather than after results arrive.
03Avoiding Common Pitfalls and Interference
2 sessions · 8 points
Session 1Novelty Effects, Interference and Segment Traps
- Distinguishing a genuine effect from a novelty or primacy effect that fades after initial exposure.
- Detecting interference between treatment and control groups in a two-sided marketplace or social product.
- Applying switchback or cluster-randomised designs when individual-level randomisation would leak between groups.
- Checking for Simpson's paradox before drawing conclusions from a segment-level breakdown of results.
Session 2Guardrails, Data Quality and Metric Gaming
- Setting guardrail metrics that automatically flag an experiment causing harm before full rollout.
- Auditing logging and instrumentation for an experiment before trusting its metric pipeline.
- Recognising when a metric improvement reflects gaming rather than a genuine behaviour shift.
- Reconciling conflicting signals between a primary metric and a longer-term guardrail metric.
04Advanced Designs and Experimentation Culture
2 sessions · 8 points
Session 1Bandits, Factorial Designs and Long-Term Holdouts
- Using multi-armed bandits to adapt traffic allocation towards better-performing variants during a test.
- Running factorial designs to test multiple independent changes within a single experiment.
- Maintaining a long-term holdout group to measure effects that only appear after sustained exposure.
- Deciding when an adaptive design is worth its added statistical and engineering complexity.
Session 2Embedding Experimentation in Product Decisions
- Setting a review process that checks experiment design and metrics before launch, not after.
- Maintaining a shared repository of past experiment results so teams do not repeat failed tests.
- Translating a statistically significant result into a clear ship, iterate or abandon recommendation.
- Building organisational trust in experimentation by reporting negative results alongside wins.
What the participant receives
4 course modules
A structured syllabus
8 training sessions
across 5 days
32 detailed points
Applied, detailed content
Accredited attendance certificate
On completing the programme
Complete your registration
We will contact you within one business day to confirm.
Ready to start?
Reserve your seat and start building the skill.
