Skip to main content
Course

Operating Platforms at Scale

Operate the internal platform hundreds of engineers depend on, and prove it earns its keep.

Advanced

Level

4

Modules

8

Lessons

4

Graded quizzes

2

Assignments

9 hours

Estimated time

What you will be able to do

  • You will be able to define platform SLIs and SLOs, set error budgets, and use them to make release and reliability decisions.
  • You will be able to instrument a platform for observability with OpenTelemetry, using metrics, logs, and traces to answer questions you did not predict.
  • You will be able to choose a multitenancy model (silo, pool, or bridge) and enforce isolation, quotas, and fairness so one tenant cannot starve the rest.
  • You will be able to allocate platform cost back to teams with showback or chargeback and manage unit economics using FinOps practices.
  • You will be able to run the platform as a product: define its users, publish paved paths, and maintain a roadmap and intake process in the open.
  • You will be able to measure platform adoption and developer experience using DORA metrics, the SPACE framework, and self-service ratios.
  • You will be able to write a platform impact review that ties reliability, cost, and adoption data to business outcomes leadership cares about.

What is inside

4 modules, 8 lessons. Each module ends in a graded quiz and most carry an assignment.

  1. 01

    Platform Reliability and Observability

    A platform at scale is a dependency for every team that builds on it, so its reliability is their reliability. This module shows you how to set service level indicators, objectives, and error budgets for platform services, then instrument them with the four golden signals, the RED and USE methods, and OpenTelemetry based distributed tracing so you can see inside the system when something breaks.

    2 lessons · 5 quiz questions

  2. 02

    Multitenancy, Isolation, and Cost

    One platform serves many tenants on shared infrastructure and produces a single cloud bill. This module teaches the three tenancy models (silo, pool, and bridge), how to isolate tenants and enforce quotas so one noisy neighbor cannot degrade everyone, and how to make the platform financially accountable using FinOps showback, chargeback, and unit economics.

    2 lessons · 5 quiz questions · assignment

  3. 03

    The Platform as a Product

    Internal platforms fail when they are built for the platform team instead of the teams that use them. Borrowing from Team Topologies, this module reframes the platform as a product with real customers: named personas, paved paths that make the right way the easy way, and a thinnest viable platform you can actually maintain. You will learn to run a roadmap and an intake process in the open, and to earn adoption instead of mandating it.

    2 lessons · 5 quiz questions

  4. 04

    Measuring Adoption and Impact

    A platform nobody adopts is a cost center, so the last job is proving it works. This module measures adoption, developer experience, and delivery outcomes using adoption funnels, DORA, the SPACE framework, and self-service ratios, then turns those signals into a platform impact review that leadership will actually read and act on.

    2 lessons · 5 quiz questions · assignment