01 Covers: The problem the rollout is meant to solve, one named accountable sponsor, and a written definition of success agreed before the first licence is bought.
Breaks when: The mandate is "roll out AI" with no named sponsor, so once the pilot's numbers turn ambiguous there is nobody left with the authority, or the obligation, to make the call.
Watch: Time between the first licence purchased and a written definition of success
02 Covers: Choosing a pilot team whose workload, codebase health, and appetite actually generalise, and naming the exit-from-pilot criteria before the pilot starts.
Breaks when: The pilot team is the platform group's best engineer working in a clean, greenfield service, so the result never predicts what a legacy team carrying real technical debt will see.
Watch: Share of exit-from-pilot criteria written down before the pilot's first week
03 Covers: The sanctioned harness and defaults, sandboxing, repo hygiene, agent instructions files, tests that actually gate a merge, the boundary of what an agent may touch, and enablement built as pairing in the team's own repo.
Breaks when: Rollout ships a licence and a slide deck with no sanctioned default, so every team improvises its own guardrailsGuardrails are the checks that keep an LLM or agent inside acceptable bounds in production. at a different level of care, and the resulting friction gets blamed on the tool.
Watch: Share of teams running on the sanctioned harness, not a bespoke one
04 Covers: Review capacity, CI throughput, spec quality, who owns the merge, and how the senior engineer's role changes once writing code stops being the bottleneck.
Breaks when: Agents multiply the diffs a team produces and review capacity never moves, so the queue that used to be "waiting on code" becomes "waiting on review," at the same total cycle time.
Watch: Review latency on agent-authored changes against the team's own pre-rollout baseline
05 Covers: The baseline captured before rollout, every adoption number paired with a quality number, and the split between the median engineer and the power user.
Breaks when: Nobody captured a pre-rollout baseline, so a genuinely improved number after rollout and a number that just looks better than a guess are indistinguishable.
Watch: Whether a dated pre-rollout baseline exists for every metric on the scoreboard
06 Covers: Plateau and regression past the launch spike, spend per team, unsanctioned use read as a signal, retiring pilot scaffolding on schedule, and a named platform owner with a refresh cadence.
Breaks when: Adoption plateaus at the enthusiasts within the first month, and the rollout reports that plateau as steady-state success because nobody asked why everyone else stopped coming back.
Watch: Active weekly usage among engineers with tool access, tracked for months past the launch spike rather than only the first one