CustomLabs
Adoption

The tools work. Getting 180 engineers to actually use them is the other half of the job.

Six places a rollout stalls, 24 practices split by who owns each move, five ways a programme fakes momentum, and the numbers that separate real adoption from seat activation. Most of what kills an agent rollout has nothing to do with the model.

Updated First published

33 min read

Markdown

A working pilot proves the harness can work somewhere. It says nothing about whether the other engineers who weren't in the room will pick it up, whether the review queue can absorb five times the diffs, or whether anyone can tell a genuine improvement from a number that just looks good next to no baseline. That gap between a proven pilot and an organisation that actually runs on it is where most rollouts quietly stall.

The Agentic Delivery Playbook covers the machine: task definition, isolation, stage gates, spend. This page is what has to be true in the organisation around that machine before it produces anything at scale.

The six surfaces

Where a rollout actually stalls.#

Each surface fails a different way. None of them substitute for each other.

Diagram in six lanes, one per adoption surface, in order: the mandate; the first team; and the paved road. The other three are absorbing the output; measurement; and sustain. Each lane holds one node that names what that surface covers. The surfaces connect in order, with a return edge from measurement back to the first team labeled expand or hold. THE MANDATE THE FIRST TEAM THE PAVED ROAD ABSORBING THE OUTPUT MEASUREMENT SUSTAIN Named sponsor Exit criteria Sanctioned harness Review capacity Pre-rollout baseline Plateau past launch expand or hold
An agent rollout moves through six stages, from the first mandate to a sustained program. Measurement results decide whether the pilot expands or holds.
01

The mandate

Covers: The problem the rollout is meant to solve, one named accountable sponsor, and a written definition of success agreed before the first licence is bought.

Breaks when: The mandate is "roll out AI" with no named sponsor, so once the pilot's numbers turn ambiguous there is nobody left with the authority, or the obligation, to make the call.

Watch: Time between the first licence purchased and a written definition of success

02

The first team

Covers: Choosing a pilot team whose workload, codebase health, and appetite actually generalise, and naming the exit-from-pilot criteria before the pilot starts.

Breaks when: The pilot team is the platform group's best engineer working in a clean, greenfield service, so the result never predicts what a legacy team carrying real technical debt will see.

Watch: Share of exit-from-pilot criteria written down before the pilot's first week

03

The paved road

Covers: The sanctioned harness and defaults, sandboxing, repo hygiene, agent instructions files, tests that actually gate a merge, the boundary of what an agent may touch, and enablement built as pairing in the team's own repo.

Breaks when: Rollout ships a licence and a slide deck with no sanctioned default, so every team improvises its own guardrailsGuardrails are the checks that keep an LLM or agent inside acceptable bounds in production. at a different level of care, and the resulting friction gets blamed on the tool.

Watch: Share of teams running on the sanctioned harness, not a bespoke one

04

Absorbing the output

Covers: Review capacity, CI throughput, spec quality, who owns the merge, and how the senior engineer's role changes once writing code stops being the bottleneck.

Breaks when: Agents multiply the diffs a team produces and review capacity never moves, so the queue that used to be "waiting on code" becomes "waiting on review," at the same total cycle time.

Watch: Review latency on agent-authored changes against the team's own pre-rollout baseline

05

Measurement

Covers: The baseline captured before rollout, every adoption number paired with a quality number, and the split between the median engineer and the power user.

Breaks when: Nobody captured a pre-rollout baseline, so a genuinely improved number after rollout and a number that just looks better than a guess are indistinguishable.

Watch: Whether a dated pre-rollout baseline exists for every metric on the scoreboard

06

Sustain

Covers: Plateau and regression past the launch spike, spend per team, unsanctioned use read as a signal, retiring pilot scaffolding on schedule, and a named platform owner with a refresh cadence.

Breaks when: Adoption plateaus at the enthusiasts within the first month, and the rollout reports that plateau as steady-state success because nobody asked why everyone else stopped coming back.

Watch: Active weekly usage among engineers with tool access, tracked for months past the launch spike rather than only the first one

The practice bank

24 practices, six surfaces.#

Filter by how far a practice has to reach (pilot, team, org), then copy the visible list as a Markdown checklist.

01 The mandate

Every agent rollout starts with an announcement, and most of them skip the part that would make the announcement mean anything: a single named owner and a written definition of success, agreed before the first licence is bought. Without that, the rollout runs on borrowed authority — a budget gets approved because a competitor shipped something, or because a demo landed well, and the actual question, what problem this is supposed to solve and how anyone will know it worked, never gets answered in writing. The gap doesn't show up immediately. It shows up months in, when the pilot's numbers are ambiguous and there's no sponsor left to make the call, and no written definition of success to make the call against.

What it covers
A single named executive accountable for whether the rollout succeeds, distinct from whoever runs the pilot day to day.
How to build it
Put a name and a title on the mandate document itself, signed before procurement starts, not assigned after the pilot already has momentum.
Produces
A single person to ask when the rollout stalls, instead of a committee nobody can point to.
Tradeoff
A sponsor who changes teams or leaves takes the mandate's authority with them. Re-confirm ownership at every leadership change; don't assume it survives one.
Prove it
The current sponsor can state the rollout's definition of success from memory, not from re-reading the deck.
What it covers
A short, specific statement of what the rollout is meant to change, agreed in writing before the first licence or seat is purchased.
How to build it
Draft it as a testable claim rather than a mission statement, with a stated target and a date, and get the sponsor to sign it.
Produces
A rollout that can be judged against what it said it would do, rather than reinterpreted after the fact to match whatever happened.
Tradeoff
A specific target invites scrutiny a vague one avoids. That scrutiny is the point.
Prove it
The definition of success predates the first purchase order, checked against the actual dates.
What it covers
A stated business problem, slow delivery, a backlog nobody can clear, an unmet SLA, that the rollout is supposed to move, ahead of any product decision.
How to build it
Write the problem statement before the vendor conversation, and route every later tool choice back through whether it moves that specific problem.
Produces
A rollout that can say no to a feature or a vendor that doesn't serve the actual problem, instead of collecting capability for its own sake.
Tradeoff
A narrow problem statement rules out adjacent wins the tool could also deliver. Note them; don't chase them yet.
Prove it
Someone outside the rollout team can restate the problem in one sentence without reading the deck.
What it covers
A short, written list of outcomes the rollout is not being judged on this cycle: headcount reduction, replacing a specific role, or a fixed seat-count target.
How to build it
Add a non-goals section to the mandate document itself, agreed by the sponsor, and repeat it in the first team-facing announcement.
Produces
Fewer engineers treating the rollout as a threat to route around quietly, because the thing they were worried about was already ruled out in writing.
Tradeoff
A non-goal stated today can look naive in a year if the plan genuinely changes. Revisit it on the same cadence as the mandate itself, and say so out loud when it does.
Prove it
A team lead can quote the non-goals list without checking, and it matches what leadership is actually saying in private.

02 The first team

The team that pilots an agent rollout decides more about the result than the tool does. A team picked because its manager was enthusiastic, its codebase was already clean, or its tickets were unusually well scoped produces a pilot that looks great and predicts almost nothing about the teams waiting behind it. The fix isn't finding the best team; it's finding a representative one, and naming what "worked" means before the pilot starts, so the decision to expand isn't made by whoever tells the best story afterward.

What it covers
A pilot team selected because its workload shape (ticket size, codebase age, test coverage) resembles the org's median, not because its manager volunteered first.
How to build it
Score candidate teams against the organisation's actual workload distribution before asking anyone if they want to go first.
Produces
A pilot result that predicts what the rest of the organisation will see, instead of a best case nobody else can reach.
Tradeoff
A representative team is harder to recruit than an enthusiastic outlier. Recruit on the strength of the mandate, not on who happens to be excited.
Prove it
The pilot team's ticket-size and test-coverage numbers sit within a stated range of the org median, checked against real data, not a guess.
What it covers
A numeric bar (acceptance rate, review latency, a rollback-rate ceiling) the pilot has to clear before the rollout expands past it.
How to build it
Agree the bar with the sponsor before the pilot's first week, and publish it to the pilot team so they know what they're being measured against.
Produces
An expansion decision made against a rule agreed in advance, instead of a judgment call once the pilot has been running a while.
Tradeoff
A bar set too high can kill a pilot that was actually working, just slowly. Calibrate it against the first team's baseline, not an aspirational number.
Prove it
The exit-criteria document predates the pilot's first week, and the expand-or-hold decision cites it directly.
What it covers
A quick read on the pilot repo's test coverage, CI speed, and documentation state, since an agent inherits whatever review and test discipline already exists there.
How to build it
Run the same health check the team would apply before any other significant investment, and treat a repo with no tests or a broken CI pipeline as a reason to fix that first, or pick a different pilot.
Produces
A pilot result attributable to the tooling, not to a codebase that happened to make everything look easy.
Tradeoff
Waiting for a healthier repo costs calendar time a sponsor may not want to spend. Weigh that against a result nobody trusts once it doesn't repeat elsewhere.
Prove it
The pilot repo's coverage and CI numbers were checked and recorded before the pilot's first change merged.
What it covers
A direct check with the candidate pilot team on their appetite for the change, distinct from their manager's enthusiasm.
How to build it
Talk to the engineers who'll do the work before the decision is finalised, and treat a genuinely reluctant team as a signal to pick a different one, not a change-management problem to push through.
Produces
A pilot that isn't quietly sandbagged by a team that never wanted it in the first place.
Tradeoff
Deferring to appetite can mean picking a less representative team. Weigh that against a pilot result nobody believes because the team's heart wasn't in it.
Prove it
The pilot team can name, unprompted, something about the rollout they are curious about, rather than only what they were told to expect.

03 The paved road

A licence and a link to the vendor's documentation is not a paved road. It's the absence of one, dressed as adoption. What actually makes an agent effective in a codebase — a sandboxed environment, a repo-local instructions file, tests that genuinely gate a merge, and a stated boundary of what the agent may touch — has to exist before a second team is asked to use any of it. Enablement that's a worked example in the team's own repo, done alongside a person who already knows the harness, beats a webinar every time, because it's the only version that survives contact with a real ticket.

What it covers
A single supported configuration, sandboxing, a default instructions file, credential scoping, that every team starts from, rather than a licence and a link to the vendor's documentation.
How to build it
Build and pilot the default harness on the first team, then package it as the thing a new team installs on day one, not the thing they eventually converge on after improvising for a month.
Produces
A second team's onboarding measured in days against a working baseline, instead of weeks spent reinventing the first team's decisions.
Tradeoff
A single sanctioned default won't fit every workload exactly. Let teams diverge deliberately once they can name why it doesn't fit, not by default.
Prove it
A new team's first week of agent-assisted changes runs on the sanctioned harness, checked against their actual configuration, not the onboarding checklist's claim.
What it covers
A file in the repository itself stating the tests that gate a merge, the boundary of what an agent may touch, and the house conventions a human reviewer would otherwise have to explain in every pull request.
How to build it
Write the first version from the pilot team's actual review comments, then require it as part of onboarding any new repo onto the paved road.
Produces
An agent, and a new human hire, that reads the same file the team actually works from, instead of a wiki page that drifted out of date months ago.
Tradeoff
A repo-local file needs its own upkeep, one more thing that can fall stale. Tie updates to the same review that catches other doc drift, not a separate process nobody owns.
Prove it
A repo picked at random has an instructions file that matches its actual current test-gating and merge rules, not last quarter's.
What it covers
An explicit, written scope of which files, systems, and actions an agent is allowed to change unsupervised, and which always route to a human.
How to build it
Start narrow, read and propose, never merge to a protected branch unattended, and widen the boundary only after the review data supports it, not on day one.
Produces
A rollout that can say precisely what went wrong when something does, instead of debating afterward whether the agent should have been allowed to do it.
Tradeoff
A narrow boundary limits how much benefit teams see early, which is real friction against the mandate's own success definition. Widen it on a measured schedule, not to quiet the complaint.
Prove it
The current boundary is written down somewhere a new team member actually reads, and the last incident review checked the action against it.
What it covers
A worked example built against the team's actual codebase and backlog, walked through with a person in the room, rather than a generic training deck.
How to build it
Pair a platform engineer with two or three engineers from the new team for their first real ticket, then let them run the second one alone.
Produces
A team that has seen the harness work on a ticket they recognise, instead of a slide deck they nodded through.
Tradeoff
Pairing doesn't scale to a large organisation by headcount alone. Train a small number of team-level champions this way, then let peer diffusion carry it past the platform team's own capacity.
Prove it
A newly onboarded engineer can name the specific ticket they paired on, rather than only that a training session happened.

04 Absorbing the output

An agent that writes code faster changes nothing on its own if the team downstream can't absorb what it produces. Review capacity, CI throughput, and spec quality were sized for the previous rate of change, and a rollout that doesn't resize them just moves the bottleneck from writing code to reviewing it, at the same overall speed. The senior engineer's job changes here too, whether anyone planned for it or not: once code stops being the scarce thing, judgment about what to build and what to approve becomes the actual work, and a team that never says so out loud ends up with its senior engineers quietly competing with agents at the one task that stopped being the bottleneck.

What it covers
A concrete plan for who reviews the extra pull requests an agent-assisted team produces, agreed before rollout, not discovered as a growing queue after.
How to build it
Model the expected diff volume from the pilot's own numbers, and check whether current reviewers can absorb it inside their existing hours, or whether the review step itself needs a design change.
Produces
A rollout that doesn't quietly convert a code-writing bottleneck into a review-queue bottleneck of the same size.
Tradeoff
Adding review capacity has a real cost, whether that is headcount or a deliberately slower ramp. Both are cheaper than a queue nobody planned for.
Prove it
The team's review latency after rollout sits at or below its pre-rollout baseline, checked directly rather than assumed from the plan.
What it covers
A stated rule for who is accountable for a merged change's correctness, regardless of whether an agent or a human wrote the diff.
How to build it
Keep the existing human-approver requirement on every merge, and say explicitly that approving an agent-authored diff carries the same accountability as approving a human-authored one.
Produces
A merge queue where "the agent wrote it" is never a legitimate answer to "who approved this."
Tradeoff
This adds no new process, which is the point, but it does mean reviewers can't treat an agent's diff as lower stakes to wave through. Say that out loud; don't leave it implied.
Prove it
A reviewer asked why they approved a specific agent-authored change gives the same kind of answer they would give for a human-authored one.
What it covers
A minimum bar for what a ticket needs to contain (acceptance criteria, the files in scope, what "done" looks like) before it is handed to an agent, matching what a competent engineer would want before starting the same ticket.
How to build it
Reject or return a ticket that fails to clear the bar at triage, the same discipline a team would apply before assigning it to a person.
Produces
Fewer plausible-looking diffs against a vague ticket, the single biggest driver of review time on agent-assisted work.
Tradeoff
A stricter triage bar slows down ticket creation, which reads as friction against exactly the speed the rollout is supposed to deliver. It is friction that pays for itself in review time saved.
Prove it
A sample of tickets from the last sprint each clear the stated bar, checked against the ticket text itself, not the triage process's own account of it.
What it covers
An explicit statement that review, spec-writing, and architectural judgment are now the senior engineer's primary output, replacing an implicit expectation that their value is measured in lines shipped.
How to build it
Update how senior engineers are evaluated and staffed to credit review quality and spec clarity alongside their own commit volume, before the incentives quietly punish the people doing the most valuable work.
Produces
Senior engineers who spend their time where the bottleneck actually moved, instead of competing with agents at the one thing agents are now faster at.
Tradeoff
Changing how a role is evaluated is a genuine management change, not a tooling rollout, and it lands slower than the harness itself. Start the conversation with the pilot team early, before the mismatch shows up in someone's review.
Prove it
A senior engineer on the pilot team can describe their own role change in their own words, not the rollout announcement's words.

05 Measurement

A number reported after rollout with no baseline before it is an anecdote wearing a dashboard. Most programmes never capture what cycle time, review latency, or change-failure rate looked like before the harness arrived, which means every improvement claim afterward is measured against a guess. The fix is procedural, not statistical: take the baseline in week one, on every team the rollout plans to compare later, and pair every adoption number that follows with the quality number that would catch it flattering the programme instead of describing it.

What it covers
A recorded measurement of every scoreboard metric (cycle time, review latency, change-failure rate) for the pilot team, taken before the harness is installed.
How to build it
Run the measurement pass as the pilot's actual first step, before the first agent-assisted change merges, and store it somewhere the post-rollout comparison can actually reach.
Produces
A real before-and-after comparison, instead of a post-rollout number with nothing honest to compare it against.
Tradeoff
Measuring a baseline delays the pilot's visible start by a sprint or two. That delay is what makes the eventual comparison mean anything at all.
Prove it
A dated baseline snapshot exists for the pilot team, and it predates the first agent-assisted merge, checked against the actual timeline.
What it covers
A rule that no adoption metric (seats activated, weekly active users, diffs merged) gets reported without the paired quality metric that would catch it flattering the programme.
How to build it
Publish the two numbers together on the same dashboard, in the same report, every time, never the adoption number alone.
Produces
A scoreboard that can't be read as "adoption is working" when the actual story is usage up and quality flat or worse.
Tradeoff
A paired number is more work to compute and can undercut a headline the sponsor wanted to lead with. Report it anyway; the alternative is a number nobody trusts six months later.
Prove it
The last quarterly rollout report shows every adoption figure with its paired quality figure on the same page.
What it covers
A reporting split between the top decile of usage and the median engineer, since a rollout's average is usually carried by a small number of enthusiasts.
How to build it
Segment every adoption metric by usage decile before reporting an org-wide average, and state both the median and the top decile explicitly.
Produces
An honest read on whether the rollout reached the organisation or just its most eager early adopters.
Tradeoff
A segmented report is a less flattering headline than a single averaged number. It is also the only version that predicts what happens when the enthusiasts change teams.
Prove it
The current adoption report states the median engineer's usage figure, not only the mean.
What it covers
A cost figure denominated in what actually shipped, a merged, reviewed, non-reverted change, rather than a per-seat licence cost that says nothing about output.
How to build it
Divide a team's total rollout spend by its count of merged agent-assisted changes over the same period, tracked on the same cadence as any other delivery metric.
Produces
A spend number that moves with actual delivery, so cost per outcome is visible before the finance review finds it first.
Tradeoff
Cost per merged change needs the same ledger discipline a fleet's own per-task spend tracking already requires. It doesn't fall out of a seat count for free.
Prove it
The last quarter's cost-per-merged-change figure for the pilot team is checked against actual billing, not estimated from seat count.

06 Sustain

Adoption that plateaus at the enthusiasts looks like success if the only number anyone tracks is whether the licence got used at all. The harder, more useful questions come later: why the other engineers stopped after week one, what the unsanctioned workarounds are actually saying about the paved road, and who owns the harness once the pilot team that built it has moved on to something else. A rollout that never revisits its own scaffolding, or never sets a cadence for refreshing the harness as the underlying tooling changes, ages into the exact rigidity it was supposed to replace.

What it covers
A standing check for shadow use, engineers routing around the sanctioned harness with a personal account or an unapproved tool, treated first as evidence the paved road is missing something.
How to build it
Ask the engineers doing it what the sanctioned harness doesn't cover, before writing a policy that blocks the workaround.
Produces
A paved road that closes the actual gap driving the workaround, instead of a ban that pushes the same behaviour further out of sight.
Tradeoff
Treating shadow use as a discovery process is slower than blocking it outright, and it means tolerating some ungoverned use while the gap gets fixed, bounded by whatever the security review already requires.
Prove it
The last confirmed instance of unsanctioned use resulted in a named change to the sanctioned harness, rather than only a warning.
What it covers
A team or role accountable for the sanctioned harness after the pilot ends, distinct from the pilot team itself, which will eventually move to other work.
How to build it
Stand up the ownership handover as a named step in the exit-from-pilot criteria, not an afterthought once the pilot team's attention has already moved elsewhere.
Produces
A paved road that survives the specific people who built it moving to a different project.
Tradeoff
A dedicated platform owner is a real, ongoing cost, not a one-time project expense. Budget it as such from the mandate stage, not as a surprise once the pilot ends.
Prove it
The current platform owner is not the same person who ran the original pilot, or an explicit reason is on record for why that hasn't changed yet.
What it covers
A stated date to remove the temporary shortcuts, manual approvals, and extra logging a pilot runs that a scaled rollout shouldn't carry indefinitely.
How to build it
Write the retirement date into the exit-from-pilot criteria alongside the metrics bar, and treat scaffolding still in place past that date as a tracked debt item, not an invisible cost.
Produces
A rollout that doesn't quietly carry pilot-era overhead into a process meant to run at many times the pilot's scale.
Tradeoff
Some scaffolding turns out to still be load-bearing once removal is attempted. Treat that discovery as a reason to make it a permanent, supported control, not a reason to leave it as an unowned pilot leftover.
Prove it
The scaffolding list from the pilot's exit criteria shows each item as either retired or explicitly promoted to a supported control.
What it covers
A recurring, scheduled review of the sanctioned default, the harness, the instructions-file template, the tool boundary, against how the underlying agent tooling has actually changed since it launched.
How to build it
Put the review on the same calendar discipline as any other platform component's support lifecycle, with a named owner and a fixed interval, not a review triggered only by a complaint.
Produces
A paved road that keeps pace with what the tooling can now do, instead of ossifying at whatever the pilot happened to ship with.
Tradeoff
A fixed refresh cadence spends platform time on a review that sometimes concludes nothing needs to change. That is a cheaper outcome than a harness nobody revisits until it is visibly behind.
Prove it
The last scheduled harness review happened on the calendar date it was due, and produced either a change or a stated reason none was needed.
Five ways a rollout lies to itself

A busy dashboard is a claim, not proof.#

Every one of these looks like a working rollout right up until the number it was supposed to produce turns out to be hollow.

Seat activation counted as adoption

Looks like: A dashboard shows every engineer with a licence assigned and at least one login, and the rollout reports that as adoption.

Costs you: A licence assigned and a login once says nothing about whether anyone landed a real merged change with it, so the rollout can report full activation while actual usage sits with a handful of engineers.

Fix: Report the paired quality number alongside every adoption figure, and segment by usage decile so the median engineer's number can't hide behind the top decile's.

See the practice: Pair every adoption number with a quality number →

The power user's numbers generalised to the org

Looks like: One enthusiastic team lead reports a large personal speedup, and the rollout report quotes that number as the expected result org-wide.

Costs you: A number that's true for the one person already motivated to make the tool work says nothing about the engineers who haven't tried it yet, and setting expectations against it makes every ordinary result look like underperformance.

Fix: Segment every metric by usage decile and state the median explicitly, alongside the number that made the best slide.

See the practice: Measure the power user and the median engineer separately →

Time saved with no throughput or quality change

Looks like: Engineers report the tool feels faster, and the rollout counts that as a win.

Costs you: If review absorbed the saved time — a bigger queue, more back-and-forth on agent-authored diffs — or the change-failure rate crept up at the same time, the time saved never reached a shipped, working outcome. It just moved the bottleneck downstream where nobody was watching.

Fix: Size review capacity to the new diff volume before rollout, and track cycle time and change-failure rate as the numbers that would catch the trade actually happening.

See the practice: Size review capacity to the new diff volume before it arrives →

A mandate with no paved road

Looks like: Leadership announces the rollout and a licence budget, and expects teams to work out the rest themselves.

Costs you: Every team invents its own guardrails at a different level of care, engineers read the resulting friction as the tool being bad rather than the harness being missing, and the mandate's credibility erodes before the paved road ever ships.

Fix: Ship one sanctioned default harness, piloted and packaged, before the mandate reaches teams beyond the pilot.

See the practice: Ship one sanctioned default harness before asking teams to adopt anything →

A pilot that depended on someone babysitting it

Looks like: The pilot clears every exit criterion, and the team credits the harness.

Costs you: If the actual reason it worked was one platform engineer quietly fixing broken configs and coaching people through failures all quarter, the result doesn't survive that person moving to something else, and the second team's onboarding fails in ways the pilot never surfaced.

Fix: Enablement that is pairing and worked examples in the team's own repo, plus a named platform owner distinct from whoever ran the original pilot.

See the practice: Name a platform owner before the pilot team disperses →
The scoreboard

Six numbers, and how each one lies.#

Every one of these is measurable today. None of them is trustworthy read alone. Pair it with the number next to it.

Share of engineers landing a merged agent-assisted change weekly

Why it matters
The most direct read on whether the rollout reached the organisation or just its early adopters.
How it misleads
A rising share can still be entirely the same enthusiasts merging more often, with no new engineers ever starting.
Pair with
Count of distinct engineers who landed their first agent-assisted merge this month

Acceptance rate of agent-authored changes

Why it matters
Whether the diffs an agent proposes are actually good enough to ship, as opposed to merely numerous.
How it misleads
A high rate on a narrow, well-scoped ticket type says nothing about the harder tickets nobody's routed to an agent yet.
Pair with
Share of ticket types actually being routed to an agent, beyond the ones the rate is measured on

Review latency and review-hours per merged change

Why it matters
Catches a code-writing bottleneck turning into a review-queue bottleneck of the same size.
How it misleads
Latency can look fine if reviewers are quietly working unpaid overtime to keep the queue moving; hours per change is what actually catches that.
Pair with
Reviewer-reported time spent per agent-authored change, tracked separately from latency

Change-failure and rollback rate on agent-authored changes

Why it matters
Whether shipping faster came at the cost of shipping worse.
How it misleads
A flat rollback rate can hide a rise in near-misses caught in review that never reached production, which the rate itself never counts.
Pair with
Defects caught in review per agent-authored change, tracked as its own number

Cost per merged change

Why it matters
The spend number that moves with real output, rather than with how many licences were assigned.
How it misleads
A falling cost per change can hide a rollback rate rising at the same time, if the change never gets checked against quality.
Pair with
Change-failure rate on the same set of changes, read on the same dashboard

A new joiner's time to first merged agent-assisted change

Why it matters
The clearest read on whether the paved road actually onboards someone who wasn't there for the pilot.
How it misleads
A fast time for a joiner who happens to land on the original pilot team says nothing about a joiner on a team several steps away from the platform team's attention.
Pair with
The same figure measured on a team outside the original pilot
What this is built from

Verifiable, not claimed.#

No invented survey statistics. Just what's already documented on this site, and how it connects.

Sources

  1. NIST - AI Risk Management Framework (AI RMF 1.0)

    The risk-management functions our governance and adoption controls map onto. Retrieved 2026-08-24.

  2. Stanford HAI - AI Index Report

    The annual report our adoption page's industry-wide statistics are drawn from. Retrieved 2026-08-24.

Not sure the pilot would actually generalise

A Ship Audit checks the paved road and the absorption math against your specific team before the rollout reaches the next 179 engineers.

Questions

Before you announce a rollout.#

What teams ask us before they take an agent programme past the first team.

01 Isn't this just change management with new labels?

Change management is the general discipline. This is the specific version of it for an agent rollout: a paved road, review-capacity math, and a scoreboard an agent programme needs that a generic change programme doesn't ship with.

Link to this answer: Isn't this just change management with new labels?
02 We already ran a successful pilot. Why isn't that enough?

A pilot proves the harness can work somewhere. It doesn't prove the result generalises to a team with a different workload, or that the organisation can absorb several times the review volume once the rollout scales past the pilot team.

Link to this answer: We already ran a successful pilot. Why isn't that enough?
03 What's the fastest way to kill a rollout without meaning to?

Skip the paved road and ship a licence with no sanctioned default. Every team improvises its own guardrails at a different level of care, the friction gets read as the tool being bad, and the mandate loses credibility before anyone built the thing that would have made it work.

Link to this answer: What's the fastest way to kill a rollout without meaning to?
04 How do we know adoption is real, rather than seat activation with nothing behind it?

Pair every adoption figure with its quality counterpart and check the median engineer alongside the top decile. Activation with no merged output and no movement in review latency is the seat-activation trap, and it looks identical to real adoption on a dashboard that only reports the average.

Link to this answer: How do we know adoption is real, rather than seat activation with nothing behind it?
05 Who should own this, the platform team or engineering leadership?

Both, at different points. Leadership owns the mandate and the definition of success; the platform team owns the paved road and its upkeep once the pilot ends. A rollout missing either one stalls in a different place depending on which is missing.

Link to this answer: Who should own this, the platform team or engineering leadership?
06 Does this still apply once we're past the pilot and scaling org-wide?

Yes. Most of what actually breaks a rollout — absorption capacity, the median-versus-power-user gap, unsanctioned use — shows up after the pilot succeeds, not during it. Revisit the surfaces above at the scaling stage even if the mandate and first-team work already happened.

Link to this answer: Does this still apply once we're past the pilot and scaling org-wide?
Find out before the org finds out

A mandate with no paved road is one of five ways a rollout stalls, and one power user rarely represents the org. A Technical Diligence review checks whether your own rollout is actually built on that road.

Source: https://customlabs.io/adoption/

navigate select esc close