# The tools work. Getting 180 engineers to actually use them is the other half of the job. Source: https://customlabs.io/adoption/ Updated: 2026-09-20 Adoption # The tools work. Getting 180 engineers to actually use them is the other half of the job. Six places a rollout stalls, 24 practices split by who owns each move, five ways a programme fakes momentum, and the numbers that separate real adoption from seat activation. Most of what kills an agent rollout has nothing to do with the model. Updated September 20, 2026 · First published August 14, 2026 · 33 min read · Key takeaways - → Seat activation counted as adoption is a named trap: a licence bought isn't a habit formed. - → A mandate with no paved road is one of five ways a rollout stalls. - → A new joiner's time to first merged agent-assisted change is one of six scoreboard numbers. - → The paved road and absorbing the output are named as separate surfaces from the mandate. - → One power user's numbers rarely generalise to the rest of the org. - → 24 practices span six surfaces, from the mandate to sustaining the rollout. A working pilot proves the harness can work somewhere. It says nothing about whether the other engineers who weren't in the room will pick it up, whether the review queue can absorb five times the diffs, or whether anyone can tell a genuine improvement from a number that just looks good next to no baseline. That gap between a proven pilot and an organisation that actually runs on it is where most rollouts quietly stall. The [Agentic Delivery Playbook](https://customlabs.io/agentic-delivery/) covers the machine: task definition, isolation, stage gates, spend. This page is what has to be true in the organisation around that machine before it produces anything at scale. The six surfaces ## Where a rollout actually stalls. Each surface fails a different way. None of them substitute for each other. An agent rollout moves through six stages, from the first mandate to a sustained program. Measurement results decide whether the pilot expands or holds. 01 ### The mandate **Covers:** The problem the rollout is meant to solve, one named accountable sponsor, and a written definition of success agreed before the first licence is bought. **Breaks when:** The mandate is "roll out AI" with no named sponsor, so once the pilot's numbers turn ambiguous there is nobody left with the authority, or the obligation, to make the call. **Watch:** Time between the first licence purchased and a written definition of success 02 ### The first team **Covers:** Choosing a pilot team whose workload, codebase health, and appetite actually generalise, and naming the exit-from-pilot criteria before the pilot starts. **Breaks when:** The pilot team is the platform group's best engineer working in a clean, greenfield service, so the result never predicts what a legacy team carrying real technical debt will see. **Watch:** Share of exit-from-pilot criteria written down before the pilot's first week 03 ### The paved road **Covers:** The sanctioned harness and defaults, sandboxing, repo hygiene, agent instructions files, tests that actually gate a merge, the boundary of what an agent may touch, and enablement built as pairing in the team's own repo. **Breaks when:** Rollout ships a licence and a slide deck with no sanctioned default, so every team improvises its own [guardrails](https://customlabs.io/glossary/guardrails/) at a different level of care, and the resulting friction gets blamed on the tool. **Watch:** Share of teams running on the sanctioned harness, not a bespoke one 04 ### Absorbing the output **Covers:** Review capacity, CI throughput, spec quality, who owns the merge, and how the senior engineer's role changes once writing code stops being the bottleneck. **Breaks when:** Agents multiply the diffs a team produces and review capacity never moves, so the queue that used to be "waiting on code" becomes "waiting on review," at the same total cycle time. **Watch:** Review latency on agent-authored changes against the team's own pre-rollout baseline 05 ### Measurement **Covers:** The baseline captured before rollout, every adoption number paired with a quality number, and the split between the median engineer and the power user. **Breaks when:** Nobody captured a pre-rollout baseline, so a genuinely improved number after rollout and a number that just looks better than a guess are indistinguishable. **Watch:** Whether a dated pre-rollout baseline exists for every metric on the scoreboard 06 ### Sustain **Covers:** Plateau and regression past the launch spike, spend per team, unsanctioned use read as a signal, retiring pilot scaffolding on schedule, and a named platform owner with a refresh cadence. **Breaks when:** Adoption plateaus at the enthusiasts within the first month, and the rollout reports that plateau as steady-state success because nobody asked why everyone else stopped coming back. **Watch:** Active weekly usage among engineers with tool access, tracked for months past the launch spike rather than only the first one The practice bank ## 24 practices, six surfaces. Filter by how far a practice has to reach (pilot, team, org), then copy the visible list as a Markdown checklist. Pilot Team Org Showing all 24 practices Copy as Markdown ### 01 The mandate Every agent rollout starts with an announcement, and most of them skip the part that would make the announcement mean anything: a single named owner and a written definition of success, agreed before the first licence is bought. Without that, the rollout runs on borrowed authority — a budget gets approved because a competitor shipped something, or because a demo landed well, and the actual question, what problem this is supposed to solve and how anyone will know it worked, never gets answered in writing. The gap doesn't show up immediately. It shows up months in, when the pilot's numbers are ambiguous and there's no sponsor left to make the call, and no written definition of success to make the call against. #### Name one accountable sponsor before the first licence is bought Org Sponsor **What it covers** A single named executive accountable for whether the rollout succeeds, distinct from whoever runs the pilot day to day. **How to build it** Put a name and a title on the mandate document itself, signed before procurement starts, not assigned after the pilot already has momentum. **Produces** A single person to ask when the rollout stalls, instead of a committee nobody can point to. **Tradeoff** A sponsor who changes teams or leaves takes the mandate's authority with them. Re-confirm ownership at every leadership change; don't assume it survives one. **Prove it** The current sponsor can state the rollout's definition of success from memory, not from re-reading the deck. #### Write the definition of success before buying anything Org Sponsor **What it covers** A short, specific statement of what the rollout is meant to change, agreed in writing before the first licence or seat is purchased. **How to build it** Draft it as a testable claim rather than a mission statement, with a stated target and a date, and get the sponsor to sign it. **Produces** A rollout that can be judged against what it said it would do, rather than reinterpreted after the fact to match whatever happened. **Tradeoff** A specific target invites scrutiny a vague one avoids. That scrutiny is the point. **Prove it** The definition of success predates the first purchase order, checked against the actual dates. #### Name the problem the rollout is for, not the tool it buys Org Sponsor **What it covers** A stated business problem, slow delivery, a backlog nobody can clear, an unmet SLA, that the rollout is supposed to move, ahead of any product decision. **How to build it** Write the problem statement before the vendor conversation, and route every later tool choice back through whether it moves that specific problem. **Produces** A rollout that can say no to a feature or a vendor that doesn't serve the actual problem, instead of collecting capability for its own sake. **Tradeoff** A narrow problem statement rules out adjacent wins the tool could also deliver. Note them; don't chase them yet. **Prove it** Someone outside the rollout team can restate the problem in one sentence without reading the deck. #### Name what the rollout is explicitly not trying to do Org Sponsor **What it covers** A short, written list of outcomes the rollout is not being judged on this cycle: headcount reduction, replacing a specific role, or a fixed seat-count target. **How to build it** Add a non-goals section to the mandate document itself, agreed by the sponsor, and repeat it in the first team-facing announcement. **Produces** Fewer engineers treating the rollout as a threat to route around quietly, because the thing they were worried about was already ruled out in writing. **Tradeoff** A non-goal stated today can look naive in a year if the plan genuinely changes. Revisit it on the same cadence as the mandate itself, and say so out loud when it does. **Prove it** A team lead can quote the non-goals list without checking, and it matches what leadership is actually saying in private. ### 02 The first team The team that pilots an agent rollout decides more about the result than the tool does. A team picked because its manager was enthusiastic, its codebase was already clean, or its tickets were unusually well scoped produces a pilot that looks great and predicts almost nothing about the teams waiting behind it. The fix isn't finding the best team; it's finding a representative one, and naming what "worked" means before the pilot starts, so the decision to expand isn't made by whoever tells the best story afterward. #### Choose the pilot for a workload that generalises, not for who's willing Pilot Platform **What it covers** A pilot team selected because its workload shape (ticket size, codebase age, test coverage) resembles the org's median, not because its manager volunteered first. **How to build it** Score candidate teams against the organisation's actual workload distribution before asking anyone if they want to go first. **Produces** A pilot result that predicts what the rest of the organisation will see, instead of a best case nobody else can reach. **Tradeoff** A representative team is harder to recruit than an enthusiastic outlier. Recruit on the strength of the mandate, not on who happens to be excited. **Prove it** The pilot team's ticket-size and test-coverage numbers sit within a stated range of the org median, checked against real data, not a guess. #### Write the exit-from-pilot criteria before the pilot starts Pilot Platform **What it covers** A numeric bar (acceptance rate, review latency, a rollback-rate ceiling) the pilot has to clear before the rollout expands past it. **How to build it** Agree the bar with the sponsor before the pilot's first week, and publish it to the pilot team so they know what they're being measured against. **Produces** An expansion decision made against a rule agreed in advance, instead of a judgment call once the pilot has been running a while. **Tradeoff** A bar set too high can kill a pilot that was actually working, just slowly. Calibrate it against the first team's baseline, not an aspirational number. **Prove it** The exit-criteria document predates the pilot's first week, and the expand-or-hold decision cites it directly. #### Check the pilot codebase's health before committing to it Pilot Platform **What it covers** A quick read on the pilot repo's test coverage, CI speed, and documentation state, since an agent inherits whatever review and test discipline already exists there. **How to build it** Run the same health check the team would apply before any other significant investment, and treat a repo with no tests or a broken CI pipeline as a reason to fix that first, or pick a different pilot. **Produces** A pilot result attributable to the tooling, not to a codebase that happened to make everything look easy. **Tradeoff** Waiting for a healthier repo costs calendar time a sponsor may not want to spend. Weigh that against a result nobody trusts once it doesn't repeat elsewhere. **Prove it** The pilot repo's coverage and CI numbers were checked and recorded before the pilot's first change merged. #### Ask the first team whether it actually wants this, and believe the answer Pilot Team lead **What it covers** A direct check with the candidate pilot team on their appetite for the change, distinct from their manager's enthusiasm. **How to build it** Talk to the engineers who'll do the work before the decision is finalised, and treat a genuinely reluctant team as a signal to pick a different one, not a change-management problem to push through. **Produces** A pilot that isn't quietly sandbagged by a team that never wanted it in the first place. **Tradeoff** Deferring to appetite can mean picking a less representative team. Weigh that against a pilot result nobody believes because the team's heart wasn't in it. **Prove it** The pilot team can name, unprompted, something about the rollout they are curious about, rather than only what they were told to expect. ### 03 The paved road A licence and a link to the vendor's documentation is not a paved road. It's the absence of one, dressed as adoption. What actually makes an agent effective in a codebase — a sandboxed environment, a repo-local instructions file, tests that genuinely gate a merge, and a stated boundary of what the agent may touch — has to exist before a second team is asked to use any of it. Enablement that's a worked example in the team's own repo, done alongside a person who already knows the harness, beats a webinar every time, because it's the only version that survives contact with a real ticket. #### Ship one sanctioned default harness before asking teams to adopt anything Team Platform **What it covers** A single supported configuration, sandboxing, a default instructions file, credential scoping, that every team starts from, rather than a licence and a link to the vendor's documentation. **How to build it** Build and pilot the default harness on the first team, then package it as the thing a new team installs on day one, not the thing they eventually converge on after improvising for a month. **Produces** A second team's onboarding measured in days against a working baseline, instead of weeks spent reinventing the first team's decisions. **Tradeoff** A single sanctioned default won't fit every workload exactly. Let teams diverge deliberately once they can name why it doesn't fit, not by default. **Prove it** A new team's first week of agent-assisted changes runs on the sanctioned harness, checked against their actual configuration, not the onboarding checklist's claim. #### Give every repo an agent instructions file, not a wiki page Team Engineer **What it covers** A file in the repository itself stating the tests that gate a merge, the boundary of what an agent may touch, and the house conventions a human reviewer would otherwise have to explain in every pull request. **How to build it** Write the first version from the pilot team's actual review comments, then require it as part of onboarding any new repo onto the paved road. **Produces** An agent, and a new human hire, that reads the same file the team actually works from, instead of a wiki page that drifted out of date months ago. **Tradeoff** A repo-local file needs its own upkeep, one more thing that can fall stale. Tie updates to the same review that catches other doc drift, not a separate process nobody owns. **Prove it** A repo picked at random has an instructions file that matches its actual current test-gating and merge rules, not last quarter's. #### Name the boundary of what an agent may touch before it touches it Team Platform **What it covers** An explicit, written scope of which files, systems, and actions an agent is allowed to change unsupervised, and which always route to a human. **How to build it** Start narrow, read and propose, never merge to a protected branch unattended, and widen the boundary only after the review data supports it, not on day one. **Produces** A rollout that can say precisely what went wrong when something does, instead of debating afterward whether the agent should have been allowed to do it. **Tradeoff** A narrow boundary limits how much benefit teams see early, which is real friction against the mandate's own success definition. Widen it on a measured schedule, not to quiet the complaint. **Prove it** The current boundary is written down somewhere a new team member actually reads, and the last incident review checked the action against it. [AI Security Review](https://customlabs.io/security-review/) #### Enablement is pairing in the team's own repo, not a webinar Team Platform **What it covers** A worked example built against the team's actual codebase and backlog, walked through with a person in the room, rather than a generic training deck. **How to build it** Pair a platform engineer with two or three engineers from the new team for their first real ticket, then let them run the second one alone. **Produces** A team that has seen the harness work on a ticket they recognise, instead of a slide deck they nodded through. **Tradeoff** Pairing doesn't scale to a large organisation by headcount alone. Train a small number of team-level champions this way, then let peer diffusion carry it past the platform team's own capacity. **Prove it** A newly onboarded engineer can name the specific ticket they paired on, rather than only that a training session happened. ### 04 Absorbing the output An agent that writes code faster changes nothing on its own if the team downstream can't absorb what it produces. Review capacity, CI throughput, and spec quality were sized for the previous rate of change, and a rollout that doesn't resize them just moves the bottleneck from writing code to reviewing it, at the same overall speed. The senior engineer's job changes here too, whether anyone planned for it or not: once code stops being the scarce thing, judgment about what to build and what to approve becomes the actual work, and a team that never says so out loud ends up with its senior engineers quietly competing with agents at the one task that stopped being the bottleneck. #### Size review capacity to the new diff volume before it arrives Team Team lead **What it covers** A concrete plan for who reviews the extra pull requests an agent-assisted team produces, agreed before rollout, not discovered as a growing queue after. **How to build it** Model the expected diff volume from the pilot's own numbers, and check whether current reviewers can absorb it inside their existing hours, or whether the review step itself needs a design change. **Produces** A rollout that doesn't quietly convert a code-writing bottleneck into a review-queue bottleneck of the same size. **Tradeoff** Adding review capacity has a real cost, whether that is headcount or a deliberately slower ramp. Both are cheaper than a queue nobody planned for. **Prove it** The team's review latency after rollout sits at or below its pre-rollout baseline, checked directly rather than assumed from the plan. #### Name who owns the merge on an agent-assisted change Team Engineer **What it covers** A stated rule for who is accountable for a merged change's correctness, regardless of whether an agent or a human wrote the diff. **How to build it** Keep the existing human-approver requirement on every merge, and say explicitly that approving an agent-authored diff carries the same accountability as approving a human-authored one. **Produces** A merge queue where "the agent wrote it" is never a legitimate answer to "who approved this." **Tradeoff** This adds no new process, which is the point, but it does mean reviewers can't treat an agent's diff as lower stakes to wave through. Say that out loud; don't leave it implied. **Prove it** A reviewer asked why they approved a specific agent-authored change gives the same kind of answer they would give for a human-authored one. #### Raise spec quality before generation, not after review Team Team lead **What it covers** A minimum bar for what a ticket needs to contain (acceptance criteria, the files in scope, what "done" looks like) before it is handed to an agent, matching what a competent engineer would want before starting the same ticket. **How to build it** Reject or return a ticket that fails to clear the bar at triage, the same discipline a team would apply before assigning it to a person. **Produces** Fewer plausible-looking diffs against a vague ticket, the single biggest driver of review time on agent-assisted work. **Tradeoff** A stricter triage bar slows down ticket creation, which reads as friction against exactly the speed the rollout is supposed to deliver. It is friction that pays for itself in review time saved. **Prove it** A sample of tickets from the last sprint each clear the stated bar, checked against the ticket text itself, not the triage process's own account of it. #### Redefine what the senior engineer's job is once writing code isn't the bottleneck Team Team lead **What it covers** An explicit statement that review, spec-writing, and architectural judgment are now the senior engineer's primary output, replacing an implicit expectation that their value is measured in lines shipped. **How to build it** Update how senior engineers are evaluated and staffed to credit review quality and spec clarity alongside their own commit volume, before the incentives quietly punish the people doing the most valuable work. **Produces** Senior engineers who spend their time where the bottleneck actually moved, instead of competing with agents at the one thing agents are now faster at. **Tradeoff** Changing how a role is evaluated is a genuine management change, not a tooling rollout, and it lands slower than the harness itself. Start the conversation with the pilot team early, before the mismatch shows up in someone's review. **Prove it** A senior engineer on the pilot team can describe their own role change in their own words, not the rollout announcement's words. ### 05 Measurement A number reported after rollout with no baseline before it is an anecdote wearing a dashboard. Most programmes never capture what cycle time, review latency, or change-failure rate looked like before the harness arrived, which means every improvement claim afterward is measured against a guess. The fix is procedural, not statistical: take the baseline in week one, on every team the rollout plans to compare later, and pair every adoption number that follows with the quality number that would catch it flattering the programme instead of describing it. #### Capture the baseline before rollout, not the week after Pilot Platform **What it covers** A recorded measurement of every scoreboard metric (cycle time, review latency, change-failure rate) for the pilot team, taken before the harness is installed. **How to build it** Run the measurement pass as the pilot's actual first step, before the first agent-assisted change merges, and store it somewhere the post-rollout comparison can actually reach. **Produces** A real before-and-after comparison, instead of a post-rollout number with nothing honest to compare it against. **Tradeoff** Measuring a baseline delays the pilot's visible start by a sprint or two. That delay is what makes the eventual comparison mean anything at all. **Prove it** A dated baseline snapshot exists for the pilot team, and it predates the first agent-assisted merge, checked against the actual timeline. #### Pair every adoption number with a quality number Org Platform **What it covers** A rule that no adoption metric (seats activated, weekly active users, diffs merged) gets reported without the paired quality metric that would catch it flattering the programme. **How to build it** Publish the two numbers together on the same dashboard, in the same report, every time, never the adoption number alone. **Produces** A scoreboard that can't be read as "adoption is working" when the actual story is usage up and quality flat or worse. **Tradeoff** A paired number is more work to compute and can undercut a headline the sponsor wanted to lead with. Report it anyway; the alternative is a number nobody trusts six months later. **Prove it** The last quarterly rollout report shows every adoption figure with its paired quality figure on the same page. #### Measure the power user and the median engineer separately Org Platform **What it covers** A reporting split between the top decile of usage and the median engineer, since a rollout's average is usually carried by a small number of enthusiasts. **How to build it** Segment every adoption metric by usage decile before reporting an org-wide average, and state both the median and the top decile explicitly. **Produces** An honest read on whether the rollout reached the organisation or just its most eager early adopters. **Tradeoff** A segmented report is a less flattering headline than a single averaged number. It is also the only version that predicts what happens when the enthusiasts change teams. **Prove it** The current adoption report states the median engineer's usage figure, not only the mean. #### Measure cost per merged change, not cost per seat Org Sponsor **What it covers** A cost figure denominated in what actually shipped, a merged, reviewed, non-reverted change, rather than a per-seat licence cost that says nothing about output. **How to build it** Divide a team's total rollout spend by its count of merged agent-assisted changes over the same period, tracked on the same cadence as any other delivery metric. **Produces** A spend number that moves with actual delivery, so cost per outcome is visible before the finance review finds it first. **Tradeoff** Cost per merged change needs the same ledger discipline a fleet's own per-task spend tracking already requires. It doesn't fall out of a seat count for free. **Prove it** The last quarter's cost-per-merged-change figure for the pilot team is checked against actual billing, not estimated from seat count. [The AI Cost Model](https://customlabs.io/cost/) ### 06 Sustain Adoption that plateaus at the enthusiasts looks like success if the only number anyone tracks is whether the licence got used at all. The harder, more useful questions come later: why the other engineers stopped after week one, what the unsanctioned workarounds are actually saying about the paved road, and who owns the harness once the pilot team that built it has moved on to something else. A rollout that never revisits its own scaffolding, or never sets a cadence for refreshing the harness as the underlying tooling changes, ages into the exact rigidity it was supposed to replace. #### Read unsanctioned use as a signal about the paved road, not a compliance problem Org Platform **What it covers** A standing check for shadow use, engineers routing around the sanctioned harness with a personal account or an unapproved tool, treated first as evidence the paved road is missing something. **How to build it** Ask the engineers doing it what the sanctioned harness doesn't cover, before writing a policy that blocks the workaround. **Produces** A paved road that closes the actual gap driving the workaround, instead of a ban that pushes the same behaviour further out of sight. **Tradeoff** Treating shadow use as a discovery process is slower than blocking it outright, and it means tolerating some ungoverned use while the gap gets fixed, bounded by whatever the security review already requires. **Prove it** The last confirmed instance of unsanctioned use resulted in a named change to the sanctioned harness, rather than only a warning. [Glossary: Shadow AI](https://customlabs.io/glossary/shadow-ai/) #### Name a platform owner before the pilot team disperses Org Sponsor **What it covers** A team or role accountable for the sanctioned harness after the pilot ends, distinct from the pilot team itself, which will eventually move to other work. **How to build it** Stand up the ownership handover as a named step in the exit-from-pilot criteria, not an afterthought once the pilot team's attention has already moved elsewhere. **Produces** A paved road that survives the specific people who built it moving to a different project. **Tradeoff** A dedicated platform owner is a real, ongoing cost, not a one-time project expense. Budget it as such from the mandate stage, not as a surprise once the pilot ends. **Prove it** The current platform owner is not the same person who ran the original pilot, or an explicit reason is on record for why that hasn't changed yet. #### Retire pilot-specific scaffolding on a schedule, not by accident Org Platform **What it covers** A stated date to remove the temporary shortcuts, manual approvals, and extra logging a pilot runs that a scaled rollout shouldn't carry indefinitely. **How to build it** Write the retirement date into the exit-from-pilot criteria alongside the metrics bar, and treat scaffolding still in place past that date as a tracked debt item, not an invisible cost. **Produces** A rollout that doesn't quietly carry pilot-era overhead into a process meant to run at many times the pilot's scale. **Tradeoff** Some scaffolding turns out to still be load-bearing once removal is attempted. Treat that discovery as a reason to make it a permanent, supported control, not a reason to leave it as an unowned pilot leftover. **Prove it** The scaffolding list from the pilot's exit criteria shows each item as either retired or explicitly promoted to a supported control. #### Set a refresh cadence for the harness as it ages Org Platform **What it covers** A recurring, scheduled review of the sanctioned default, the harness, the instructions-file template, the tool boundary, against how the underlying agent tooling has actually changed since it launched. **How to build it** Put the review on the same calendar discipline as any other platform component's support lifecycle, with a named owner and a fixed interval, not a review triggered only by a complaint. **Produces** A paved road that keeps pace with what the tooling can now do, instead of ossifying at whatever the pilot happened to ship with. **Tradeoff** A fixed refresh cadence spends platform time on a review that sometimes concludes nothing needs to change. That is a cheaper outcome than a harness nobody revisits until it is visibly behind. **Prove it** The last scheduled harness review happened on the calendar date it was due, and produced either a change or a stated reason none was needed. No practices match that combination. Clear a filter to see more. Five ways a rollout lies to itself ## A busy dashboard is a claim, not proof. Every one of these looks like a working rollout right up until the number it was supposed to produce turns out to be hollow. ### Seat activation counted as adoption **Looks like:** A dashboard shows every engineer with a licence assigned and at least one login, and the rollout reports that as adoption. **Costs you:** A licence assigned and a login once says nothing about whether anyone landed a real merged change with it, so the rollout can report full activation while actual usage sits with a handful of engineers. **Fix:** Report the paired quality number alongside every adoption figure, and segment by usage decile so the median engineer's number can't hide behind the top decile's. [See the practice: Pair every adoption number with a quality number →](https://customlabs.io/adoption/#pair-every-adoption-number-with-a-quality-number) ### The power user's numbers generalised to the org **Looks like:** One enthusiastic team lead reports a large personal speedup, and the rollout report quotes that number as the expected result org-wide. **Costs you:** A number that's true for the one person already motivated to make the tool work says nothing about the engineers who haven't tried it yet, and setting expectations against it makes every ordinary result look like underperformance. **Fix:** Segment every metric by usage decile and state the median explicitly, alongside the number that made the best slide. [See the practice: Measure the power user and the median engineer separately →](https://customlabs.io/adoption/#measure-power-users-and-the-median-separately) ### Time saved with no throughput or quality change **Looks like:** Engineers report the tool feels faster, and the rollout counts that as a win. **Costs you:** If review absorbed the saved time — a bigger queue, more back-and-forth on agent-authored diffs — or the change-failure rate crept up at the same time, the time saved never reached a shipped, working outcome. It just moved the bottleneck downstream where nobody was watching. **Fix:** Size review capacity to the new diff volume before rollout, and track cycle time and change-failure rate as the numbers that would catch the trade actually happening. [See the practice: Size review capacity to the new diff volume before it arrives →](https://customlabs.io/adoption/#size-review-capacity-to-the-new-diff-volume) ### A mandate with no paved road **Looks like:** Leadership announces the rollout and a licence budget, and expects teams to work out the rest themselves. **Costs you:** Every team invents its own guardrails at a different level of care, engineers read the resulting friction as the tool being bad rather than the harness being missing, and the mandate's credibility erodes before the paved road ever ships. **Fix:** Ship one sanctioned default harness, piloted and packaged, before the mandate reaches teams beyond the pilot. [See the practice: Ship one sanctioned default harness before asking teams to adopt anything →](https://customlabs.io/adoption/#ship-one-sanctioned-default-harness) ### A pilot that depended on someone babysitting it **Looks like:** The pilot clears every exit criterion, and the team credits the harness. **Costs you:** If the actual reason it worked was one platform engineer quietly fixing broken configs and coaching people through failures all quarter, the result doesn't survive that person moving to something else, and the second team's onboarding fails in ways the pilot never surfaced. **Fix:** Enablement that is pairing and worked examples in the team's own repo, plus a named platform owner distinct from whoever ran the original pilot. [See the practice: Name a platform owner before the pilot team disperses →](https://customlabs.io/adoption/#name-a-platform-owner-before-the-pilot-team-moves-on) The scoreboard ## Six numbers, and how each one lies. Every one of these is measurable today. None of them is trustworthy read alone. Pair it with the number next to it. ### Share of engineers landing a merged agent-assisted change weekly **Why it matters** The most direct read on whether the rollout reached the organisation or just its early adopters. **How it misleads** A rising share can still be entirely the same enthusiasts merging more often, with no new engineers ever starting. **Pair with** Count of distinct engineers who landed their first agent-assisted merge this month ### Acceptance rate of agent-authored changes **Why it matters** Whether the diffs an agent proposes are actually good enough to ship, as opposed to merely numerous. **How it misleads** A high rate on a narrow, well-scoped ticket type says nothing about the harder tickets nobody's routed to an agent yet. **Pair with** Share of ticket types actually being routed to an agent, beyond the ones the rate is measured on ### Review latency and review-hours per merged change **Why it matters** Catches a code-writing bottleneck turning into a review-queue bottleneck of the same size. **How it misleads** Latency can look fine if reviewers are quietly working unpaid overtime to keep the queue moving; hours per change is what actually catches that. **Pair with** Reviewer-reported time spent per agent-authored change, tracked separately from latency ### Change-failure and rollback rate on agent-authored changes **Why it matters** Whether shipping faster came at the cost of shipping worse. **How it misleads** A flat rollback rate can hide a rise in near-misses caught in review that never reached production, which the rate itself never counts. **Pair with** Defects caught in review per agent-authored change, tracked as its own number ### Cost per merged change **Why it matters** The spend number that moves with real output, rather than with how many licences were assigned. **How it misleads** A falling cost per change can hide a rollback rate rising at the same time, if the change never gets checked against quality. **Pair with** Change-failure rate on the same set of changes, read on the same dashboard ### A new joiner's time to first merged agent-assisted change **Why it matters** The clearest read on whether the paved road actually onboards someone who wasn't there for the pilot. **How it misleads** A fast time for a joiner who happens to land on the original pilot team says nothing about a joiner on a team several steps away from the platform team's attention. **Pair with** The same figure measured on a team outside the original pilot What this is built from ## Verifiable, not claimed. No invented survey statistics. Just what's already documented on this site, and how it connects. - The Agentic Delivery Playbook is the machine half of this: task definition, isolation, stage gates, and spend, once a team is running a fleet. This page is what has to be true in the organisation around it for that machine to actually get used. [The Agentic Delivery Playbook](https://customlabs.io/agentic-delivery/) - The AI Release Path picks up once a change is ready to ship. This page is the layer above it: whether the organisation absorbing that change actually has the review capacity and the paved road to ship it well. [The AI Release Path](https://customlabs.io/release/) - The AI Cost Model's attribution surface is what the cost-per-merged-change practice above draws on directly: the same discipline for turning ledger noise into a number engineering and finance both trust. [The AI Cost Model](https://customlabs.io/cost/) - AI Security Review is the standing check the unsanctioned-use practice above routes back into: what a workaround actually has to clear before it becomes the sanctioned default. [AI Security Review](https://customlabs.io/security-review/) - The Delivery Record is the evidence this studio has for the absorption surface above: real defects a review and verify stage caught, and the build-time check each one left behind. [The Delivery Record](https://customlabs.io/delivery-record/) ### Sources - [NIST - AI Risk Management Framework (AI RMF 1.0)](https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf) The risk-management functions our governance and adoption controls map onto. Retrieved 2026-08-24. - [Stanford HAI - AI Index Report](https://aiindex.stanford.edu/report/) The annual report our adoption page's industry-wide statistics are drawn from. Retrieved 2026-08-24. Not sure the pilot would actually generalise A Ship Audit checks the paved road and the absorption math against your specific team before the rollout reaches the next 179 engineers. [Book a Ship Audit →](https://customlabs.io/diagnostic/ship-audit/) [See the delivery playbook →](https://customlabs.io/agentic-delivery/) Questions ## Before you announce a rollout. What teams ask us before they take an agent programme past the first team. 01 Isn't this just change management with new labels? + Change management is the general discipline. This is the specific version of it for an agent rollout: a paved road, review-capacity math, and a scoreboard an agent programme needs that a generic change programme doesn't ship with. 02 We already ran a successful pilot. Why isn't that enough? + A pilot proves the harness can work somewhere. It doesn't prove the result generalises to a team with a different workload, or that the organisation can absorb several times the review volume once the rollout scales past the pilot team. 03 What's the fastest way to kill a rollout without meaning to? + Skip the paved road and ship a licence with no sanctioned default. Every team improvises its own guardrails at a different level of care, the friction gets read as the tool being bad, and the mandate loses credibility before anyone built the thing that would have made it work. 04 How do we know adoption is real, rather than seat activation with nothing behind it? + Pair every adoption figure with its quality counterpart and check the median engineer alongside the top decile. Activation with no merged output and no movement in review latency is the seat-activation trap, and it looks identical to real adoption on a dashboard that only reports the average. 05 Who should own this, the platform team or engineering leadership? + Both, at different points. Leadership owns the mandate and the definition of success; the platform team owns the paved road and its upkeep once the pilot ends. A rollout missing either one stalls in a different place depending on which is missing. 06 Does this still apply once we're past the pilot and scaling org-wide? + Yes. Most of what actually breaks a rollout — absorption capacity, the median-versus-power-user gap, unsanctioned use — shows up after the pilot succeeds, not during it. Revisit the surfaces above at the scaling stage even if the mandate and first-team work already happened. Find out before the org finds out A mandate with no paved road is one of five ways a rollout stalls, and one power user rarely represents the org. A Technical Diligence review checks whether your own rollout is actually built on that road. [Book a Technical Diligence review →](https://customlabs.io/diagnostic/technical-diligence/) [Talk to us about strategy →](https://customlabs.io/services/strategy-architecture/)