# Are you ready to ship AI? Source: https://customlabs.io/tools/ai-readiness/ Updated: 2026-09-05 Tools # Are you ready to ship AI? Ten questions across data, evals, hardening, and ownership. Answer honestly, and you'll get a maturity tier and the 2–3 gaps actually stalling your pilot, not a generic score. Question 1 of 10 ← Back ## ← Back to last question ## Your readiness score is ready Enter your work email to see your tier, your dimension scores, and the 2–3 things stalling your pilot. Work email See my results → Where you're stalling A Ship Audit turns these gaps into a scored, board-ready readiness report and a pilot plan. [Book a Ship Audit →](https://customlabs.io/diagnostic/ship-audit/) How it scores ## What this measures The scorecard runs ten questions across six weighted dimensions, then rolls each one into a single score out of 100. A heavier weight moves the overall score harder when that dimension is weak. That is why a gap in eval maturity costs more than the same gap in team ownership. Every answer maps to a point value. The points roll up into a sub-score for that dimension, and the sub-scores combine into the score on the results screen. The scorecard also names the weakest dimensions on their own, since one overall number hides the detail a reader needs to act on. ### The six dimensions Every question belongs to exactly one dimension, and every dimension has a fixed number of points on offer. A heavier weight moves the overall score more per point than a lighter one. That is a deliberate choice, not an accident of question count. Dimension Weight Questions Max points Why it stalls a pilot What closes the gap Use-case clarity 1× 2 6 A fuzzy problem statement means no one can tell whether the pilot succeeded, so it never graduates. Write a one-page spec covering the user and the workflow. Name the single metric that defines success. Data readiness 1× 2 6 The model demos on a dataset that won't exist in production. Real data must be accessible, clean, and cleared for use. Inventory the data the use case needs and confirm access + rights before building anything. Eval & quality maturity 1.2× 2 6 Without evals you're shipping on vibes: every prompt or model change is a silent regression risk, so teams freeze. Stand up a small labelled eval set and score every change against it. Production hardening 1.1× 2 6 A prototype that ignores timeouts and provider outages breaks the first week real traffic hits it. Add retries and a fallback path before you widen traffic, and monitor cost and latency. Team & ownership 0.9× 1 3 With no clear owner, AI work stalls between teams and no one is accountable for getting it to production. Name a single accountable owner and give them the time to ship it. Security & governance 0.9× 1 3 Unresolved data-handling and access questions block the pilot at the security review, not the demo. Compliance is the usual reason a review says no. Get data-handling and access in front of security now, not at launch. ### The maturity tiers The weighted overall score lands in one of four bands, each with its own read on where the work actually stands. A team can move up a band by closing its weakest dimension rather than every dimension at once. Tier Score range What the band means Exploring 0–34 You're still mapping the problem. Most of what makes a pilot succeed hasn't started yet. Piloting 35–59 You've got the makings of a pilot, but the muscle that turns a demo into production (evals, hardening, ownership) is still thin. Scaling 60–79 The fundamentals are in place and you're pushing past a single pilot. What's left is what bites you at scale, not at the demo. Production-ready 80–100 You've got the use case, the data, the evals, and the ownership to run this in production responsibly. What this does not do A dimension counts as a gap once its own score falls at or below 60. The scorecard always names between 2 and 3 gaps, even when every dimension ties. The reader always leaves with something concrete to fix first. It trusts the answers given: it does not read eval logs or watch a team work. A generous self-assessment produces a generous score. Treat the tier as a starting hypothesis, not an audit finding. Two teams can land on the same tier for different reasons, and the dimension breakdown above is where those reasons show up. Answering the ten questions takes a few minutes. Closing even one real gap usually takes weeks, and the weightiest gap is where the scorecard points first. Prefer a human read? The scorecard is a free, self-serve preview of the Ship Audit, a fixed-scope, fixed-fee engagement that turns these same six dimensions into a written, board-ready readiness report and a scoped pilot plan. [Book a Ship Audit →](https://customlabs.io/diagnostic/ship-audit/)