CustomLabs
Choosing a partner

Choosing an AI delivery partner is a procurement decision as much as a technical one.

Five routes for staffing AI work, 24 questions to ask any of them across six gates, five named ways an engagement fails, and the cases where we're honestly the wrong choice. Filter to what matters, copy the questions into your own RFP.

Updated First published

28 min read

Markdown

Every guide on this site so far is written for an engineer deciding how to build something. This one is for whoever signs the contract deciding who builds it: hire in-house, a big consultancy, an offshore dev shop, a freelancer, or a specialist studio. That decision gets made constantly and gets a real checklist almost never.

This page is that checklist. Five routes, written straight, including the three we compete against. Twenty four questions any of them should be able to answer, with what a passing answer contains and what a failing one sounds like. Where we fit, and the plain cases where we don't.

The five routes

Five ways to staff this, honestly compared.#

Each one has a real strength. None of them wins every time, including us.

01

Hire in-house

Hiring one or more engineers directly onto your own payroll to own the AI work full-time, inside your existing team structure.

Best at
Institutional knowledge compounds. Someone who sits in your standups, knows your data model, and carries context across projects for years is a real advantage no outside party can match, and it only grows the longer they stay.
Where it breaks
Hiring takes months in a market where strong AI engineers have their pick of offers, and a single hire has no second opinion on their own architecture calls until the team grows past one person.
Time to first ship
3-6 months, once sourcing, interviewing, and onboarding are counted before the first meaningful production code.
Six-week tell
Six weeks in, the requisition still isn't filled, or it is and the new hire is still ramping on your codebase with nothing shipped yet — a signal worth naming in a leadership review, not a failure on its own.
02

Big consultancy or systems integrator

A large, multi-practice firm that staffs your engagement from a bench of consultants, usually with a partner-level relationship sitting above the team doing the actual work.

Best at
Scale and coverage. If the work spans a dozen legacy systems, needs a change-management program running alongside the code, or has to be staffed to twenty people next quarter, that is exactly the shape a large integrator is built to absorb.
Where it breaks
The people who ran the sales conversation are rarely the people writing the code, and staffing decisions can follow bench availability across the firm's other clients as much as fit for your specific problem.
Time to first ship
2-4 months to first production ship, once the statement of work, staffing, and onboarding for a multi-week engagement are counted.
Six-week tell
Six weeks in, the person who ran the pitch is nowhere near the daily calls, and the current team roster doesn't match the résumés circulated at kickoff.
03

Offshore or generalist dev shop

A software agency, often offshore or nearshore, that builds whatever a client asks for across a wide range of technology, with AI as one more line on the menu rather than the specialty.

Best at
Cost-effective throughput on well-specified work. Handing over a clear, detailed spec and getting a competent implementation back at a lower day rate than a specialist commands is a real, legitimate trade a generalist shop is built to deliver.
Where it breaks
AI work rewards judgment on ambiguous, under-specified problems more than execution on a fixed spec, and a shop staffed for throughput on clear tickets has little incentive, or practice, at pushing back on a weak brief.
Time to first ship
4-8 weeks to a first demo, often three to four times that to something that survives real production traffic, because eval discipline and guardrails are usually the first things a fixed-spec engagement skips.
Six-week tell
The demo looks finished, but nobody can answer what happens when the model is wrong, and there's no eval set behind the accuracy number they quoted.
04

Freelancer or fractional lead

An independent contractor, often a senior engineer moonlighting or between roles, engaged directly for a defined chunk of work or an ongoing fractional commitment.

Best at
Speed and cost per hour. A good freelancer with real production AI experience can move faster than any team fielding a kickoff deck and a project manager, at a fraction of a studio's or integrator's rate.
Where it breaks
One person is a single point of failure for both delivery and continuity, nobody reviews an architecture call before it ships, and a freelancer juggling several clients has a hard ceiling on how many hours are actually yours.
Time to first ship
2-6 weeks — a wide range, because it depends almost entirely on the individual and how much of their week is genuinely committed to you.
Six-week tell
You genuinely can't say what happens to the project if this person is unavailable for two weeks: no documentation, and no second person who could pick it up.
05

Specialist AI studio

A small team that works only on applied AI, engaged for a scoped build or an ongoing retainer, sized between a freelancer and a full consultancy practice.

Best at
Judgment on the ambiguous parts: the build-buy-or-skip call, the architecture fork, the eval discipline, because that judgment is the entire business, not one service line competing for staffing against ten others.
Where it breaks
Small teams have a real capacity ceiling: a studio sized for a handful of engagements at a time cannot absorb a client that needs twenty engineers next quarter, and a bad pick with no track record carries as much risk as a bad hire.
Time to first ship
4-8 weeks to a working, evaluated pilot, because AI judgment is what a studio optimizes for over raw headcount.
Six-week tell
There's still no eval set and no named production owner for the pilot — the same tell that sinks a bad engagement on any of the other four routes.
Question bank

24 questions, six gates.#

Filter to the gates that matter for your situation, then copy the visible list as a Markdown checklist for your RFP.

Diagram in six lanes, one per partner gate: evidence and track record; technical fit; and ownership and exit. The other three are delivery model and velocity; security and compliance; and commercials and risk. Each lane holds one node that names what that gate checks for. The gates connect in order, from evidence to commercials. EVIDENCE & TRACK RECORD TECHNICAL FIT OWNERSHIP & EXIT DELIVERY MODEL & VELOCITY SECURITY & COMPLIANCE COMMERCIALS & RISK Checkable evidence Eval discipline Transfers at handover Operating model Before or during review Risk gets priced
Six gates run in order before a contract gets signed. Each one tests something the last gate did not.

01 Evidence & track record

A sales deck proves a firm can present. It doesn't prove a firm can deliver. This gate replaces the pitch with something checkable: a named system still running today, the people who'll actually do the work, and a reference willing to talk about what went wrong.

Can you show me a named, live production system you built that's still running today?

Why you ask it
A portfolio slide with client logos proves a sales relationship, not a shipped system. You're checking whether the work survives contact with real users and real data after the invoice is paid.
What passes
A specific system, named, live today, with roughly when it shipped and what it does — a real reference you can call, not a case-study PDF.
What fails
A list of client logos with no named system, or "we can't say due to NDA" offered as the whole answer instead of a redacted but specific description.
The artifact
A reference call with the team that operates the system now, not the team that sold it.

What happens when your own delivery goes wrong, and can you show me an example?

Why you ask it
Every delivery team ships a defect eventually. The honest answer is whether they can name one and what changed afterward, not whether they claim a flawless record.
What passes
A specific, named incident with the root cause and what changed in their own process afterward, offered without being pressed for it.
What fails
A claim of a flawless track record, or a generic answer about continuous improvement with no specific incident behind it.
The artifact
A published incident record or postmortem, or a specific verbal example you can check with a reference.

Who specifically will be doing the work, and can I see their own production AI experience?

Why you ask it
The person in the sales call and the person writing the code aren't always the same person. You're checking the actual named individuals, not the firm's aggregate résumé.
What passes
Named individuals, with their specific prior production AI work described, before the contract is signed rather than after.
What fails
A generic "our team has years of combined experience" with nobody named, or a roster where the senior names roll off right after the sale.
The artifact
Named CVs or profiles for the actual staffed team, checked against who shows up to the first working session.

Have you shipped something at roughly our scale and data sensitivity before, or would this be a first for you too?

Why you ask it
A team's best work is usually close to what they've already done. A partner whose biggest prior engagement was a tenth your scale is taking on real, unstated risk on your project.
What passes
A specific precedent at comparable scale or sensitivity, or an honest "this would be a first for us at this scale" said up front rather than discovered later.
What fails
Confidence with no specific precedent offered, or scale questions answered with generic capability claims instead of a named comparable engagement.
The artifact
A description of the most similar prior engagement, detailed enough for you to judge the comparison yourself.

02 Technical fit

Model-agnostic architecture, real eval discipline, and production experience past launch day separate a partner who can ship a demo from one who can run a system for years. This gate checks for the second kind.

Is your architecture tied to one model provider, or can you actually swap models without a rewrite?

Why you ask it
A provider outage, price change, or deprecation isn't hypothetical. You're checking whether that day is a config change or a rewrite.
What passes
A named abstraction layer between the application and any single model provider, with a specific past example of a model swap they actually did.
What fails
Everything hardcoded to one provider's SDK, or "we'd just rebuild that part" offered as an acceptable answer to a provider outage.
The artifact
The architecture diagram showing the provider abstraction layer.

How do you know your system works before you ship it, and how do you know it still works after the next change?

Why you ask it
Anyone can demo a system that works once, on a case chosen to show it off. You're checking for a repeatable, checkable measurement, not a demo.
What passes
A labelled eval set with a specific pass/fail condition per case, run in CI on every change, with a current pass rate they can show you.
What fails
"It works well in our testing" with no eval set named, or evals described as a one-time activity before launch rather than a running gate.
The artifact
An eval report, or a description specific enough that you could ask to see one before signing.

What's the most complex production AI system you've operated after launch, past the day it shipped?

Why you ask it
Building a demo and operating a system through months of real traffic are different skills. This checks for the second one specifically.
What passes
A specific system with a description of what broke after launch and how it was caught — which only exists if someone actually operated it.
What fails
Every example stops at launch day, with no mention of what happened afterward or how a regression would even be noticed.
The artifact
A description of the observability or monitoring setup on a real post-launch system.

Walk me through a build-versus-buy or agent-versus-pipeline call you made recently, and why it went that way.

Why you ask it
The reasoning behind a past architecture decision reveals more about judgment than any credential. You're checking whether the fork was decided on purpose or defaulted into.
What passes
A specific past decision with the actual tradeoff named, including a case where the answer was the simpler option rather than the flashiest one.
What fails
Every past decision defaults to the most complex or most current option, with no example of choosing the simpler path when it was the right call.
The artifact
A specific past project where the simpler option won, described in enough detail to check.

03 Ownership & exit

The point of paying someone to build this is that you end up owning it. This gate checks what actually transfers at handover, and what it costs to leave if the relationship stops working.

Do we own the source code, prompts, and eval sets outright at handover, or are we licensing something back from you?

Why you ask it
The difference between owning the work and licensing access to it is the difference between real independence and a vendor relationship. You're checking which one you're signing.
What passes
A written clause stating source code, prompts, and eval sets transfer to you in full at handover, with no ongoing dependency on the vendor's own tooling to run or modify them.
What fails
Ownership described verbally but absent from the contract, or a clause that transfers code but not the eval sets or infrastructure-as-code needed to actually run it.
The artifact
The ownership clause in the signed engagement contract.

If we wanted to walk away and hire someone else tomorrow, what would that actually take?

Why you ask it
A partner confident in their own value has an easy answer to this. One relying on lock-in to keep the relationship does not.
What passes
A specific, short answer: handover documentation, a knowledge-transfer session, and the code and credentials you already own — nothing else required.
What fails
Hesitation, a vague answer, or a discovery that key infrastructure or credentials sit in the vendor's own accounts rather than yours.
The artifact
A documented handover checklist, ideally reviewed before the engagement starts, not written after you ask to leave.

What documentation and runbooks do we get, and can someone outside your team actually operate this from them?

Why you ask it
Undocumented systems create a dependency on the specific people who built them, whether or not that was the intent. You're checking whether the knowledge actually left their heads.
What passes
Runbooks and documentation specific enough that a competent engineer who never worked on the project could run and modify the system from them alone.
What fails
"We'll walk your team through it" offered as a substitute for written documentation, or documentation that turns out to be a slide deck rather than an operational runbook.
The artifact
The actual documentation, reviewed for specificity before the engagement closes, not promised for later.

What happens the week after handover if something breaks, and is that covered or a new invoice?

Why you ask it
The gap between the engagement ending and being fully on your own is exactly where an undocumented dependency shows up. You're checking the terms of that gap before you're in it.
What passes
A stated support window after handover, with clear terms for what's covered and what triggers a new engagement, agreed before work starts.
What fails
No stated post-handover terms at all, discovered only when something breaks and a new statement of work appears.
The artifact
The support or warranty clause in the contract.

04 Delivery model & velocity

Most AI engagements don't fail on the model. They fail on process: no named owner, no visibility into a bad week, scope that drifts with no way to renegotiate it. This gate checks the operating model, not the technology.

Once this ships, who is the named owner on our side and on yours, and does that person actually have the authority to act?

Why you ask it
A pilot with no owner is a pilot that quietly stalls after the demo. You're checking that a specific, accountable person exists on both sides before it ships, not after it stalls.
What passes
A named individual on each side with real decision authority, agreed before the pilot starts rather than assigned informally after launch.
What fails
Ownership described as "the team" rather than a person, or a named owner with no real authority to approve changes.
The artifact
The ownership line in the project charter or kickoff document.

How will I see progress week to week, and what does a bad week actually look like when it happens?

Why you ask it
A partner who only ever shows you good news has already decided what you're allowed to see. You're checking for a cadence that surfaces a problem the same way it reports one going well.
What passes
A specific, regular cadence, such as a demo, a written update, or a dashboard, that would show a stalled or blocked week as clearly as a good one.
What fails
Updates that are exclusively verbal reassurance with no artifact, or a cadence that only exists when things are going well.
The artifact
A sample status update or dashboard from a past engagement.

When the scope needs to change mid-engagement, because it usually does, what actually happens?

Why you ask it
Scope drift is normal on real AI work. What matters is whether there's a defined, fast path to renegotiate it, or whether it just quietly expands the original quote.
What passes
A named change-control process with a stated turnaround time, agreed before the engagement starts.
What fails
No defined process, with scope changes handled ad hoc and disputes about what was "really" agreed only surfacing after the fact.
The artifact
The change-control clause in the statement of work.

What does the path from pilot to production actually look like, and who decides the pilot is done?

Why you ask it
A pilot with no graduation criteria can run indefinitely, or die quietly, with nobody able to say which. You're checking that both outcomes have a name and an owner.
What passes
Stated success criteria for the pilot, agreed up front, with a named decision point and decision-maker for graduating to production or stopping.
What fails
"We'll know it when we see it" as the entire graduation criteria, with no stated metric or decision point.
The artifact
The pilot success criteria, written down before the pilot starts.

05 Security, privacy & compliance

Whatever a partner builds eventually meets your own InfoSec, privacy, and procurement review — the same four gates covered on our security review page. This gate checks whether that work happens before the review, or during a rejected one.

Could this system pass an internal security or privacy review today, and if not, what specifically is missing?

Why you ask it
A system that works in a demo but was never built with review in mind usually needs weeks of retrofitting before it can ship. You're checking whether that work already happened or is still ahead of you.
What passes
A specific, honest list of what's ready and what isn't, mapped to the actual gates your own reviewers will run.
What fails
An assumption that security review is a formality to handle later, or unfamiliarity with what your own reviewers will actually ask.
The artifact
A completed pre-flight artifact list against your own review gates.

Where does our data actually go once it leaves our systems, and is that written down?

Why you ask it
Model providers, subprocessors, and logging destinations form a real chain, and a vague answer here usually means nobody has mapped it, not that the chain is simple.
What passes
A named data-flow description covering every model provider and subprocessor your data reaches, checked against your own compliance requirements.
What fails
"It's all handled securely" with no specifics, or a chain that turns out to include a subprocessor nobody mentioned.
The artifact
The data-flow diagram and subprocessor list.

If this system can take actions as well as answer, what stops it from doing something irreversible by mistake?

Why you ask it
An agent that can write, send, or delete inherits real consequences from a hallucinated call or a successful prompt injection. This checks whether that risk was designed around or just assumed away.
What passes
Named tool scoping to least privilege, with a human checkpoint on anything irreversible, described specifically rather than asserted generally.
What fails
A claim that "the model is pretty reliable" offered as the actual control, with no scoping or checkpoint named.
The artifact
The tool/permission matrix.

Which compliance framework does our industry actually map to, and how does that change what you build?

Why you ask it
GDPR, SOC 2, and industry-specific rules change what has to be logged, retained, and disclosed. A partner who can't name yours specifically is guessing at your actual requirements.
What passes
The specific framework named for your situation, with a concrete example of how it changes a design decision, not a generic list of acronyms.
What fails
Every industry gets the same generic compliance answer, with no framework named specifically for your situation.
The artifact
A written mapping of your compliance requirements to specific design decisions.

06 Commercials & risk

Pricing model, real production cost, and a reference call with a client who left are where an engagement's actual risk gets priced, not in the rate card.

Is this priced fixed-fee, time-and-materials, or value-based, and does that match who's actually holding the risk?

Why you ask it
A pricing model quietly assigns risk: time-and-materials rewards slow delivery, and a fixed fee on an unscoped problem punishes honest scope discovery. A mismatch shows up later as friction.
What passes
A pricing model matched to the actual scope certainty, explained plainly, with the tradeoff named rather than glossed over.
What fails
A pricing model presented as the only option with no explanation of the tradeoff, or open-ended time-and-materials billing on a poorly scoped problem.
The artifact
The pricing clause in the proposal, with the model and the reasoning behind it.

What's the actual cost per completed task if this runs in production, and what's the worst case if it misbehaves?

Why you ask it
A sticker-price estimate routinely undercounts real spend once retries, fallbacks, and agent loops are counted. You're checking whether the number you're quoted is the real one.
What passes
A modeled cost per successful outcome, not per API call, plus a stated ceiling on worst-case spend given an enforced budget.
What fails
A per-token estimate presented as the total cost, with no mention of retries, fallbacks, or a spend ceiling.
The artifact
The cost model, including the worst-case scenario.

What specifically triggers a change order versus what's included in the original quote?

Why you ask it
A quote with no stated boundary invites scope disputes later. You're checking that the line between included and extra is named up front, not discovered mid-engagement.
What passes
A specific, written boundary between what's included and what triggers additional cost, agreed before work starts.
What fails
"That wasn't included" surfacing for the first time after the work is already underway, against a quote with no stated boundary.
The artifact
The scope-boundary section of the statement of work.

Can I talk to a past client who didn't renew or extend, not just one who did?

Why you ask it
A reference list curated entirely from happy renewals tells you about the best case, not the honest average — the willingness to offer a mixed reference is itself informative.
What passes
A willingness to offer, or at least discuss, a reference from an engagement that ended rather than renewed, with an honest account of why.
What fails
Every reference offered is a renewal or expansion, with no acknowledgment that an engagement has ever ended for a reason worth naming.
The artifact
A reference call with a past client, including one that didn't continue.
Five ways it goes wrong

These sink an engagement, not a model.#

Each one is a contract problem with a contract fix, not a technology problem.

The pilot with no production owner

Symptom: A pilot ships, gets a good reaction in the room, and then sits untouched for months because nobody on either side is actually responsible for pushing it into production.

Root cause: The kickoff named a project, not a person. Accountability was implied by "the team is on it" rather than assigned to someone with the authority to act.

The clause that prevents it: Name a production owner, on both sides, in the kickoff document itself, before the pilot starts, with an explicit decision date for graduate-or-stop.

Most common in Big consultancy or systems integratorOffshore or generalist dev shop

Demo-driven scope

Symptom: The system looks finished in a conference room on curated examples, then falls apart on the first messy, real-world input a week after launch.

Root cause: The engagement optimized for the thing that gets sign-off, a good demo, rather than the thing that survives production, because nobody wrote down what "done" actually meant.

The clause that prevents it: Define the eval set and its pass/fail conditions in the statement of work itself, before a single demo is built.

Most common in Offshore or generalist dev shopFreelancer or fractional lead

The Eval Stack

The black-box handover

Symptom: The engagement ends, the invoice is paid, and the client discovers nobody in-house actually understands how the system works or can safely change it.

Root cause: Documentation and knowledge transfer were treated as a nice-to-have at the end of the project rather than a deliverable specified at the start.

The clause that prevents it: Make documentation and a knowledge-transfer session an explicit, named deliverable in the contract, with its own sign-off separate from the code shipping.

Most common in Big consultancy or systems integratorFreelancer or fractional lead

Model lock-in by omission

Symptom: A provider raises prices or deprecates the exact model the system depends on, and switching turns out to mean a rewrite, not a config change.

Root cause: Nobody made model-agnosticism an explicit requirement, so the fastest path to a working demo, calling one provider's SDK directly everywhere, became the permanent architecture.

The clause that prevents it: Require a model-agnostic abstraction layer as a named technical requirement in the contract, not an assumed best practice.

Most common in Offshore or generalist dev shopFreelancer or fractional lead

Glossary: Model-Agnostic Architecture

Seniority bait-and-switch

Symptom: The senior names who ran the pitch and impressed everyone in the room are nowhere near the team actually doing the work three weeks later.

Root cause: Staffing on the proposal was aspirational, or a sales asset, and nothing in the contract tied the named individuals to the delivered work.

The clause that prevents it: Name the actual staffed individuals in the contract, with a defined process, and client sign-off, for any substitution.

Most common in Big consultancy or systems integrator

A technical sibling to several of these already exists on our failure modes reference.

Our answers

Our answers, with receipts.#

The sharpest handful of our own 24 questions, each pointing at something already published rather than a new claim.

When we're wrong for you

Honest disqualifiers.#

If any of these describe you, one of the other routes above, or the resource we point at, fits better than an engagement with us.

This page describes the questions a buyer typically needs answered before signing an AI delivery contract, and the evidence that tends to answer them well. It is not legal or procurement advice; your own legal and procurement functions make the actual sign-off call for your organization.

Ready to run this against a specific partner

A Ship Audit turns an unscoped idea into a written plan before you take it to any of the five routes above, including us.

Questions

Before you sign anything.#

What buyers ask us before they run this checklist against a partner.

01 Is this page trying to talk me out of hiring you?

No, it's trying to give you a real decision framework. Two of the five routes and several of the disqualifiers point away from us on purpose — a guide that concludes "hire us" regardless of the question isn't a guide, it's an ad.

Link to this answer: Is this page trying to talk me out of hiring you?
02 Do all 24 questions apply to every engagement?

No — filter to the gates that matter for your situation. A ten-day proof of concept doesn't need the full commercials gate; a system touching regulated data needs every question in the security gate answered before anyone signs anything.

Link to this answer: Do all 24 questions apply to every engagement?
03 What if a partner can't answer one of these yet?

A specific "we don't have this yet, here's the plan and the date" is a better answer than a vague or invented one — the same standard our security review page holds reviewers to.

Link to this answer: What if a partner can't answer one of these yet?
04 Can I use this if I've already picked a partner?

Yes. Run the 24 questions against the partner you've already chosen. Catching a gap before the contract is signed is cheaper than catching it three months in.

Link to this answer: Can I use this if I've already picked a partner?
05 Why publish the routes where you're not the best fit?

Because a page that only ever concludes "hire us" isn't useful to anyone making a real decision. The honesty is the actual point of the wrong-choice section below.

Link to this answer: Why publish the routes where you're not the best fit?
Check the record, not the pitch

Seniority bait-and-switch and a black-box handover are two of the five ways an engagement fails. Our delivery record shows the actual commits and merge requests behind every project, not a highlight reel.

Source: https://customlabs.io/choosing-a-partner/

navigate select esc close