# CustomLabs > The senior-only applied-AI studio for teams burned by demos that don't survive production. We ship eval-tested, model-agnostic systems in weeks, and you own every line. > Production-grade applied AI. Eval-tested, model-agnostic, yours to own — shipped in weeks, not quarters. CustomLabs is an agentic AI consultancy specializing in production-grade applied AI. Model-agnostic, no vendor lock-in. This file is a curated map of the site for LLMs and agents. ## Services - [AI Integrations](https://customlabs.io/services/ai-integrations/): Ship AI features into the stack you already run: production-grade, eval-tested, and yours to own. - [Custom Development](https://customlabs.io/services/custom-development/): Bespoke agents and applications built to your workflow, production-grade, eval-tested, and yours to own. - [Strategy & Architecture](https://customlabs.io/services/strategy-architecture/): Honest, technically credible roadmaps for what to build, what to buy, and what to skip. - [Readiness & Diligence](https://customlabs.io/services/readiness-diligence/): A written, defensible read on whether the AI is really ready, and what the risks are. ## Diagnostics - [AI Readiness / Ship Audit](https://customlabs.io/diagnostic/ship-audit/): A fixed-scope, fixed-fee read on whether you’re actually ready to ship AI. - [AI Technical Diligence](https://customlabs.io/diagnostic/technical-diligence/): An investor-grade read on a target’s AI claims, before you commit capital. - [Diagnostic](https://customlabs.io/diagnostic/): The low-risk first step before a full engagement. ## Capabilities - [Conversational interfaces](https://customlabs.io/capabilities/): Domain-tuned chat and voice systems with proper retrieval, guardrails, and a paper trail you can audit. - [Agentic workflows](https://customlabs.io/capabilities/): Tool-using agents that run real operations — book, route, file, decide — with deterministic fallbacks and human-in-the-loop where it matters. - [Document & data extraction](https://customlabs.io/capabilities/): Structured output from invoices, contracts, forms, and long-form docs at production scale and cost. - [Evaluation & observability](https://customlabs.io/capabilities/): Eval suites, trace pipelines, and dashboards so quality stops being a vibes check and starts being a number. ## Tools - [AI Cost Calculator](https://customlabs.io/tools/cost-calculator/): What AI actually costs to run in production — not the sticker price. - [AI Readiness Scorecard](https://customlabs.io/tools/ai-readiness/): A fast, honest read on whether your data, infra, and process are ready to ship AI. - [Tools](https://customlabs.io/tools/): Free, self-serve tools for scoping AI work before you commit budget. ## Products - [CodeHerder](https://customlabs.io/products/codeherder/): Command your fleet of coding agents. (codeherder.com) - [CostMon](https://customlabs.io/products/costmon/): One trusted number for all your spend. — coming soon (costmon.com) - [FreeTier](https://customlabs.io/products/freetier/): Ship on a budget. (freetier.co) - [GreatAPIs](https://customlabs.io/products/greatapis/): A field guide to the programmable web. (greatapis.com) - [Beemy](https://customlabs.io/products/beemy/): The super-organized version of you. — coming soon (beemy.co) - [CustomHosted](https://customlabs.io/products/customhosted/): Hosting, run like a utility. (customhosted.com) - [Reserver](https://customlabs.io/products/reserver/): Reserved Instances, managed properly. — coming soon (reserver.io) - [txtfetch](https://customlabs.io/products/txtfetch/): Any document in. Clean text out. — coming soon (txtfetch.com) - [BizBinder](https://customlabs.io/products/bizbinder/): Your whole business, one binder. — coming soon (bizbinder.com) ## Case Studies - [The AI-Native Target That Wasn't: A Two-Week Diligence](https://customlabs.io/case-studies/ai-diligence-flagged-replatform-cost/): A growth-equity firm had a term sheet out on an 'AI-native' SaaS target. A two-week technical readiness assessment surfaced $1.4M of undisclosed re-platform cost and a vendor dependency that would triple at projected volume, and repriced the deal. - [From Notebook to Production: An Extraction Model You Could Trust](https://customlabs.io/case-studies/extraction-notebook-to-production/): A mid-market healthtech had a document-extraction model that dazzled in a notebook and hallucinated on real intake traffic. An eval harness and confidence gating took hallucinations from 12% to 3% and shipped it in 8 weeks. - [Cutting Inference Spend 40% Without Betting on One Vendor](https://customlabs.io/case-studies/model-agnostic-routing-cut-inference-spend/): A Series B fintech was scaling LLM spend faster than revenue on a single provider. A model-agnostic routing layer cut unit cost, killed vendor concentration risk, and met a data-residency requirement its incumbent couldn't. - [Retrieval Pipeline That Actually Cut Support Load](https://customlabs.io/case-studies/retrieval-pipeline-cut-support-load/): A mid-market SaaS support platform replaced a keyword search widget with a tenant-isolated retrieval pipeline, cutting escalations and first-response time without a headcount increase. ## Insights - [Prompt Injection Is a Data Problem: A Threat Model You Can Ship Against](https://customlabs.io/insights/prompt-injection-threat-model/): Prompt injection can't be filtered away: the model can't reliably tell instructions from data. Here's the actual threat model and the controls that hold up. (by CustomLabs Engineering, updated 2026-07-24) - [Why Your AI Agent Works in the Demo and Stalls in Production](https://customlabs.io/insights/ai-agents-stall-in-production/): An agent that nails the demo stalls in production because reliability compounds across steps. Here's the math, the real failure modes, and how to ship one anyway. (by CustomLabs Engineering, updated 2026-07-19) - [What AI Actually Costs in Production, by Workload](https://customlabs.io/insights/ai-cost-benchmark-by-workload/): A modeled, reproducible benchmark of cost-per-successful-outcome across four common AI workloads, with every token assumption, price, and overhead multiplier shown. (by CustomLabs Engineering, updated 2026-07-11) - [Model-Agnostic by Design](https://customlabs.io/insights/model-agnostic-by-design/): Models change under you every few months: price, quality, and capability. Here's why we never hardcode a single provider into a client's feature. (by CustomLabs Engineering, updated 2026-07-11) - [Your RAG Demo Lied to You](https://customlabs.io/insights/your-rag-demo-lied/): Retrieval that looks flawless on ten clean PDFs falls apart on a real corpus. Here's why, and what evaluating retrieval quality actually requires. (by CustomLabs Engineering, updated 2026-07-11) - [What an AI Feature Actually Costs in Production](https://customlabs.io/insights/what-ai-actually-costs/): Token costs that look trivial in a demo compound fast at scale. Here's how to make cost a first-class metric instead of a surprise on the invoice. (by CustomLabs Engineering, updated 2026-07-11) - [Evals Before You Ship: Why AI Features Need Tests Too](https://customlabs.io/insights/evals-before-you-ship/): Shipping an AI feature without an eval suite in CI means every prompt tweak is a guess. Here's what an eval suite actually needs to cover. (by CustomLabs Engineering, updated 2026-07-11) - [Build, Buy, or Skip: A Framework for AI Decisions](https://customlabs.io/insights/build-buy-or-skip/): A technically honest framework for deciding whether an AI initiative should be built in-house, bought off the shelf, or skipped entirely this year. (by CustomLabs Engineering, updated 2026-07-11) - [From Notebook to Production: What Actually Breaks](https://customlabs.io/insights/notebook-to-production/): The gap between a working AI demo and a production feature is auth, latency, cost and fallbacks. Here's how we close it without a rewrite. (by CustomLabs Engineering, updated 2026-07-11) ## Comparisons - [Agents vs Pipelines: When Autonomy Is Worth the Reliability Cost](https://customlabs.io/compare/agents-vs-pipelines/): Agentic autonomy vs Deterministic pipelines — Constrain to a pipeline until evals prove the workflow genuinely needs an agent's judgment. Default to less agent, not more. - [Open-Weight Models vs Frontier APIs: The Real Cost/Control Tradeoff](https://customlabs.io/compare/open-weight-vs-frontier-api/): Self-hosted open-weight models vs Frontier API models — Default to a frontier API until volume, compliance, or control needs force the switch. Self-hosting is an operations commitment first and a cost decision second. - [RAG vs Fine-Tuning: Which One Actually Solves Your Problem](https://customlabs.io/compare/rag-vs-fine-tuning/): Retrieval-augmented generation (RAG) vs Fine-tuning — Default to retrieval. Fine-tune only when you need a fixed output form or a latency budget retrieval can't hit, not just 'better answers.' - [Vector Database vs pgvector: Do You Actually Need a Dedicated Store](https://customlabs.io/compare/vector-database-vs-pgvector/): Dedicated vector database vs Postgres + pgvector — Start on pgvector. Migrate to a dedicated vector database only when you hit a measured scale, latency, or feature ceiling it can't clear, not preemptively. ## Glossary - [Retrieval-Augmented Generation (RAG)](https://customlabs.io/glossary/retrieval-augmented-generation/): Retrieval-Augmented Generation (RAG) is an architecture that retrieves relevant documents or passages at query time and feeds them into an LLM's context so it answers from your data instead of its training memory. It trades a training-time knowledge problem for a retrieval-quality problem — chunking, embeddings, and ranking now decide whether the answer is grounded or a plausible-sounding guess. Most "AI doesn't know our data" complaints are RAG pipeline defects, not model limitations. - [Embeddings](https://customlabs.io/glossary/embeddings/): Embeddings are numeric vector representations of text (or images) positioned so semantically similar content sits close together in vector space. They're the substrate under vector search and RAG: an embedding model converts a query and a corpus into vectors, then similarity search finds the nearest matches. Embedding model choice and dimensionality directly bound retrieval quality — a mismatched or stale embedding model is a common silent RAG failure. - [Context Window](https://customlabs.io/glossary/context-window/): The context window is the maximum amount of text, measured in tokens, a model can consider at once — spanning the prompt, retrieved documents, conversation history, and its own output. A larger window doesn't mean it should be filled: cost scales with tokens processed, and good retrieval still beats brute-force stuffing every relevant document in. Sizing the context window is a cost and latency decision as much as a capability one. - [Token](https://customlabs.io/glossary/token/): A token is the basic unit an LLM reads and writes — roughly three-quarters of a word in English, though the exact split varies by model and tokenizer. Usage-based pricing, context window limits, and latency are all denominated in tokens, so token counts are the unit every cost or capacity conversation in this space eventually collapses to. - [Model-Agnostic Architecture](https://customlabs.io/glossary/model-agnostic-architecture/): A model-agnostic architecture puts an abstraction layer between your application and any single LLM provider, so you can swap or route between models — Anthropic, OpenAI, open-weight, self-hosted — without rewriting the application. It protects against price changes, deprecations, and capability shifts from any one vendor, and lets different workloads route to whichever model is cheapest or best suited. It's a deliberate design decision, not a default — most teams retrofit it after their first vendor lock-in scare. - [Fine-Tuning (vs RAG)](https://customlabs.io/glossary/fine-tuning/): Fine-tuning further trains a model's weights on your own examples so it changes behavior — tone, format, a narrow skill — baked into the model itself, rather than supplying facts at query time the way RAG does. The two solve different problems: fine-tuning teaches a model how to respond, RAG gives it what to respond with. Reaching for fine-tuning to fix a knowledge or freshness problem that RAG (or better prompting) would solve cheaper and faster is a common build-vs-buy mistake. - [Vector Search (Semantic Search)](https://customlabs.io/glossary/vector-search/): Vector search finds the nearest matches to a query by comparing embeddings in vector space, rather than matching exact keywords the way traditional full-text search does. It's what lets a RAG pipeline surface a passage about "cancel my plan" when the user typed "how do I stop being billed." In production it's usually paired with metadata filters and a reranker, since raw nearest-neighbor results alone are rarely precise enough to hand straight to the model. - [Chunking](https://customlabs.io/glossary/chunking/): Chunking is splitting source documents into smaller passages before embedding them, so retrieval returns focused, relevant text rather than an entire document. Chunk size and overlap are load-bearing decisions — chunks too large dilute relevance and blow the context budget, chunks too small lose the surrounding context a passage needs to make sense. Most RAG accuracy problems trace back to chunking, not the model. - [Agent (Agentic AI)](https://customlabs.io/glossary/agent/): An agent is an LLM given a loop, memory, and a set of tools it can call, so it can plan multi-step work and take actions rather than just returning a single answer. "Agentic" describes a pattern along a spectrum — from a single tool call to a fully autonomous multi-step loop — not a strict on/off feature. The engineering that matters is less the model and more the guardrails, observability, and evals around the loop, since an agent that can act is also an agent that can act wrongly. - [Tool Calling](https://customlabs.io/glossary/tool-calling/): Tool calling — also called function calling — is the mechanism that lets a model request a structured action, like calling an API or running a query, instead of only generating text, with the calling application executing the action and returning the result. It's the primitive underneath every agent: an agent is essentially a model given a set of callable tools and a loop to call them in. Reliability here comes down to tight tool schemas and validating what the model actually asks for before executing it. - [Model Context Protocol (MCP)](https://customlabs.io/glossary/model-context-protocol/): The Model Context Protocol (MCP) is an open standard for connecting LLM applications to external tools, data sources, and other agents through a common interface, rather than every integration being a bespoke one-off. It's the plumbing that lets a single tool or data connector be written once and reused across different agents and applications. Adopting it early is part of building model-agnostic, vendor-neutral agent systems instead of tools wired to one specific assistant — though MCP standardizes the wire format for that connection, not the catalog, contract, or permission decisions behind it. - [Context Engineering](https://customlabs.io/glossary/context-engineering/): Context engineering is deciding what occupies a model's context window at every step — tool definitions, retrieved content, conversation history, and the task itself — and in what order, not just what to write in a system prompt. It treats the window as a fixed, competed-for resource: every token spent on one slice is a token unavailable to another, so the allocation has to be a deliberate decision rather than whatever's left over once everything else is assembled. Prompt engineering is one input to it, not the whole discipline. - [Tool Contract](https://customlabs.io/glossary/tool-contract/): A tool contract is the schema a tool exposes to a model — its parameter names, types, required fields, and whatever enums or patterns constrain them. A loose contract lets a model fill an ambiguous field with a plausible-looking guess instead of a real value; a tight one, validated server-side, is what actually stops a hallucinated argument from reaching execution. It's the one piece of documentation a model reads before every call, so it has to carry everything a new hire would otherwise ask about in person. - [Idempotency](https://customlabs.io/glossary/idempotency/): An idempotent operation produces the same result no matter how many times it runs with the same input — calling it twice does nothing a single call didn't already do. For a tool that writes, sends, or deletes, idempotency (usually enforced with a unique key passed on every retry of the same logical call) is what makes a retry safe: without it, a timeout followed by an automatic retry can execute the same mutation twice. - [Structured Output](https://customlabs.io/glossary/structured-output/): Structured output constrains a model's response to a defined schema — JSON, an enum, a typed object — instead of free-form prose, so the calling code can parse it reliably without brittle regex or string matching. It's foundational to tool calling and agent loops, where a wrongly-shaped response breaks the next step in the chain. Enforcing and validating the schema, not just asking nicely for JSON, is what makes it dependable in production. - [Guardrails](https://customlabs.io/glossary/guardrails/): Guardrails are the checks placed around a model's input and output — content filters, schema validation, permission scoping, human-approval gates — that keep an LLM or agent inside acceptable bounds in production. They matter most for agents with real tool access, where an ungrounded or manipulated response can translate directly into an unwanted action, not just a bad chat reply. Guardrails are a design layer added deliberately, not a property models come with by default. - [Eval Suite (Evals)](https://customlabs.io/glossary/eval-suite/): An eval suite is a repeatable, versioned set of test cases and scoring criteria used to measure whether an LLM system's outputs are actually good — accurate, on-format, safe — before and after every change. Unlike traditional unit tests, evals often score graded or probabilistic quality rather than strict pass/fail, which is why most teams pair automated scoring with LLM-as-judge or periodic human review. Shipping an LLM feature without an eval suite means every prompt or model change is a guess about whether quality went up or down. - [LLM-as-Judge](https://customlabs.io/glossary/llm-as-judge/): LLM-as-judge is an evaluation technique that uses a — typically stronger or differently-configured — LLM to score another model's outputs against a rubric, at a scale human review can't match. It's useful for grading subjective qualities like tone, relevance, or faithfulness to a source document, but it inherits the judging model's own biases and blind spots, so it's normally calibrated against a smaller human-labeled sample rather than trusted blind. Treat it as one signal in an eval suite, not the whole suite. - [Observability](https://customlabs.io/glossary/observability/): Observability, in an LLM context, means capturing traces of every prompt, retrieval, tool call, and response so a team can debug why a specific output happened and track quality, latency, and cost over time — not just whether the request returned a 200. Without it, a regression after a prompt or model change surfaces as a vague complaint that "the AI got worse," with no way to pinpoint which step changed. It's the operational counterpart to an eval suite: evals catch regressions before ship, observability catches them after. - [Hallucination](https://customlabs.io/glossary/hallucination/): A hallucination is a confident, fluent output that is factually wrong or unsupported by any real source — the model completing a plausible-sounding answer rather than admitting it doesn't know. It's not a bug that gets patched out; it's a structural property of how these models generate text, which is why grounding (RAG) and evals, not a better prompt, are the actual mitigations. A demo that never surfaces a hallucination almost always means the test set was too easy, not that the system is hallucination-free. - [Prompt Injection](https://customlabs.io/glossary/prompt-injection/): Prompt injection is an attack where untrusted input — a document, a webpage, a user message — contains instructions crafted to override a model's original system prompt or task, hijacking its behavior. It matters most once a model can retrieve untrusted content or call tools, since a successful injection can turn a summarization task into an unauthorized action. Defending against it takes input/output evals and guardrails, not just a stricter system prompt, because the system prompt is exactly what's being attacked. - [Inference Cost](https://customlabs.io/glossary/inference-cost/): Inference cost is what it costs to run a trained model on a request — usually priced per token for hosted APIs — as distinct from training cost, a one-time or periodic expense most teams building on foundation models never pay directly. It scales with token volume, model choice, and context size, and is the line item that turns a good demo into an uneconomical product if it isn't modeled before shipping. Model-agnostic routing — sending easy requests to a cheaper model and hard ones to a stronger one — is one of the most direct levers for controlling it. - [Notebook-to-Production](https://customlabs.io/glossary/notebook-to-production/): Notebook-to-production describes the gap between a working data-science notebook or prototype and a system that runs reliably, observably, and cheaply in production — error handling, retries, monitoring, cost controls, and deployment infrastructure a notebook doesn't need. It's usually a bigger lift than the original prototype, which is why it's often underestimated in project timelines. Treating it as a distinct phase, with its own scope and budget, is what separates a demo that ships from one that stalls. - [Data Processing Agreement (DPA)](https://customlabs.io/glossary/data-processing-agreement/): A Data Processing Agreement is the contract that names a vendor as a processor (or subprocessor) of personal data on your behalf, and sets the terms for how they handle it — retention, subprocessors, breach notification, deletion. Routing personal data to a third-party model API almost always brings that provider into this category, whether or not anyone thought of it as "processing personal data" at the time. Privacy and legal reviewers check that the specific model provider and product you're calling is actually named in the DPA, not covered by a generic reference to "AI features." - [Data Residency](https://customlabs.io/glossary/data-residency/): Data residency is where data is physically processed and stored, as distinct from where your company or your users are located. A model API call can route through infrastructure in a different region than the rest of your stack, so the residency question has to be answered per provider, not assumed from your primary hosting region. It matters most once a privacy notice, a customer contract, or a framework like GDPR makes a specific residency commitment — at that point residency stops being an infrastructure detail and becomes something reviewers check against what you actually promised. - [Red Teaming](https://customlabs.io/glossary/red-teaming/): Red teaming is deliberately attacking your own AI system — planting adversarial documents, crafting injection payloads, probing for actions a user shouldn't be able to trigger — to find what breaks before an attacker, or a curious user, finds it first. Unlike a golden-set eval, which checks whether the system gets normal cases right, a red-team suite checks whether it fails safely on cases designed to make it fail. Run once, it is a point-in-time audit; run in CI on every change, it is the same regression protection a golden-set gate gives ordinary quality, applied to security. ## Failure modes - [Why does our AI cite a policy we deleted six months ago?](https://customlabs.io/failure-modes/stale-index-serves-deleted-content/): Your retrieval index was built once at ingest and never told the source changed. When a document is edited or deleted, nothing re-embeds the new version or tombstones the old chunk, so the stale vector keeps scoring well and keeps getting served — confidently, and with no signal to the reader that it is out of date. - [Why can't retrieval find an answer that is definitely in the document?](https://customlabs.io/failure-modes/chunk-boundary-splits-the-answer/): The answer exists in the source, but a fixed-size chunker cut it in half at ingest time — a table row split from its header, a procedure split from its trigger condition. Each half scores weakly on its own, the ranker drops both, and retrieval reports nothing when the document plainly contains the answer. - [Why does retrieval return confidently wrong but similar documents?](https://customlabs.io/failure-modes/similarity-is-not-relevance/): Cosine similarity rewards topical resemblance, not correctness — it can rank a document about the wrong product, the wrong date, or the negated version of a claim above the one that actually answers the query, because embeddings represent "about the same thing" far more reliably than they represent identifiers, negation, or numbers. - [Why did our agent run for 40 minutes and produce nothing?](https://customlabs.io/failure-modes/unbounded-agent-loop/): The agent has no step budget, no token budget, and no way to recognize it is stuck — a failing tool call stays in its context and keeps looking like a reasonable next thing to try, so it keeps trying variations of the same failed approach until something external (a timeout, a bill, a human) stops it. - [Why is our agent calling tools with IDs that don't exist?](https://customlabs.io/failure-modes/tool-argument-hallucination/): Loose tool schemas — free-form string IDs, everything optional — give the model room to fill a gap with something plausible-looking instead of something real, and with no server-side validation catching the mismatch before execution, a confidently invented ID reaches a system that expects a real one. - [Why is our agent confidently wrong right after a tool call failed?](https://customlabs.io/failure-modes/silent-tool-failure/): The tool returned HTTP 200 with an error message in the body, or an empty result set, and the agent read the absence of data as evidence rather than as a failure — because nothing in the response forced a distinction between "nothing matched" and "something broke." - [Why does our agent forget its instructions halfway through a long task?](https://customlabs.io/failure-modes/context-overflow-drops-the-task/): As the conversation grows, a naive truncation strategy drops the oldest messages to stay under the context window — and the oldest messages are exactly where the system prompt and the original task state usually live, so the agent keeps running with no memory of what it was actually supposed to do. - [Why did quality drop after a prompt tweak nobody thought was risky?](https://customlabs.io/failure-modes/vibes-based-prompt-regression/): Without a labelled eval set, the change was graded against whatever two or three examples the author happened to have open — which is not a test, it's an anecdote. A regression anywhere outside that narrow, unrepresentative sample ships straight to production undetected. - [Why does our LLM-as-judge say everything passes?](https://customlabs.io/failure-modes/judge-prefers-its-own-output/): A judge from the same model family as the generator tends to rate that family's output favorably — self-preference bias — and a single vague rubric ('is this good?') collapses almost everything to a passing score, so the eval suite stops being able to tell a real regression from noise. - [Why is our inference bill three times the estimate?](https://customlabs.io/failure-modes/retry-amplified-spend/): The estimate priced the happy path — one clean call per outcome. Production reality includes retries on malformed or rate-limited calls, fallbacks to a larger model when the first attempt fails, and agent loops that make several calls per completed task, and every one of those multiplies calls per successful outcome without multiplying the original per-token estimate. - [Why aren't we getting prompt-caching discounts?](https://customlabs.io/failure-modes/prompt-cache-never-hits/): A dynamic prefix — a timestamp, a per-user greeting, a reordered tool list, retrieved chunks placed before the static instructions — changes the start of the prompt on every call, and prompt caching only pays off when the shared prefix is byte-identical across requests. One volatile token near the front is enough to bust the whole cache. - [Can a document in our own knowledge base hijack our agent?](https://customlabs.io/failure-modes/injection-via-retrieved-content/): Yes — retrieved content arrives on the same channel as instructions, so a document, ticket, or webpage crafted (or compromised) to contain commands can have the model execute them with its real tool permissions, and the system has no built-in way to tell 'instruction from us' apart from 'text we retrieved.' ## Patterns - [Intent router to specialists](https://customlabs.io/patterns/intent-router-to-specialists/): A cheap, fast classifier reads the incoming request first and routes it to one of several narrow, single-purpose agents — each holding only the tools, context, and instructions its job needs — instead of a single god-agent carrying every tool definition and every rule for every possible request. - [Bounded agent loop](https://customlabs.io/patterns/bounded-agent-loop/): An agent loop runs under an explicit budget — a maximum step count, a token ceiling, and a wall-clock limit — plus a termination contract that forces every run to end in one of a small number of named states: success, failure, or escalation. - [Structure-aware chunking](https://customlabs.io/patterns/structure-aware-chunking/): Chunk boundaries follow the document's own structure — headings, table rows, list items, section boundaries — instead of a fixed token count, and each chunk carries its parent heading or identifying context in its own body. - [Retrieve-then-rerank](https://customlabs.io/patterns/retrieve-then-rerank/): A cheap, high-recall first pass — vector search, optionally fused with keyword search — pulls a wide candidate set of 50 to 100 documents likely to contain the right answer somewhere. - [Change-data-capture ingest](https://customlabs.io/patterns/change-data-capture-ingest/): The ingest pipeline subscribes to the actual change events of the source system — a webhook, a CMS publish hook, a database trigger — instead of re-crawling on a fixed schedule. - [Typed tool contract](https://customlabs.io/patterns/typed-tool-contract/): Every tool argument is defined by a strict JSON schema — enums for known value sets, validated patterns for IDs, required fields wherever the tool genuinely needs them — with no free-text catch-all surface. - [Human checkpoint before irreversible actions](https://customlabs.io/patterns/human-checkpoint-before-irreversible/): Every tool the agent can call is scoped to the narrowest permission the task genuinely needs, and any action that can't be cleanly undone — a refund, a delete, an external message — requires an explicit human confirmation before it executes, not just a plausible-looking model decision. - [Golden-set gate in CI](https://customlabs.io/patterns/golden-set-gate-in-ci/): A fixed, human-labelled set of real cases — each with a specific, checkable expected property, not a vibe — runs automatically in CI on every prompt or model change. - [Trace-first observability](https://customlabs.io/patterns/trace-first-observability/): One trace ID follows a single request across every hop it takes — retrieval, every model call, every tool call — logged with enough detail to reconstruct exactly what happened after the fact. - [Model cascade](https://customlabs.io/patterns/model-cascade/): A cheap, fast model attempts every request first. - [Stable-prefix prompt caching](https://customlabs.io/patterns/stable-prefix-prompt-caching/): The prompt is ordered with everything invariant across calls first — system instructions, tool definitions in a fixed serialization order, few-shot examples — and everything that changes per request — retrieved chunks, user input, timestamps — placed last. ## Handbook - [01 Decide](https://customlabs.io/handbook/decide/): Should we build this with AI at all — build, buy, or skip? - [02 Design](https://customlabs.io/handbook/design/): What shape is the system: pipeline, agent, retrieval, or none of the above? - [03 Build](https://customlabs.io/handbook/build/): How do we get from a working notebook to a deployable service? - [04 Evaluate](https://customlabs.io/handbook/evaluate/): How do we know it works, and how do we keep knowing after every change? - [05 Operate](https://customlabs.io/handbook/operate/): What breaks in production, and how do we see it before the user does? - [06 Cost](https://customlabs.io/handbook/cost/): What will this actually cost to run, and where does the spend hide? - [The Applied AI Handbook](https://customlabs.io/handbook/): The six-stage map of the studio's full knowledge base. ## Security review - [AI Security Review](https://customlabs.io/security-review/): The questions InfoSec, Privacy & Legal, Risk & Compliance, and Procurement ask before an AI feature ships, and the evidence that answers them. - InfoSec — AppSec & Security Engineering: What can be reached and what happens when someone abuses it — the actual blast radius of a compromised prompt, key, or session. - Privacy & Legal: Lawful basis for processing personal data, data-subject rights, and who is contractually on the hook when a processor mishandles it. - Risk & Model Governance: Auditability and accountability — whether anyone can reconstruct why the system did what it did, and who signs off that it was allowed to. - Procurement — Vendor Risk: The supplier chain behind this feature, and whether you can leave it without losing your data or your work. ## Agentic delivery - [The Agentic Delivery Playbook](https://customlabs.io/agentic-delivery/): The operating model for running real software delivery with a fleet of coding agents, first-hand from the studio's own delivery system. - Task definition: What acceptance criteria a task carries before an agent starts, and whether 'done' has a checkable definition - Isolation: Whether two agents can ever touch the same mutable state at the same time - Stage gates: The sequence a task moves through, and what a fresh context inherits at each hand-off - Review throughput: How many tasks can be in flight without the review queue growing unboundedly - Durable memory: Whether a decision made once has to be re-litigated by every subsequent session - Spend & blast radius: The ceiling on any single loop's steps, tokens, wall-clock, and cost, and the gate every landing path must pass through ## Evals - [The Eval Stack](https://customlabs.io/evals/): Six eval layers, 24 named checks, five ways an LLM judge lies, and the numbers worth putting on a dashboard. - Contract checks: Malformed output — broken JSON, a missing field, a tool call with the wrong argument shape. - Golden-set task evals: A prompt, model, or retrieval change that quietly makes real cases worse. - Retrieval evals: The retriever handing the generator the wrong context, or none at all. - Trajectory evals: An agent that loops, calls the wrong tool, or never terminates. - Online signals: Everything the offline suite never saw, because this is real traffic. - Adversarial & safety: Prompt injection, jailbreaks, PII leakage, and tool misuse — under deliberate attack. ## Tool design - [The Agent Tool Interface](https://customlabs.io/tool-design/): Six interface surfaces, 24 named design rules, and five named failures — how to design the tools and context an agent actually works through. - Tool catalog & naming: Which tools exist in a session, and what they are named. - Tool contracts (the schema): The shape of a single tool's inputs. - Tool responses (what comes back): What the model actually reads on the way back from a call. - Context assembly & budget: What actually occupies the token budget at each step: tool defs, retrieved content, history, the task. - Permissions & blast radius: What a tool call is actually allowed to do once it executes. - Transport & operations: How a tool call crosses a process or team boundary, and whether anyone can reconstruct what happened afterward. ## Delivery record - [The Delivery Record](https://customlabs.io/delivery-record/): Six real defects the review and verify stages caught across 126 tasks merged in 31 days — what shipped, what caught it, and the build-time check that now fails if it comes back. - Nine product pages published a "what it's built with" stack that was inferred, not confirmed. Caught by review stage. - 26 identical, blank Open Graph social cards shipped to production. Caught by human inspection after deploy. - The light theme drew every focus ring in the lime accent — 1.08:1 on paper, so no visible indicator at all. Caught by verify stage, after merge. - Filter chips left five empty category headings behind, and two strings shipped literal markdown backticks. Caught by review stage. - The header search trigger navigated to /search/ and opened the command palette on top of it. Caught by browser check after merge. - Nine product pages rendered nine content sections with no real heading — one h1, one h2, per page. Caught by review stage. ## Topics - [Topics](https://customlabs.io/topics/): Insights and case studies grouped by topic — retrieval, cost, evals, production, and strategy. ## Library - [The Library](https://customlabs.io/library/): A single, filterable index of everything CustomLabs has published — the best crawl entry point on the site. ## Process - [Our process](https://customlabs.io/process/): 01 Discovery → 02 Brief & estimate → 03 Build & ship → 04 Handover ## Pricing - [Pricing & engagement](https://customlabs.io/pricing/): the three-stage engagement path — fixed-fee diagnostic, value-priced phased build, monthly evals & roadmap retainer ## Company - [About](https://customlabs.io/about/): The studio's background and operating principles - [FAQ](https://customlabs.io/faq/): common questions on cost, ownership, data - [Contact](https://customlabs.io/contact/): start a brief