CustomLabs
Cost & ops

Inference Cost

Inference cost is what it costs to run a trained model on a request.

It is priced per token for hosted APIs and is distinct from training cost.

It scales with token volume, model choice, and context size.

Left unmodeled before shipping, it turns a good demo into an uneconomical product.

Model-agnostic routing, sending easy requests to a cheaper model, is a direct lever for controlling it.

← Back to the full glossary

Source: https://customlabs.io/glossary/inference-cost/

navigate select esc close