All terms

Enterprise AI

LLM pricing

Also known as: token pricing, model pricing

LLM pricing is how providers charge for access to large language models, most commonly per token of input and output, with separate rates for each model. Subscription tiers, per-seat plans and discounts for cached or batched requests are also common. Understanding it lets teams estimate the running cost of an AI feature before they ship it.

What it is

Token pricing splits usage into what you send the model, the input or prompt, and what it returns, the output or completion. Output tokens usually cost more than input tokens, and more capable models cost more than smaller or distilled ones. Alongside API pricing, most vendors sell consumer and business subscriptions that bundle usage into a monthly fee.

Why it matters

Cost per interaction decides which AI features are viable at scale, so pricing shapes product design as much as model quality does. Long prompts, large retrieved contexts and verbose outputs can multiply spend quickly once traffic grows. For marketing teams, it also governs how much automated content research, classification or prompt monitoring you can afford to run.

How it works

Teams estimate token counts for a typical request, multiply by expected volume, then test cheaper models on the same task to see where quality drops. Common levers include trimming system prompts, reusing cached context, capping output length, batching non-urgent jobs and routing simple requests to a smaller model. Usage dashboards and per-feature tagging keep spend attributable.

When it applies

It applies whenever you build on a model API, buy assistant seats for a team, or budget for recurring AI workloads such as content generation, support triage or visibility tracking.

Examples

  • A support team compares the cost per resolved ticket on a large model against a smaller one before rolling out an AI first responder
  • A marketing team runs weekly brand visibility prompts in batch mode overnight to reduce API spend
  • A product team caches a long system prompt so repeated calls are charged at the cheaper cached input rate

How it is measured

  • Cost per request and cost per completed task
  • Average input and output tokens per interaction
  • Monthly model spend against budget, split by feature
  • Spend per resolved outcome, such as a booked demo or closed ticket

Related terms in Enterprise AI

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.