The Cost of Running an AI SaaS: Infrastructure Line by Line

7 minUpdated:
The Cost of Running an AI SaaS: Infrastructure Line by Line

Running an AI SaaS costs a mix of fixed and variable lines: model inference, embeddings and vector storage scale with usage, while hosting, database, observability, email and tooling are mostly fixed. Estimate each line from cost per user action multiplied by expected volume, then add a safety margin.

What are the main cost lines of an AI SaaS?

An AI product has the usual SaaS bill plus a variable layer that grows with every request. The trick is to separate them, because they behave differently as you grow.

The table below lists the lines most teams end up paying for. Your stack may skip some, for example if you self-host models or use a single platform that bundles database and auth.

Cost lineFixed or variableWhat drives itHow to estimate
LLM inference (API)VariableInput and output tokens per request, model choiceLog tokens per action in beta, multiply by provider price and volume
Self-hosted GPU inferenceMostly fixedGPU hours, utilization, redundancyInstances needed for peak throughput times hours per month
EmbeddingsVariableDocuments ingested, re-indexing frequencyTotal tokens embedded per month times price
Vector databaseFixed then steppedNumber of vectors, dimensions, replicasVectors times dimension size, then the tier or node size that fits
App hosting and databaseMostly fixedTraffic, background jobs, storageStart from one small instance plus managed or self-hosted Postgres
File and object storageVariableUploads, generated files, backupsGB stored plus egress per month
Observability and LLM tracingFixed or per eventLog volume, traces, retentionEvents per request times requests, or self-host
Email, SMS, notificationsVariableMessages sentMessages per user per month times users
Third-party APIsVariableSearch, OCR, speech, scraping callsCalls per action times price
PaymentsVariableRevenue processedProcessor percentage plus fixed fee per transaction

How do you estimate model inference cost?

Inference is usually the largest variable line, and the one most founders guess wrong. Do not estimate it from a single test prompt. Real requests include system prompts, retrieved context, tool results, retries and conversation history.

Instrument every model call with the model name, input tokens, output tokens and the user action that triggered it. After a week of beta traffic you can compute average and 90th percentile cost per action.

Then build the monthly estimate: actions per active user per month, multiplied by cost per action, multiplied by active users. Run it for light, typical and heavy users separately, because averages hide the accounts that break margins.

API models have no fixed cost and scale to zero, which suits early products with uneven traffic. You pay a premium per token for not running infrastructure.

Self-hosting an open-weight model with a server such as vLLM or llama.cpp turns inference into a mostly fixed GPU bill. It pays off only when utilization is high and steady, and it adds operations work: scaling, monitoring, updates and on-call.

Many teams mix both. They self-host embeddings or a small classification model where volume is high, and call an API for the complex reasoning steps.

What hidden costs do teams forget?

  • Agent loops: an agent that calls tools repeatedly can use many times the tokens of a single answer.
  • Retries and fallbacks when a provider times out or returns malformed output.
  • Re-embedding the whole corpus after changing chunking or embedding model.
  • Evaluation runs: every regression test suite that calls a model costs real money.
  • Free tier and trial usage, which consumes the same resources as paying users.
  • Staging and development environments that call production-grade models.
  • Log retention for prompts and completions, which grows faster than typical app logs.
  • Backups and egress, especially for products that store user documents.

How do you build a cost model step by step?

  • List every user action that triggers a paid resource: chat message, document upload, report generation, agent run.
  • For each action, record the resources it uses: model calls, embedding tokens, storage, third-party API calls.
  • Measure the unit cost of each resource from provider pricing pages or your own invoices.
  • Estimate actions per user per month for each plan and segment.
  • Add fixed lines: hosting, database, monitoring, domains, email service, developer tools.
  • Compute cost per active user and cost per plan, then compare with plan revenue.
  • Add a buffer for retries, spikes and growth, and revisit the model monthly with real invoices.

Where can you cut costs without hurting the product?

LeverAffectsEffortWatch out for
Route easy requests to a smaller modelInferenceMediumQuality drop on edge cases; evaluate first
Prompt and response cachingInferenceLow to mediumStale answers when source data changes
Trim context and historyInferenceLowLosing information the model needed
Batch non-urgent jobsInference, third-party APIsMediumSlower results for users
Self-host embeddingsEmbeddingsMediumExtra infrastructure to maintain
Tighter log retentionObservability, storageLowLess data for debugging and evals

Common mistakes in AI SaaS cost planning

  • Estimating from list prices without real token logs.
  • Using averages only and ignoring the heaviest ten percent of accounts.
  • Forgetting to budget for evaluations and internal testing.
  • Ignoring egress and storage for user-uploaded files.
  • Assuming model prices only go down; plan for switching to a pricier model when quality demands it.
  • Not setting per-account spending limits, so one abusive account becomes a large bill.

How does open source change the cost picture?

Open-source components remove licence fees for things like the vector store, observability, workflow engine and chat UI, but they add hosting and maintenance time. Count engineering hours as a cost line, even if nobody invoices you for them.

A self-hosted stack on one or two servers is often cheapest at small scale and becomes a trade-off only when you need high availability and a team on call. Revisit that decision at each growth stage instead of locking it in on day one.

How do costs change as you grow?

Cost lines do not scale together. Early on, fixed lines such as hosting, monitoring and developer tools dominate, because there are few users to spread them over. As usage grows, variable lines take over and inference usually becomes the largest item.

At larger scale, stepped costs appear. A vector database outgrows its node, a single app server needs replicas, or a provider rate limit forces a higher tier. Map the thresholds where each line jumps so they do not surprise you.

StageDominant costsWhat to watch
PrototypeYour time, small hosting, API experimentsTest calls in development leaking into production keys
BetaInference, free usage, basic monitoringCost per action and the heaviest users
Early revenueInference, database, support toolingGross margin per plan
ScalingInference, vector storage, redundancy, staff on callStepped infrastructure costs and provider limits

Which costs are people, not infrastructure?

Infrastructure invoices are only part of running an AI SaaS. Someone has to answer support, review failed runs, update prompts when models change and keep evaluations current. These tasks grow with the customer base.

Budget time for model migrations. When a provider retires a model version or a better option appears, prompts, tests and sometimes parsing code need adjusting. Treat this as a recurring maintenance line rather than an exception.

Include compliance work where it applies: data processing agreements, security questionnaires from customers and privacy reviews. For B2B products these appear early and consume founder time.

Tag every paid call with the account, plan and feature that triggered it. With that tagging in place, one query answers the questions that matter: which feature is most expensive, which accounts are unprofitable and whether a prompt change raised costs.

Reconcile your internal estimates with provider invoices each month. Gaps usually point to untracked calls, such as background jobs, retries or test scripts using production keys.

Set alerts on daily spend per provider and per account. A sudden jump is often a bug, such as an agent stuck in a loop, and catching it within hours rather than at month end saves real money.

  • Support hours per hundred active accounts.
  • Prompt and evaluation maintenance per model change.
  • Security reviews and customer questionnaires.
  • Incident response and on-call rotation.

Frequently asked questions

What is usually the biggest cost for an AI SaaS?
For most products built on API models, inference is the largest variable cost, followed by hosting and third-party APIs. Products that store large document collections can also see storage and vector database costs grow quickly. Measure your own traffic rather than relying on general rules.
How do I estimate costs before I have users?
Run your product internally with realistic tasks and log tokens and API calls per action. Build three usage scenarios, light, typical and heavy, and multiply by expected users. Treat the result as a range and replace it with real data once beta users arrive.
Is self-hosting an LLM cheaper than using an API?
Only when utilization is high and steady enough to keep GPUs busy. At low or uneven volume, APIs are usually cheaper because they scale to zero. Self-hosting also adds operations work, which should be counted as a real cost.
How do I stop one customer from causing a huge bill?
Set per-account rate limits and monthly spending caps, alert on unusual usage, and cap agent loop iterations and context size. For usage plans, let customers set their own budget limits. These controls protect margin and prevent surprise invoices on both sides.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides