8 min read

How much does an AI agent cost to build and run?

AI agent cost splits into a one-time build and a recurring run bill. Here is what drives each, illustrative ranges with their assumptions stated, and how to keep cost per task visible.

Also available inالعربيةDeutschEspañolFrançais中文

AI agent cost has two parts: a one-time build cost driven by integrations, write actions, and evaluation work, and a recurring run cost driven by model calls per task, human review time, and maintenance. A narrow single-workflow agent is a modest project; an agent that writes into several business systems is a platform decision.

This guide explains the drivers behind each number, gives illustrative ranges with their assumptions stated, and shows how to measure cost per task before you commit. The ranges are planning aids, not quotes. Your workflow, your data, and your tolerance for mistakes move them more than any vendor rate card does.

Why AI agent cost is hard to estimate

A chatbot answers once. An agent plans, calls a tool, reads the result, and decides whether to continue. The number of model calls per task is not fixed; it depends on the input. Two support tickets that look alike can produce a three-step run and a twelve-step run, and the second one costs several times more to process.

Build cost is hard to pin down for a different reason. The demo is the cheap part. A prototype that calls a model and one API can be assembled in days. The expensive work sits at the edges: systems without a test environment, records that disagree with each other, approval screens, permission scoping, evaluation, monitoring, and the recovery path when a write half-succeeds. None of that shows up in a demo, and all of it decides whether the agent can stay switched on.

What teams often get wrong

The most common mistake is pricing the model instead of the workflow. Tokens are the most visible line item, so budgets anchor on them. In many workflows they are not the largest cost. The time a person spends reviewing drafts, the engineering hours spent on integrations, and the monthly upkeep when a model version is retired usually matter more.

Other recurring errors:

  • Treating the prototype as most of the build, then discovering that the production work is the larger share.
  • Leaving out review time, even though every approval step converts model output into paid human minutes.
  • Forgetting maintenance. Models are deprecated, provider APIs change versions, and every prompt change needs a regression run.
  • Running without a ceiling. An agent with no step limit or budget per run can loop on a failing tool and keep spending until someone notices.
  • Comparing proposals on day rate alone, when the real difference is whether evaluation, monitoring, and handover are included.

A practical way to estimate the build

Estimate the build from the workflow outward. Write down the systems the agent reads from, the systems it writes to, the decisions it makes, and the worst realistic mistake. That list, not the choice of model, is what a developer can price.

The build cost drivers

  • Systems touched: each integration needs authentication, error handling, and tests. A system with no API or no sandbox adds work or forces an explicit human step.
  • Read versus write: a read-only assistant is far cheaper to make safe than an agent that changes records, sends messages, or moves money.
  • Approval experience: a reviewer needs to see the proposed change, the evidence behind it, and a way to edit before approving.
  • Data preparation: retrieval over messy documents produces confident wrong answers. Cleaning the slice of data one workflow depends on is often required.
  • Evaluation set: real historical cases with known correct outcomes, plus adversarial inputs, so quality can be measured before launch and after every change.
  • Observability and limits: per-run logs, cost tracking, alerts, and ceilings.
  • Execution model: a short request-response agent can run in a web request; long runs need a queue, a worker, and checkpoints.

Illustrative build ranges

The ranges below assume a small senior team, existing systems with usable APIs, one production environment, and a client who can supply real historical examples for evaluation. They exclude software licenses, broad data cleanup beyond the workflow's own slice, and compliance work that needs legal review. They are orders of magnitude for a budgeting conversation, not prices.

  • Read-only assistant over your own documents or one system, with citations and an evaluation set: roughly $8,000 to $25,000.
  • Single-workflow agent that drafts actions into one system, with an approval step, an audit log, and monitoring: roughly $20,000 to $60,000.
  • Multi-system agent with several write integrations, durable execution, per-tool permissions, and a shared tool layer: often $60,000 to $150,000 or more, usually delivered in stages.

A fixed quote needs a scoped workflow, sample data, and a written definition of a correct outcome. Without those, any number is a guess dressed up as a price.

How to estimate the run cost per task

Run cost is easiest to reason about per task, then multiplied by volume. Use a simple model: cost per task equals model cost, plus tool and API fees, plus infrastructure, plus human review time, plus a share of monthly maintenance.

Here is a worked example with placeholder prices; check your provider's current price sheet before using real figures. Suppose a model costs $1 per million input tokens and $5 per million output tokens. A typical task makes eight model calls, each sending about 6,000 input tokens and receiving about 500 output tokens. That is 48,000 input tokens, or about $0.048, and 4,000 output tokens, or $0.02, so roughly $0.07 per task. At 20,000 tasks a month, model spend is around $1,360.

Now add review. If a person spends two minutes approving each draft, 20,000 tasks is about 667 hours of review a month. In this example the reviewer's time is a much larger line than the tokens. That is why approval design, covered in human-in-the-loop agents, is a cost decision as much as a safety decision.

Maintenance and change costs

Budget for recurring work even when nothing is broken. Model versions are retired and replacements behave differently. Connected systems change their APIs. New edge cases appear in real traffic and become new evaluation cases. Each prompt or model change needs a regression run before release. Teams usually cover this with a monthly support arrangement or a defined share of an internal engineer's time.

Implementation considerations that change the bill

Several engineering choices move run cost directly:

  • Step limits: cap the number of reasoning steps per run so a confused agent stops instead of looping.
  • Budgets per run: stop or escalate when token use for one task crosses a threshold.
  • Output shaping: truncate or summarize large tool results before they go back into the model's context.
  • Model routing: use a smaller model for classification and extraction steps, and reserve the larger model for planning or drafting.
  • Bounded retries: retry only on errors that can succeed on retry, with backoff. Every retry is another paid call.
  • Usage logging: record token usage per step so cost per task comes from data, not estimates.
  • Caching: reuse retrieved context and stable system instructions where the provider supports it.

Trade-offs

A cheaper model can raise total cost if it needs more steps, more retries, or more human correction. Measure the whole task, not the price per token.

Approval steps cost reviewer time but lower the expected cost of an incident. Keep them where a mistake is expensive and remove them where evidence shows the agent is reliable.

Building durable execution, a tool-call ledger, and evaluation up front raises the initial build. Adding them after an agent is in production usually costs more, because the prototype's shortcuts have spread through the code. The staged approach in migrating existing systems to agents keeps the first stage small while preserving that structure.

Lessons from ImadDhin work

The observations below come from the code of this portal's Agent workspace. They are implementation and code-level notes, not client outcomes or production cost figures.

  • Cost exposure is a product decision. Free chat is capped per visitor on the server, and research tools that call a paid external API return a payment-required response unless the visitor has the premium tier. The expensive capabilities are gated before they are called, not after the invoice arrives.
  • Counters need the right consistency. A usage cap is only as strong as the way its counter is updated. A read-then-write counter lets concurrent requests slip past the limit, which is a common finding when reviewing AI-built products. That may be tolerable for a small free allowance; anything tied to billing or credits should use an atomic increment or a transaction.
  • Runs are bounded in several places. The workflow runtime limits a run to a fixed number of graph steps, rejects stored checkpoints above a size budget, and truncates tool output before returning it to the model. Each limit stops a different kind of runaway.
  • Usage is recorded per step. Each reasoning step writes the model's usage metadata into the run's event log, which makes cost per task something you can compute from records.
  • Retries are spend. The research wrapper makes at most three attempts, honors the provider's retry-after header on rate limits, backs off on timeouts and server errors, and never retries bad-request or authentication failures.

Common mistakes to test for

Before launch, test the cost behavior as deliberately as the answers:

  • Send an unusually long or adversarial input and confirm token use per run stays within the ceiling.
  • Make a tool fail repeatedly and confirm the agent stops or escalates instead of looping.
  • Fire concurrent requests at any usage limit and confirm it holds where it must.
  • Simulate a provider outage and check what the fallback path costs and how it behaves.
  • Cancel a run mid-way and confirm spending stops.
  • Confirm you can report cost per task for each workflow, not just a monthly invoice total.

When a simpler solution is better

If the task always follows the same steps, a deterministic workflow with one model call for classification or drafting is cheaper to build, cheaper to run, and easier to test than an agent. If volume is low, a person with a good template may be the cheapest option of all.

An agent earns its cost when inputs vary enough that the path must be decided case by case, and when the volume is high enough that the build is amortized. If you are unsure which side of that line a workflow sits on, the AI Readiness Scan is a low-cost way to find out before committing to a build.

Get a cost model for your own workflow

A useful estimate starts from the workflow, the systems it touches, and the worst realistic mistake, then separates build cost from cost per task. If you want that model built around your actual process, see AI agent development or bring the workflow and its current volumes to a 30-minute call.

Frequently asked questions

What is the biggest driver of AI agent cost?

For the build, it is the number of systems the agent writes to and how safely those writes must be handled. For running costs, human review time and maintenance often outweigh model tokens, especially when every action needs approval.

Can I get a fixed price for an AI agent?

Yes, once the workflow is scoped, sample data is available, and a correct outcome is written down. Before that, a fixed price either carries a large contingency or leaves out evaluation, monitoring, and handover.

How do I estimate monthly running costs?

Estimate cost per task: model calls times tokens per call at your provider's prices, plus tool fees, infrastructure, reviewer minutes, and a share of maintenance. Multiply by expected monthly volume, then validate against logged usage during a shadow period.

Does a cheaper model always lower cost?

No. A cheaper model that needs more steps, retries, or human corrections can cost more per completed task. Compare models on the full task using your evaluation set.

What controls keep agent spending predictable?

Step limits per run, token budgets per task, truncated tool output, bounded retries on retryable errors only, usage logging per step, and alerts when cost per task drifts.

Build a cost model around your real workflow

Bring one workflow and its current volume; leave with the main build and run cost drivers.

Book a 30-minute call

Scoped builds with evaluation, approvals, audit trails, and cost ceilings included.

See AI agent development

Check whether a workflow needs an agent or a simpler automation.

Run the AI Readiness Scan

Keep reading