AI
AI Agent Cost Optimization: Budget Per Task, Not Per Call
October 202611 min read

An agent runs a loop, and each turn resends the earlier prompt, tool calls and tool results as input. Even with previous_response_id, OpenAI bills every previous input token in the chain as input again. A ticket that takes eight turns can therefore cost many times the price of one request.
After a run, read result.context_wrapper.usage, which totals requests, input, output and cached tokens across every model call. Price the fresh input, cached input and output at your model's rates, and record whether the run completed. Divide total spend over all runs by the number of completed runs, so failed runs are counted.
The Agents SDK raises MaxTurnsExceeded when a run passes its max_turns limit. You can pass error_handlers with a max_turns handler that returns a RunErrorHandlerResult, for example an escalation message, instead of an exception. Setting include_in_history to False keeps that fallback out of the session history.
Not directly. It caps how many subagents are active at the same time, with a default of 3, but the Responses API sets no fixed limit on the total number of subagents or the tree depth. max_tool_calls is also unsupported with multi-agent, so put a spending ceiling in your own code before the request starts.
Functions marked with defer_loading stay out of the context until the model searches for them, so unused tool schemas are not sent every turn. Loaded tools are injected at the end of the context window, which keeps the cached prefix from earlier requests valid. Tool search requires gpt-5.4 or later in the Responses API.

Key Takeaway
AI agent cost optimization starts by measuring cost per completed task, not per API call, because every agent turn re-bills the growing context as input. Price each run from its usage totals, divide by completed tasks, then cap max_turns, limit subagent concurrency, defer rarely used tools, compact long chains and route simple steps to cheaper models.
Consider an ERP helpdesk agent that answers questions such as why an invoice is still unpaid. The spreadsheet that approved it priced one request: a 6,000-token prompt and a 300-token answer, roughly Rp322 at gpt-5.4 rates. The first monthly invoice from the API provider comes in an order of magnitude higher, and nobody can say which tickets caused it.
The gap is not a pricing error. An agent is a loop, and each turn resends everything that came before. This post works through AI agent cost optimization for that loop using the OpenAI Agents SDK and the Responses API: how to measure cost per completed task, and the five controls the official documentation provides for holding it down, with a worked rupiah example you can rerun with your own numbers.
A chat completion has one input and one output. An agent run has many, and they are not independent. Turn two includes turn one's prompt, its tool call and the tool's result. Turn eight includes all seven before it. The OpenAI conversation-state guide is explicit about the server-side version: even with previous_response_id, all previous input tokens in the chain are billed as input tokens again. Not sending the history does not mean not paying for it.
So the unit that matters is the completed task: one resolved ticket, one reconciled bank statement, one drafted purchase order. That is also the unit the business understands. A finance manager can decide whether Rp2,000 per resolved ticket is worth it; nobody can decide anything with a price per million tokens.
Assume a helpdesk agent with a 6,000-token prefix of instructions and tool definitions. Each turn adds about 1,500 tokens of tool output and 300 tokens of model output, and a typical ticket takes eight turns. Prices are the Standard-tier rates on the OpenAI pricing page: gpt-5.4 at 2.50 USD input, 0.25 USD cached input and 15 USD output per million tokens; gpt-5.4-mini at 0.75, 0.075 and 4.50. The exchange rate is an assumed Rp16,500 per USD. The cached rows assume each turn reuses the previous turn's whole input as a cached prefix, which is the best case rather than a guarantee.
| Scenario | Input tokens | Output tokens | Cost per task |
|---|---|---|---|
| The spreadsheet estimate: one request | 6,000 | 300 | Rp322 |
| 8 turns, gpt-5.4, no cache hits | 98,400 | 2,400 | Rp4,653 |
| 8 turns, gpt-5.4, cached prefix | 98,400 of which 79,800 cached | 2,400 | Rp1,690 |
| 8 turns, gpt-5.4-mini, cached prefix | 98,400 of which 79,800 cached | 2,400 | Rp507 |
| 12 turns, hits the cap, gpt-5.4, cached | 190,800 of which 165,000 cached | 3,600 | Rp2,636 |
The uncached eight-turn run costs about 14 times the spreadsheet estimate, and it does so without any bug. Caching recovers most of that, which is why the controls later in this post are mostly about keeping the prefix stable. Note what the twelve-turn row means: a run that loops to its limit and then gives up costs more than a run that succeeds.
Now run a month. Out of 1,000 tickets, suppose 800 complete in eight turns and 200 hit a twelve-turn cap and are escalated to a person. Spend is 800 times Rp1,690 plus 200 times Rp2,636, about Rp1.88 million. Divided by the 800 completed tasks, the real cost per resolved ticket is about Rp2,349, roughly 39 percent above the happy-path figure. That is the number to budget against.
The Agents SDK accumulates usage across every model call in a run. After Runner.run returns, result.context_wrapper.usage holds requests, input_tokens, output_tokens and total_tokens, plus input_tokens_details.cached_tokens and output_tokens_details.reasoning_tokens, and request_usage_entries gives the per-request breakdown. That is everything needed to price a run, without scraping the provider dashboard.
from agents import Agent, Runner, RunErrorHandlerInput, RunErrorHandlerResult
# USD per 1M tokens, Standard tier, OpenAI pricing page (October 2026).
PRICES = {
"gpt-5.4": {"input": 2.50, "cached": 0.25, "output": 15.00},
"gpt-5.4-mini": {"input": 0.75, "cached": 0.075, "output": 4.50},
}
IDR_PER_USD = 16_500 # assumption: use your finance team's booking rate
def task_cost_idr(usage, model: str) -> float:
p = PRICES[model]
cached = usage.input_tokens_details.cached_tokens
fresh = usage.input_tokens - cached # cached tokens are NOT free,
usd = (fresh * p["input"] # just cheaper; reasoning tokens
+ cached * p["cached"] # are already inside output_tokens
+ usage.output_tokens * p["output"]) / 1_000_000
return usd * IDR_PER_USD
helpdesk = Agent(
name="ERP helpdesk",
model="gpt-5.4",
instructions="Answer ERP questions with the tools. Never post documents.",
tools=[get_invoice, list_payments, get_customer], # your function tools
)
def run_ticket(ticket_text: str) -> dict:
hit_cap = False
def on_max_turns(_data: RunErrorHandlerInput[None]) -> RunErrorHandlerResult:
nonlocal hit_cap
hit_cap = True
return RunErrorHandlerResult(
final_output="Escalated to a human: the agent ran out of steps.",
include_in_history=False, # keep the fallback out of session storage
)
result = Runner.run_sync(
helpdesk, ticket_text,
max_turns=12, # p95 of observed turns, plus headroom
error_handlers={"max_turns": on_max_turns},
)
usage = result.context_wrapper.usage
return {
"completed": not hit_cap,
"requests": usage.requests, # one per model call in the loop
"cost_idr": task_cost_idr(usage, "gpt-5.4"),
}
# The number to put on a dashboard and in a budget:
# sum(cost_idr over ALL runs) / count(runs where completed)
# A capped run still cost money; dividing by attempts hides it.Two details in the code matter. Cached tokens are subtracted from input and priced at the cached rate, because they are cheaper rather than free. And the run records whether it completed, so the aggregate can divide total spend over all runs by completed runs only. Dividing by attempts would quietly reward an agent that gives up early.
Store one row per task: ticket id, completed flag, requests, the token fields and the computed rupiah. With a few hundred rows you can read the turn distribution directly, see which ticket types produce the long tail, and set the caps in the next section from data instead of intuition.
max_turns limits how many agent-loop iterations a run may take, where one turn is one model call. When a run exceeds it, the SDK raises MaxTurnsExceeded. Passing max_turns=None disables the limit, which is the one setting a cost-sensitive production agent should never ship with.
An exception is a poor user experience, so every Runner entry point also accepts error_handlers, a dict keyed by error kind. The max_turns handler returns a RunErrorHandlerResult with a final_output of your choosing, such as an escalation message, and include_in_history=False keeps that fallback out of the result history and session storage. The user gets a clear answer, the ticket goes to a person, and the spend stops.
Set max_turns from the measured distribution, not from a round number. A cap near the 95th percentile of turns on completed tasks, plus a little headroom, stops runaway loops while touching only the tasks that were unlikely to finish anyway. Review it when you add tools, because new tools change how many steps a task takes.
The cap is also a quality signal. A ticket type that regularly hits it usually means a missing tool, an ambiguous instruction or a task that needs a person. Fixing that is cheaper than raising the limit, because every extra turn costs more than the one before it.
The Responses API multi-agent feature lets a root agent spawn subagents within a single request. Its one width control is max_concurrent_subagents, which caps how many subagents are active at once across the whole tree, children and deeper descendants included but not the root. The default is 3, which the guide recommends for most workloads. Three other facts from the same guide shape the budget:
from openai import OpenAI
client = OpenAI()
DAILY_BUDGET_IDR = 250_000
spent_today_idr = ledger.sum_today("reconciliation") # your own store
# Multi-agent has no fixed limit on total subagents or tree depth, and
# max_tool_calls is not supported with it. Refuse to START work you cannot afford.
if spent_today_idr >= DAILY_BUDGET_IDR:
raise BudgetExhausted("reconciliation agents paused until tomorrow")
response = client.beta.responses.create(
model="gpt-6.1-sol",
input="Reconcile September bank lines against open AR invoices for PT Contoh. "
"Use at most two subagents: one per bank account.",
multi_agent={
"enabled": True,
"max_concurrent_subagents": 2, # default is 3; this caps width, not total
},
betas=["responses_multi_agent=v1"],
)
ledger.record("reconciliation", response.usage) # price it like any other taskTreat max_concurrent_subagents as rate control, not spend control. It decides how fast tokens are consumed, not how many. Put the actual ceiling outside the request: scope the task narrowly in the input, record each response's usage in your own ledger, and refuse to start new multi-agent work once a daily budget is spent.
In practice that makes multi-agent a tool for tasks with naturally parallel, bounded parts, such as one subagent per bank account in a reconciliation. For open-ended work, a single agent with max_turns, or agents-as-tools in the Agents SDK, where each tool agent can be given its own, cheaper model, keeps the budget enforceable.
Tool definitions are part of the prompt prefix. An ERP agent with forty tools carries every schema on every turn, and the tool search guide notes that unused definitions occupy context. Tool search, available on gpt-5.4 and later in the Responses API, lets you mark functions with defer_loading so they stay out of the context until the model searches for them.
erp_namespace = {
"type": "namespace",
"name": "erp",
"description": "ERP tools for invoices, payments, stock and purchasing.",
"tools": [
{ # used by almost every ticket: load it up front
"type": "function",
"name": "get_invoice",
"description": "Fetch an AR invoice by number.",
"parameters": {
"type": "object",
"properties": {"invoice_no": {"type": "string"}},
"required": ["invoice_no"],
"additionalProperties": False,
},
},
{ # needed by a minority of tickets: defer it
"type": "function",
"name": "list_stock_moves",
"description": "List stock moves for an item and warehouse.",
"defer_loading": True,
"parameters": {
"type": "object",
"properties": {
"sku": {"type": "string"},
"warehouse": {"type": "string"},
},
"required": ["sku", "warehouse"],
"additionalProperties": False,
},
},
# ...more deferred tools; the guide suggests under 10 per namespace
],
}
response = client.responses.create(
model="gpt-5.4", # tool_search needs gpt-5.4 or later
input="Why is INV-2026-0912 still unpaid?",
tools=[erp_namespace, {"type": "tool_search"}],
)
# Wrong: appending list_stock_moves to tools mid-task. The tool block sits in
# the prefix, so the change invalidates the cache for everything after it.
# Right: let tool search load it; loaded tools go to the END of the context.The cost detail is where loaded tools go. The guide states that all tools, hosted or client-executed search alike, are loaded at the end of the context window, so the cached prefix from earlier requests survives. Adding a tool to the tools array mid-task does the opposite: the tool block sits early in the prompt, so the cache is lost for everything after it. The guide also warns that changing the loaded tool set later breaks the cache from that point forward, so load once and leave it.
Group tools into namespaces, which the guide says models are primarily trained to search, and keep each namespace under about ten functions. Leave the two or three tools nearly every task needs undeferred, so the common path never pays for a search step.
Long tasks eventually hit the second problem: even with a perfect cache, the context keeps growing, and cached tokens are cheap but not free. Server-side compaction addresses this. You pass context_management with an entry of type compaction and a compact_threshold, and when the rendered token count crosses it, the server compacts the context into an encrypted compaction item that carries prior state forward in fewer tokens.
The threshold is a cost dial, not just a safety net. The compaction guide's example uses 200,000 tokens, but in the twelve-turn ticket above the input was already close to 26,000 tokens per turn by the end. Pick a threshold from your own usage log, at the point where per-turn input starts to dominate the cost of a task.
prev_id = None
for step in ticket_steps: # your own loop over tool results
response = client.responses.create(
model="gpt-5.4",
previous_response_id=prev_id, # only the NEW input is sent...
input=[{"role": "user", "content": step}],
# ...but the whole chain is billed as input every turn. Compact
# well before the context window, not at it.
context_management=[{"type": "compaction", "compact_threshold": 60_000}],
)
prev_id = response.id
log_usage(response.usage) # watch input_tokens fall after compactionCompaction replaces earlier context with a shorter representation, so the request after a compaction cannot reuse the old cached prefix. Compacting too often trades a large context for repeated cache misses. And the compaction item is opaque and not human-readable, so log anything you need for audit, such as which invoice was checked, before it is compacted away.
If you chain turns with previous_response_id, remember the billing rule from the first section: the whole chain is billed as input every turn. Compaction is the lever that shortens the chain. Without it, a long-running chain grows until the context window, not the budget, stops it.
Not every turn needs the strongest model. Extracting an invoice number, choosing a tool or formatting a reply are small steps; diagnosing a payment mismatch is not. The price gap between tiers is wide enough that routing matters as much as caching. The rates below are the Standard-tier prices per million tokens from the OpenAI pricing page in October 2026.
| Model | Input, USD per 1M | Cached input, USD per 1M | Output, USD per 1M |
|---|---|---|---|
| gpt-5.4 | 2.50 | 0.25 | 15.00 |
| gpt-5.4-mini | 0.75 | 0.075 | 4.50 |
| gpt-5.4-nano | 0.20 | 0.02 | 1.25 |
In the worked example, the same eight-turn ticket drops from Rp1,690 to Rp507 on gpt-5.4-mini. Whether mini resolves the same share of tickets is the question your usage log and evals must answer, because a cheaper model that completes fewer tasks can cost more per completed task. Pulled together, a per-task budget looks like this:
The rule to carry is simple: an agent's price is cost per completed task, and it includes the runs that failed. Measure that first, from the usage object rather than the dashboard, and the controls stop being guesses. max_turns bounds the tail, a concurrency cap bounds the rate but not the total, tool search and compaction keep the context cheap, and routing picks the cheapest model that still finishes the job.