AI
WhatsApp AI Agent Indonesia: UMKM Order Bot with Agents SDK
October 202612 min read

Yes, when the AI supports the business's own customer service. Meta's WhatsApp Business Platform terms prohibit AI Providers from using the platform where AI is the primary function, such as a general ChatGPT-style assistant, but a shop agent that answers questions about its own products and orders is a secondary, ancillary use. Check the current terms before launch, because they have been revised several times.
Since 1 July 2025 Meta charges per delivered template message. Non-template replies and utility templates sent inside the 24-hour customer service window are free, so an agent that only replies to customers who have just messaged pays no messaging fee. Marketing templates are always charged, and the per-country rate card on Meta's pricing page has the actual rupiah figures.
Use an OpenAI Agents SDK session keyed by the customer's WhatsApp id, for example SQLAlchemySession on Postgres or RedisSession. Every worker that processes a message from that number reads and appends to the same history. Add a per-number lock so several quick message bubbles are processed in order rather than in parallel.
Give the triage agent an Agents SDK handoff with an on_handoff callback. When the model calls it, the callback sets a human-mode flag in the database and notifies the owner, and the worker skips the agent while that flag is set. The admin then replies from your own inbox through the same Cloud API number, inside the 24-hour window.
Meta retries any webhook that does not receive HTTP 200, for up to seven days, and those retries can deliver the same message more than once. If your endpoint waits for the AI model before returning, a slow model causes a retry and a duplicate answer. Return 200 immediately, process the message in a worker, and dedupe on the WhatsApp message id.

Key Takeaway
A WhatsApp AI agent for an Indonesian UMKM should answer only inside the 24-hour customer service window, where Meta charges nothing for non-template replies. Read stock and prices from tools rather than the prompt, keep one Agents SDK session per phone number, and hand payment, shipping-cost and complaint questions to a human through a handoff.
Picture a coffee-bean shop in Bandung that sells through WhatsApp. At nine in the evening the owner has forty unread chats, and most of them are the same four questions: is the Gayo still in stock, how much is the 500 gram bag, can I pay by transfer, and when will it ship. Each one waits until morning, and some of those customers buy from the shop that answered first.
This guide builds a WhatsApp AI agent for that kind of Indonesian UMKM on the WhatsApp Cloud API and the OpenAI Agents SDK for Python: a webhook that survives Meta's retries, tools that read real stock and prices, one session per phone number, prompts written in Bahasa Indonesia, and a handoff to a human admin. The raw Cloud API setup is covered in the existing WhatsApp Business API integration post on this site; this one is about the agent on top, and every pricing and platform claim is taken from Meta's and OpenAI's own documentation.
Yes, for this use. Meta's terms for the WhatsApp Business Platform prohibit AI Providers from using the platform to distribute AI technology when that technology is the primary function of the service, and Meta's AI-provider pricing page dates that policy from 15 January 2026. The terms separate a primary function from one that is secondary or ancillary. A shop whose agent answers questions about its own products and takes its own orders is a business using AI in a support workflow, which is the ancillary case. A number that offers to answer anything at all, ChatGPT-style, is the case the rule was written to stop.
That distinction is not only legal housekeeping; it shapes the design. The agent's instructions scope it to the shop's catalogue and orders, off-topic requests get a polite refusal, and the tools it has can only touch that shop's data. A narrow agent is easier to defend under the terms, cheaper to run, and less likely to say something the owner has to apologise for.
Check the current terms before launch, not this article. The Meta Terms for WhatsApp Business Platform page was last updated on 23 September 2026 and the AI-provider rules have changed several times since late 2025. If your product is an assistant that many shops resell, rather than one shop's own customer service, read the AI Providers section with a lawyer, because that is the scenario it targets.
Since 1 July 2025 Meta bills the WhatsApp Business Platform per message, and you are charged only when a template message is delivered. Every message a customer sends opens or refreshes a 24-hour customer service window, and inside it your non-template replies cost nothing. That single rule decides most of the architecture: an agent that only replies to inbound messages, promptly, runs at zero messaging cost.
| What you send | Inside the 24-hour window | Outside the window | Design consequence |
|---|---|---|---|
| Free-form reply from the agent | Free | Not allowed at all; only templates can be sent | Check the window before sending, not before running the model |
| Utility template, such as an order confirmation | Free | Charged per delivered message | Send order and shipping updates while the customer is still chatting |
| Marketing template, such as a promo blast | Charged | Charged | Never give the agent a tool that sends one |
| Any reply to a chat opened from a Click-to-WhatsApp ad or Facebook Page button | Free, and if you reply within 24 hours a 72-hour free entry point window opens | Normal rules resume | Ad-driven chats can take a longer, slower sales conversation |
Authentication templates are charged as well, but a shop agent has no reason to send them. The practical rules are short: the agent sends free-form text only, templates are sent only by deliberate code paths a person configured, and the per-market rate card on Meta's pricing page is the only place to take an actual rupiah figure from. Rates change, so a cost estimate belongs in a spreadsheet that links to that page rather than hard-coded in this post.
Meta signs every webhook POST with an HMAC-SHA256 of the raw body using your app secret and sends it in the X-Hub-Signature-256 header. If your endpoint answers anything other than 200, Meta retries with decreasing frequency for up to seven days, and the documentation warns that retries can produce duplicate notifications. Both facts point at the same shape: the webhook does no AI work at all.
# webhook.py - FastAPI. The only job here: prove it is Meta, dedupe, enqueue, return 200.
import hashlib, hmac, json, os
from fastapi import FastAPI, Request, Response, HTTPException
app = FastAPI()
APP_SECRET = os.environ["META_APP_SECRET"].encode()
VERIFY_TOKEN = os.environ["WA_VERIFY_TOKEN"]
@app.get("/webhook/whatsapp")
async def verify(request: Request):
q = request.query_params
if q.get("hub.mode") == "subscribe" and q.get("hub.verify_token") == VERIFY_TOKEN:
# Raw challenge as the body - not JSON, not quoted.
return Response(content=q.get("hub.challenge"), media_type="text/plain")
raise HTTPException(status_code=403)
@app.post("/webhook/whatsapp")
async def receive(request: Request):
raw = await request.body() # sign the RAW bytes, never re-serialised JSON
sent = request.headers.get("X-Hub-Signature-256", "").removeprefix("sha256=")
expected = hmac.new(APP_SECRET, raw, hashlib.sha256).hexdigest()
if not hmac.compare_digest(sent, expected):
raise HTTPException(status_code=401)
payload = json.loads(raw)
for entry in payload.get("entry", []):
for change in entry.get("changes", []):
value = change.get("value", {})
for msg in value.get("messages", []):
# Meta retries non-200 deliveries for up to 7 days, so the same
# message id WILL arrive twice. SET NX makes the second one a no-op.
if await redis.set(f"wa:seen:{msg['id']}", 1, nx=True, ex=8 * 86400):
await queue.enqueue("handle_inbound", msg, value["metadata"])
# value["statuses"] (sent/delivered/read) goes to a separate, cheap handler.
# Return before any LLM call. A slow model must never turn into a retry storm.
return Response(status_code=200)The model call happens in a worker. If the endpoint waited for the model and the model was slow, Meta would see a timeout, retry, and you would answer the customer twice; the SET NX on the message id is what turns that second delivery into a no-op. The statuses array, which reports sent, delivered and read, arrives on the same webhook and is useful for an admin dashboard, but it should never trigger an agent run.
The most expensive mistake a shop agent can make is quoting a wrong price, because the customer screenshots it. So prices and stock never live in the prompt. The agent gets two tools: one that looks up products in the shop's own table and returns price and stock as of this second, and one that creates an order as a draft. The run context carries the shop id and the customer's WhatsApp id, so the model cannot ask about another shop or create an order for another number.
# tools.py - the model names a product; the database answers. Prices never come from the prompt.
from dataclasses import dataclass
from agents import RunContextWrapper, function_tool
@dataclass
class ShopContext:
shop_id: int
wa_id: str # the customer's WhatsApp id, e.g. "6281234567890"
customer_name: str
@function_tool
async def cek_stok_harga(ctx: RunContextWrapper[ShopContext], nama_produk: str) -> str:
"""Cari produk di katalog toko dan kembalikan harga serta stok saat ini.
Args:
nama_produk: Nama atau kata kunci produk yang ditanyakan pelanggan.
"""
rows = await db.fetch(
"""SELECT sku, nama, varian, harga_idr, stok
FROM produk
WHERE shop_id = $1 AND aktif AND nama ILIKE '%' || $2 || '%'
LIMIT 5""",
ctx.context.shop_id, nama_produk,
)
if not rows:
return "TIDAK_DITEMUKAN" # an explicit token the prompt tells the agent how to handle
rupiah = lambda n: f"Rp{n:,}".replace(",", ".") # Rp125.000, the way customers write it
return "\n".join(
f"{r['sku']} | {r['nama']} {r['varian'] or ''} | {rupiah(r['harga_idr'])} | stok {r['stok']}"
for r in rows
)
@function_tool
async def buat_draft_pesanan(
ctx: RunContextWrapper[ShopContext], sku: str, jumlah: int, alamat: str
) -> str:
"""Buat DRAFT pesanan. Pesanan baru diproses setelah admin mengonfirmasi pembayaran.
Args:
sku: SKU persis seperti yang dikembalikan cek_stok_harga.
jumlah: Jumlah barang, minimal 1.
alamat: Alamat pengiriman lengkap dari pelanggan.
"""
if jumlah < 1 or jumlah > 50:
return "JUMLAH_TIDAK_VALID"
# Status 'draft' on purpose: the agent can take an order, it cannot accept money.
draft_id = await db.fetchval(
"""INSERT INTO pesanan (shop_id, wa_id, sku, jumlah, alamat, status)
VALUES ($1, $2, $3, $4, $5, 'draft') RETURNING id""",
ctx.context.shop_id, ctx.context.wa_id, sku, jumlah, alamat,
)
return f"DRAFT-{draft_id}"Two details matter more than they look. The tools return explicit tokens such as TIDAK_DITEMUKAN and JUMLAH_TIDAK_VALID, and the prompt says what to do with each, so a miss becomes a clear sentence instead of an improvised guess. And the order is created with status draft: the agent can collect the product, quantity and address, but turning a draft into a paid order happens only after a person has seen the transfer. The docstrings are written in Indonesian because they become the tool descriptions the model reads, in the same language as the conversation.
The Agents SDK keeps conversation history in a session object keyed by a session id, and its built-in backends include SQLiteSession, RedisSession and SQLAlchemySession. For a WhatsApp agent the natural key is the customer's WhatsApp id, so each phone number gets its own memory, and any worker that picks up the next message reads the same history from Postgres.
# worker.py - one queue job per inbound message
# pip install "openai-agents[sqlalchemy]" asyncpg
import os
from datetime import datetime, timedelta, timezone
from agents import Runner
from agents.extensions.memory import SQLAlchemySession
from sqlalchemy.ext.asyncio import create_async_engine
engine = create_async_engine(os.environ["DATABASE_URL"]) # postgresql+asyncpg://...
WINDOW = timedelta(hours=24)
async def handle_inbound(msg: dict, metadata: dict):
wa_id = msg["from"]
if msg["type"] != "text":
return await send_text(wa_id, "Maaf Kak, untuk saat ini kami baru bisa membaca pesan teks ya.")
# Customers send three bubbles in a row. A per-number lock keeps those turns
# in order instead of three agents racing on the same history.
async with redis.lock(f"wa:lock:{wa_id}", timeout=120):
if await is_with_human(wa_id):
return # an admin owns this chat; the bot stays silent
await mark_read_with_typing(msg["id"])
session = SQLAlchemySession(
f"wa:{wa_id}", # one session per phone number
engine=engine,
create_tables=True,
ensure_ascii=False, # keep "Rp", emoji and Indonesian text readable in the DB
)
ctx = ShopContext(shop_id=SHOP_ID, wa_id=wa_id, customer_name=await name_for(wa_id))
result = await Runner.run(triage_agent, msg["text"]["body"], session=session, context=ctx)
# A backlog can outlive the free window. Check before sending, not before running.
sent_at = datetime.fromtimestamp(int(msg["timestamp"]), tz=timezone.utc)
if datetime.now(timezone.utc) - sent_at > WINDOW:
return await flag_for_template_followup(wa_id, result.final_output)
await send_text(wa_id, result.final_output)
async def mark_read_with_typing(message_id: str):
# One call: blue ticks plus a typing indicator (dismissed on reply or after 25 s).
await graph_post(f"/{PHONE_NUMBER_ID}/messages", {
"messaging_product": "whatsapp",
"status": "read",
"message_id": message_id,
"typing_indicator": {"type": "text"},
})The lock is the part people skip. Indonesian customers often type a question across three short bubbles, which arrive as three webhooks a second apart; without a per-number lock, three workers run three agents on the same history and reply in a random order. The window check sits after the run on purpose: if a queue backlog ever delays a job past 24 hours, the reply is held for a person to send as a template rather than failing at the API. The read call doubles as a typing indicator, which WhatsApp dismisses when you reply or after 25 seconds.
Pass ensure_ascii=False to SQLAlchemySession. It only changes how history is serialised in the database, but it means a support person reading a stored conversation sees Kak, Rp125.000 and the customer's emoji instead of escaped unicode, which matters the first time someone has to investigate a complaint.
The instructions are written in Indonesian, for the same reason the tool descriptions are: the model mirrors the language and register it is given. A prompt written in English and told to reply in Indonesian tends to produce stiff, translated sentences that customers notice. Four things are worth stating explicitly.
INSTRUKSI_TOKO = """
Kamu adalah admin chat Toko Kopi Sari Rasa di Bandung. Kamu HANYA membantu soal
produk toko ini: stok, harga, varian, cara pesan, dan status pesanan.
Gaya bahasa:
- Panggil pelanggan "Kak". Ramah, singkat, maksimal 3 kalimat per balasan.
- Pelanggan sering menulis singkat ("brp kak", "ready?", "bs cod?"). Pahami, jangan koreksi.
- Ikuti bahasa pelanggan: kalau mereka menulis dalam bahasa Inggris, balas dalam bahasa Inggris.
Aturan keras:
- Harga dan stok WAJIB dari tool cek_stok_harga. Jangan pernah menebak angka.
- Kalau tool mengembalikan TIDAK_DITEMUKAN, katakan produknya belum tersedia.
- Ongkir, diskon, dan konfirmasi pembayaran BUKAN wewenangmu: serahkan ke admin.
- Pertanyaan di luar toko (PR sekolah, politik, resep umum): tolak dengan sopan,
arahkan kembali ke produk.
"""The prompt also carries the scope rule from the first section: questions outside the shop get a polite refusal and a nudge back to the products. That one paragraph is what keeps the agent on the ancillary side of Meta's terms, and it is worth testing with a few deliberately off-topic messages before launch.
A handoff in the Agents SDK is exposed to the model as a tool, and calling it transfers the conversation to another agent. Customising it with handoff() lets you rename that tool, give it a description the model will follow, require structured input, and run an on_handoff callback when it fires. That callback is where the real handover happens, in code.
# agents_setup.py
from pydantic import BaseModel
from agents import Agent, RunContextWrapper, handoff
from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX
class AlasanHandover(BaseModel):
alasan: str # "minta ongkir", "komplain barang rusak", "bukti transfer"
async def serahkan_ke_admin(ctx: RunContextWrapper[ShopContext], data: AlasanHandover):
# Side effects live in code, not in the model's goodwill:
# flip the chat to human mode and ping the shop owner's phone.
await set_with_human(ctx.context.wa_id, reason=data.alasan)
await notify_owner(f"Chat {ctx.context.customer_name} perlu admin: {data.alasan}")
admin_manusia = Agent[ShopContext](
name="Admin manusia",
instructions=(
"Beri tahu pelanggan bahwa admin akan membalas sebentar lagi, "
"jam kerja 08.00-20.00 WIB. Satu kalimat. Jangan menjawab pertanyaannya sendiri."
),
)
triage_agent = Agent[ShopContext](
name="CS Toko",
instructions=f"{RECOMMENDED_PROMPT_PREFIX}\n{INSTRUKSI_TOKO}",
tools=[cek_stok_harga, buat_draft_pesanan],
handoffs=[
handoff(
admin_manusia,
tool_name_override="serahkan_ke_admin",
tool_description_override=(
"Pakai untuk ongkir, diskon, komplain, bukti transfer, "
"atau jika pelanggan minta bicara dengan orang."
),
on_handoff=serahkan_ke_admin,
input_type=AlasanHandover,
)
],
)When the triage agent decides a question needs a person, it calls serahkan_ke_admin with a reason. The callback marks the chat as human-owned and notifies the owner, and the admin agent sends a single sentence telling the customer someone will reply during working hours. From then on the worker's is_with_human check skips the agent entirely, so the bot cannot talk over the admin.
The admin's replies should go out through the same Cloud API number from a simple inbox your application provides, and they follow the same pricing rule: a free-form reply is free and allowed only while the customer's 24-hour window is open. The SDK's result.last_agent tells you which agent finished a run, which is useful for logging, but human mode is a database flag rather than an agent so it survives restarts. Release it with an explicit admin action or after a fixed quiet period, and log both.
Most failures of a WhatsApp agent are not model failures. They are ordinary distributed-systems problems that show up in a chat, where the customer sees them immediately.
| Symptom the customer sees | Usual cause | Fix |
|---|---|---|
| The same answer arrives twice | Meta retried a webhook that was slow to return 200 | Return 200 before any model call and dedupe on the message id |
| Replies arrive in the wrong order or contradict each other | Three message bubbles were processed concurrently | A per-number lock around the whole run |
| A price from yesterday is quoted today | The model reused a tool result from earlier in the session history | Instruct it to call the price tool on every turn that mentions a price, and keep history short |
| The bot keeps answering after the admin has joined | Human mode lives in memory or in the prompt, not in the database | Check a persisted flag before every run |
| An unexpected messaging bill at the end of the month | An automation sent templates outside the window, such as follow-ups to quiet chats | Give the agent no template tool and review every template-sending code path |
Before launch, run the agent against a list of real questions from the shop's chat history: stock checks, price questions in shorthand, an order with a full address, a complaint, a request for shipping cost and two off-topic messages. Each one should end in either a correct tool-backed answer or a handoff, and none should end in an invented number.
A good UMKM agent on WhatsApp is narrow on purpose. It replies only when the customer has spoken, so it stays inside the free window; it reads every price from the shop's own data; it remembers each phone number separately; and it knows the questions it must hand to a person. Build those four constraints first and the model choice becomes the least important decision in the project.