ERP
ERP AI Agent Tool Design: Drafts, Not Direct Writes
October 202612 min read

In most ERP designs, no. Give the agent read tools and tools that create documents in draft status, and leave submitting, approving, posting and sending to a person in the ERP under the existing approval matrix. A wrong draft is deleted; a wrong posted journal needs a reversal and may be impossible in a closed period.
ERP documents already have a state machine where draft is the only freely editable state. Writing into draft means the agent's output lands in the review queue the business already runs, with no stock movement, ledger entry or external message until a person acts. It turns the agent into a fast clerk rather than an unsupervised approver.
Give every write tool an idempotency key built from the agent run id, which your executor owns, plus a client reference the model reuses when retrying. Claim the key in the same database transaction that creates the draft, return the existing draft when the key is replayed, and reject reuse of a key with different arguments.
No. Resolve the branch, the user and their scopes from the authenticated request and filter every query by them. If branch_id is an argument, a prompt-injected note in a vendor record can change it. The MCP tools specification also allows the tool list itself to vary by the caller's authorization, so users only see tools their scopes permit.
Return them as tool execution errors with isError set to true, and state what failed, what was not done, and which call to make next, using document numbers rather than internal ids. The MCP specification describes these errors as actionable feedback that lets the model self-correct. An opaque code such as 422 usually makes the model resend the same call.

Key Takeaway
Good ERP AI agent tool design gives the model three tiers: read tools, tools that create draft documents, and posting tools it never holds. Every write takes an idempotency key, branch and user come from the token rather than the arguments, errors tell the model what to do next, and each document records which agent run created it.
Consider a distributor with three branches that wants an agent to turn reorder alerts into purchase orders. The first prototype is usually a single tool wrapping the ERP's REST API: method, path, body. It works in the demo. Then a timeout makes the agent retry, two identical orders reach the vendor, and nobody can say whether a person or the agent pressed the button. The fault is not the model. It is the tool surface it was given.
ERP AI agent tool design is the discipline of deciding what the model can see, what it can change, and in what state it leaves a document. This post sets out an opinionated design for that surface, drawn from how ERP modules already work: documents have a draft state, posting is a separate act, and every change is attributable. The guidance cites OpenAI's function calling and tool search docs, the MCP 2026-07-28 tools specification and Anthropic's engineering notes on agent tools.
Sort every operation the agent might want into one of three tiers before writing a single schema. The tier decides the approval policy, the credential, and whether the tool exists for the agent at all.
| Tier | What it may change | Examples | Approval policy |
|---|---|---|---|
| Read | Nothing. Returns data the caller could already see in the ERP screens. | Search vendors, get a purchase order, stock on hand, open invoices for a customer | No per-call approval; limited by scope and branch |
| Draft | Creates or edits a document in draft status only. No stock movement, no journal entry, nothing sent outside. | PO draft, sales quotation draft, journal voucher draft, stock adjustment request | Approval on the tool call where the action is unusual; review of the document always |
| Post | Submits, approves, posts, sends or cancels. Creates ledger entries or external side effects. | Post a journal, approve a PO, send an invoice to the customer, void a payment | Not exposed to the agent. A person does it in the ERP, under the existing approval matrix |
The third row is the design decision. Posting tools are left out of the agent's tool list entirely rather than gated behind a confirmation prompt, because a confirmation prompt is one click a tired reviewer makes without reading. OpenAI's remote MCP guide tells developers to use allowed_tools and require_approval so that sensitive actions go through an approval flow; with posting removed, that flow only has to cover the draft tier, where a mistake costs a deleted draft rather than a reversal journal.
ERP documents already have a state machine: draft, submitted, approved, posted, sometimes cancelled. A draft can be edited or deleted freely; a posted document can only be reversed with another document, and in a closed period not even that. Writing into the draft state means the agent produces exactly what a junior clerk produces, and it lands in the review queue the business already runs. The schema below is the whole contract for one such tool.
// One business action, one tool. The model can describe a draft;
// it cannot choose a status, a branch, a price or a posting date.
export const createPurchaseOrderDraft = {
type: "function",
name: "purchasing_create_po_draft",
description:
"Create a DRAFT purchase order for the caller's branch. A draft is not " +
"sent to the vendor and does not touch stock or the ledger until a person " +
"submits it in the ERP. Prices come from the vendor price list. Returns " +
"the draft number and server-computed totals.",
strict: true,
parameters: {
type: "object",
additionalProperties: false,
// Strict mode: every property is listed as required;
// optional ones accept null instead of being omitted.
required: ["vendor_code", "needed_by", "client_ref", "note", "lines"],
properties: {
vendor_code: {
type: "string",
description: "Vendor code from purchasing_search_vendors, e.g. V-00231",
},
needed_by: { type: "string", description: "Date needed, YYYY-MM-DD" },
client_ref: {
type: "string",
description: "Your own id for this order within the run. Reuse it when retrying.",
},
note: {
type: ["string", "null"],
description: "Why this order is needed. Shown to the reviewer.",
},
lines: {
type: "array",
items: {
type: "object",
additionalProperties: false,
required: ["item_code", "qty", "uom"],
properties: {
item_code: { type: "string" },
qty: { type: "number" },
// An enum, not free text: "kg", "Kg" and "kilo" are three bugs.
uom: { type: "string", enum: ["PCS", "BOX", "KG", "LTR"] },
},
},
},
// Deliberately absent: branch_id, status, unit_price, posting_date.
},
},
} as const;Notice what the schema leaves out. There is no branch_id, so the model cannot file the order under another branch. There is no status, so it cannot create anything except a draft. There is no unit_price, because prices come from the vendor price list on the server; a model that can type a price can also type a wrong one. OpenAI's function calling guide recommends strict mode, which requires additionalProperties set to false and every property listed as required, with null as the way to say a field is optional. It also recommends enums and object structure to make invalid states impossible, which is why uom is an enum instead of a string.
The description does as much work as the schema. It says what a draft is, what it does not do, and what comes back. The same guide proposes an intern test: could a person use the function correctly given only what the model was given? A description that says Create PO fails that test, because it does not say whether the vendor is notified or whether stock is reserved.
A generic endpoint tool hands the model the whole API and moves every permission decision into string parsing. A generic update_document tool is the same mistake in a nicer shape: one tool, every document type, every field. Replace both with one tool per business action, named with a module prefix, and grouped into one namespace per ERP module.
// Wrong: one tool that can do anything the REST API can do.
// The model now chooses between GET /items and POST /journal-entries/123/post,
// and every permission check has to be re-derived from a free-form path.
{ type: "function", name: "erp_api",
parameters: { method: "string", path: "string", body: "object" } }
// Wrong, more subtly: a generic writer keyed by document type.
{ type: "function", name: "update_document",
parameters: { doctype: "string", id: "string", fields: "object" } }
// Right: one namespace per ERP module, each under 10 functions,
// loaded only when the task needs that module.
const tools = [
{
type: "namespace",
name: "purchasing",
description:
"Purchasing: search vendors and items, read purchase orders, create PO drafts. " +
"Cannot approve, send or cancel orders.",
tools: [
{ type: "function", name: "purchasing_search_vendors", defer_loading: true, /* ... */ },
{ type: "function", name: "purchasing_get_po", defer_loading: true, /* ... */ },
{ type: "function", name: "purchasing_create_po_draft", defer_loading: true, /* ... */ },
],
},
{
type: "namespace",
name: "inventory",
description: "Inventory: stock on hand and reorder points per warehouse. Read only.",
tools: [/* inventory_get_stock, inventory_list_below_reorder */],
},
{ type: "tool_search" },
];OpenAI's function calling guide advises keeping the initially available functions small for accuracy, aiming for fewer than 20 at the start of a turn, and combining functions that are always called in sequence. Its tool search guide goes further for large surfaces: put functions in namespaces, mark them with defer_loading so the model sees only the namespace name and description until it searches, and keep each namespace to fewer than 10 functions. An ERP maps onto that naturally, since purchasing, inventory, sales and finance are already separate modules with separate owners.
Anthropic's guidance on writing agent tools makes the same two points from the other side: consolidate several API calls into one tool that matches the task, and use prefixes such as purchasing_ to keep boundaries clear when many tools are loaded. The namespace description should also state what the module cannot do. Cannot approve, send or cancel orders costs a dozen tokens and stops the model from searching for a tool that does not exist.
Agents retry. The HTTP call to the tool server times out after the draft was committed; the runtime resends it. A run crashes and resumes from its last checkpoint; the model emits the same call again. Without a key, each retry is a new purchase order. The fix is the same one payment APIs use: claim a key in the same transaction that creates the document, and return the original result when the key is seen again.
-- Every write tool claims its key before it does anything else.
CREATE TABLE agent_write_keys (
idempotency_key text PRIMARY KEY, -- agent_run_id || ':' || client_ref
tool_name text NOT NULL,
request_hash text NOT NULL, -- sha256 of the canonical arguments
document_type text,
document_id bigint,
created_at timestamptz NOT NULL DEFAULT now()
);export async function createPoDraft(ctx: ToolContext, args: PoDraftArgs) {
// The run id comes from the executor, so two runs can never collide,
// and a retry inside one run always lands on the same key.
const key = `${ctx.agentRunId}:${args.client_ref}`;
const hash = sha256(canonicalJson(args));
return db.transaction(async (tx) => {
const claimed = await tx.query(
`INSERT INTO agent_write_keys (idempotency_key, tool_name, request_hash)
VALUES ($1, 'purchasing_create_po_draft', $2)
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING idempotency_key`,
[key, hash],
);
if (claimed.rowCount === 0) {
// A concurrent duplicate blocks on the primary key until the first
// transaction commits, so document_id is already filled in here.
const prior = await tx.one(
`SELECT request_hash, document_id FROM agent_write_keys
WHERE idempotency_key = $1`,
[key],
);
if (prior.request_hash !== hash) {
return toolError(
`client_ref ${args.client_ref} was already used in this run for a different ` +
`purchase order. Use a new client_ref for a new order, or resend the original ` +
`arguments to get the existing draft back.`,
);
}
return draftSummary(tx, prior.document_id, { replayed: true });
}
const po = await insertDraftPo(tx, ctx, args); // status = 'draft', always
await tx.query(
`UPDATE agent_write_keys SET document_type = 'purchase_order', document_id = $2
WHERE idempotency_key = $1`,
[key, po.id],
);
return draftSummary(tx, po.id, { replayed: false });
});
}The key is built from the agent run id, which the executor owns, plus a client_ref the model supplies and is told to reuse on retry. Neither part alone is enough. A key the model invents with no run scope can collide with another run, and a key from the run id alone allows only one write per run. The request hash catches the remaining case: the same key sent with different arguments, which is almost always the model reusing a reference by accident. That gets an error explaining the choice, not a silent overwrite.
A replay returns the existing draft with a replayed flag instead of an error, because from the model's point of view the call succeeded and it should carry on. The flag goes into the audit row so that a run with many replays stands out when someone reviews why it was slow or expensive.
In a multi-branch ERP the branch is a security boundary, not a filter. If branch_id is a tool argument, it is something a prompt-injected vendor note can change. Resolve the branch, the user and the scopes from the authenticated request into a context object the model never touches, and filter every query by it.
// Built from the authenticated request. Nothing here comes from model output.
interface ToolContext {
agentRunId: string;
onBehalfOf: { userId: number; roles: string[] };
branchId: number; // resolved from the user's token or session
branchCode: string; // "SBY-01", for messages the model can read
scopes: Set<string>; // "purchasing:read", "purchasing:draft", ...
}
// Wrong: branch_id as an argument. A vendor note that says
// "file this under branch 7" is now an instruction the model can follow.
// Right: no branch argument at all; every query is filtered by ctx.branchId.
export async function getPurchaseOrder(ctx: ToolContext, args: { po_number: string }) {
const po = await db.maybeOne(
`SELECT * FROM purchase_orders WHERE po_number = $1 AND branch_id = $2`,
[args.po_number, ctx.branchId],
);
if (!po) {
// One message for "does not exist" and "belongs to another branch",
// so the tool cannot be used to probe what other branches hold.
return toolError(
`No purchase order ${args.po_number} in branch ${ctx.branchCode}. ` +
`Use purchasing_search_pos to find the right number.`,
);
}
return summarisePo(po);
}
// The tool list itself follows the caller's grants:
// a user without purchasing:draft never sees the draft tool.
export function listTools(ctx: ToolContext) {
return ALL_TOOLS.filter((tool) =>
tool.requiredScopes.every((scope) => ctx.scopes.has(scope)),
);
}The MCP 2026-07-28 tools specification allows the tools/list result to vary by the authorization on the request, for example returning only the tools the caller's granted scopes permit. Use that: a warehouse clerk's agent session should not even see purchasing_create_po_draft. The same specification's note on stateful tools says a handle such as a document number is a name, not a capability, and that the server should validate the caller's authorization against it on every call. A draft number returned an hour ago is still checked against the branch today.
Do not enforce any of this with tool annotations or descriptions. The MCP specification says clients must treat tool annotations as untrusted unless they come from trusted servers, and a description that says read only is a hint to the model, not a control. The read tier is read only because its database role has no write grant, and the draft tier cannot post because no code path in it changes status.
A tool error is read by a model that will decide its next call from the text alone. The MCP specification separates protocol errors, such as an unknown tool, from tool execution errors returned in the result with isError set to true, and describes the latter as actionable feedback the model can use to self-correct and retry. Clients should pass execution errors back to the model. So write them for that reader.
// Wrong: accurate, and useless to a model. It will resend the same call.
{ "content": [{ "type": "text", "text": "ERR_VALIDATION 422" }], "isError": true }
// Right: what failed, the allowed values, what to do next, and what did NOT happen.
{
"content": [{
"type": "text",
"text": "Item BRG-0042 is not stocked in KG. Allowed units for BRG-0042: PCS, BOX (1 BOX = 24 PCS). Resend the line with uom PCS or BOX. No draft was created."
}],
"isError": true
}| Situation | Opaque error | Written for a model |
|---|---|---|
| Vendor is blocked | 403 Forbidden | Vendor V-00231 is blocked for new orders by finance since 12 September. No draft was created. Choose another vendor for item BRG-0042 or tell the user the vendor needs unblocking. |
| Accounting period closed | PERIOD_LOCKED | August 2026 is closed in branch SBY-01. Journal drafts can only be dated 1 September 2026 or later. Resend with a date in the open period. |
| Draft already submitted | 409 Conflict | PO-SBY-2026-0918 was submitted by a user and can no longer be edited by the agent. Create a new draft or ask the user to return it to draft. |
| Item code not found | Item not found | No item BRG-042 in this branch. Close matches: BRG-0042 Kardus 40x40, BRG-0420 Lakban Coklat. Call inventory_search_items to confirm. |
Every good message in that table states what was not done, because a model that is unsure whether a write happened will often try again. They also use document numbers and names instead of internal ids. Anthropic's guidance reports that resolving opaque UUIDs to meaningful identifiers improves precision, and the same holds for reviewers reading the transcript later. Keep internal ids in the structured result for chaining calls, and put the human-readable version in the text.
An ERP audit log usually records a user and a timestamp. With an agent in the loop that answer is incomplete: the user did not type the order, the agent did, on the user's behalf, in a specific run, with a specific prompt version. Record all of it, from the executor, for every call including the failed and denied ones.
-- One row per tool call, written by the executor, never by the model.
CREATE TABLE agent_tool_calls (
id bigserial PRIMARY KEY,
agent_run_id text NOT NULL,
tool_call_id text NOT NULL, -- id of the model's function call item
tool_name text NOT NULL,
on_behalf_of integer NOT NULL REFERENCES users(id),
branch_id integer NOT NULL,
arguments jsonb NOT NULL,
outcome text NOT NULL
CHECK (outcome IN ('ok', 'tool_error', 'denied', 'replayed')),
document_type text,
document_id bigint,
model text NOT NULL,
prompt_version text NOT NULL,
created_at timestamptz NOT NULL DEFAULT now()
);
-- Provenance on the document itself, so the reviewer sees it without a join
-- and it survives when tool-call logs are rotated.
ALTER TABLE purchase_orders
ADD COLUMN drafted_by_agent_run text,
ADD COLUMN drafted_on_behalf_of integer REFERENCES users(id);The tool_call_id ties the row to the exact function call item in the model transcript, so a reviewer can open the run and see the reasoning that led to the document. The two columns on purchase_orders keep provenance with the document, which matters because tool-call logs are often kept for weeks while documents are kept for years. Show it in the UI too: a Drafted by agent badge with a link to the run tells the approver to read the lines rather than trust the totals.
Return server-computed results from every draft tool: the document number, line totals after pricing, tax, and any warnings such as a quantity far above the usual order. Then have the agent restate those numbers to the user. A reviewer comparing what the agent said with what the ERP calculated catches mistakes that neither would catch alone.
Run each new tool through these questions before it reaches an agent's tool list. A no to any of them is a design change, not a documentation fix.
The model is the least controllable part of an ERP agent, so the tool surface has to carry the controls. Give it read tools and draft tools, keep posting with people, and make every write idempotent, branch-scoped and attributable. Then a retry produces a replay, an injected instruction finds no argument to change, and an auditor can trace every agent-made document to a run.
Related posts on this site
Sources