AI
Mastra TypeScript Agent Framework Tutorial: Workflows and Evals
October 202613 min read

Mastra is an open-source TypeScript framework for AI agents, workflows, memory, evals and tracing, built by the team behind Gatsby. Version 1.0.0 of @mastra/core was published on npm on 20 January 2026, and the founders reported production use at companies such as Replit, PayPal and Sanity. The 1.x line releases often, so pin versions in production.
Yes. Mastra is licensed under Apache 2.0 for everything outside directories named ee in the repository. Those ee directories, which include some auth and Agent Builder code, carry a separate enterprise licence, so check the LICENSE file before relying on them.
An agent is a tool-calling loop: the model decides which tools to call and when to stop, which suits open-ended tasks. A workflow is a set of typed steps chained with then, parallel, branch and foreach, where your code decides the order. Workflows can suspend and resume from storage, which makes them the right place for approval rules and human-in-the-loop steps.
Install @mastra/evals and call runEvals from Vitest or another ESM test runner with a dataset, a target resolved through mastra.getAgent, and quick checks such as calledTool, noToolErrors and maxToolCalls. Checks listed under gates must average 1.0 or the verdict is failed. Quick checks need no LLM, so they are fast and free to run on every commit.
Use the AI SDK alone when the feature is a chat interface with a few tools and the work ends when the stream ends. Use Mastra when you also need durable workflows that suspend for a person, built-in evals and a local Studio. Mastra depends on AI SDK provider packages and its Next.js guide streams through AI SDK UI, so the two can be combined.

Key Takeaway
Mastra is an Apache 2.0 TypeScript agent framework that reached 1.0 in January 2026. Use its agents for open-ended tool calling, its workflows for deterministic steps that can suspend and resume from storage, and its quick-check evals in Vitest to gate tool usage in CI. Keep approval rules in workflow code, not prompts.
Consider a procurement team on an ERP built with Next.js. Buyers raise purchase orders all day, and someone wants an assistant that reads a supplier's delivery history before a PO goes to a manager. The first prototype is usually one agent with a long prompt that says when to approve. It works in a demo, and then the auditor asks which line of code decides that a 48 million rupiah order skips the manager, and the honest answer is a paragraph of English the model reads differently on a bad day.
This Mastra TypeScript agent framework tutorial builds that feature the other way round with Mastra 1.x: a typed tool, an agent that only describes risk, a workflow that owns the approval rule and pauses for a human, evals that run in Vitest, and the local Studio. Every API name comes from the current Mastra docs and npm metadata, and the last section compares Mastra with the Vercel AI SDK and the OpenAI Agents SDK for a team that writes TypeScript end to end. Python-first options such as LangGraph and CrewAI are a separate discussion.
Mastra is an open-source TypeScript framework from the team behind Gatsby, a Y Combinator W25 company. Version 1.0.0 of @mastra/core was published on npm on 20 January 2026, and the founders' launch post reported more than 300,000 weekly npm downloads, 19.4k GitHub stars and an Apache 2.0 licence. The 1.x line has moved quickly since: @mastra/core is at 1.73.0 on npm at the time of writing. Four facts shape everything that follows:
The split between agents and workflows is the main design decision. The docs put it plainly: use an agent when the steps are not known in advance, and a workflow when the process is defined up front. A purchase-order approval is mostly the second kind, with one judgement call in the middle. Note also that @mastra/core declares Node.js 22.13.0 or later in its engines field and accepts zod 3.25 or zod 4 as a peer dependency.
Mastra can run as its own server, but for an ERP that already has a Next.js App Router codebase, importing it directly into route handlers is the shortest path. The CLI's init command adds a src/mastra folder with an example agent, a tool and an index file to the project you already have.
# Inside an existing Next.js App Router project.
# @mastra/core 1.x declares engines: node >= 22.13.0.
npx mastra@latest init # creates src/mastra/{index.ts,agents,tools}
npm install @mastra/libsql @mastra/evals zod
npm install -D vitest
# .env
# The model router reads the provider key from the model string:
# any 'openai/...' model id needs OPENAI_API_KEY, no provider import.
OPENAI_API_KEY=sk-...The Mastra instance is the registry. Agents, workflows and storage are registered once, and everything else looks them up by key so it shares the same logger, storage and tracing. Two lines in this file prevent most early confusion: the Next.js config flag, and an absolute storage path.
// next.config.ts
import type { NextConfig } from 'next';
const nextConfig: NextConfig = {
// Documented for Next.js on Vercel: keep Mastra out of the bundle.
serverExternalPackages: ['@mastra/*'],
};
export default nextConfig;
// src/mastra/index.ts
import { Mastra } from '@mastra/core';
import { LibSQLStore } from '@mastra/libsql';
import { procurementAgent } from './agents/procurement-agent';
import { purchaseOrderWorkflow } from './workflows/purchase-order';
export const mastra = new Mastra({
agents: { procurementAgent },
workflows: { purchaseOrderWorkflow },
// Suspended workflow snapshots and eval scores are written here.
// Absolute path on purpose: `next dev` and `mastra dev` start in
// different working directories, and a relative file: URL gives each
// process its own database. LibSQL needs a real filesystem, so this
// suits a VPS or container, not a serverless function.
storage: new LibSQLStore({ url: 'file:/srv/erp-app/data/mastra.db' }),
});LibSQLStore writes a local file, and Mastra's deployment docs say to remove it before deploying to a serverless platform. On a VPS or a Docker container with a mounted volume it works, but a suspended workflow is only as durable as that file. If the volume is not persisted, a redeploy silently discards every PO waiting for a manager. Move to a database-backed store, such as the @mastra/pg package, before real approvals depend on it.
A Mastra tool is created with createTool and has an id, a description, a Zod inputSchema, an optional outputSchema and an execute function. The docs are emphatic that execute has exactly one signature, inputData first and an execution context second, and that plain-object tool definitions silently fail to execute. The context carries requestContext, the tracing context and an abort signal.
// src/mastra/tools/supplier-tool.ts
import { createTool } from '@mastra/core/tools';
import { z } from 'zod';
import { db } from '../../lib/db';
export const getSupplierHistory = createTool({
id: 'get-supplier-history',
description:
'Look up one supplier by code, for example SUP-0193: open purchase ' +
'orders, late deliveries in the last 12 months and payment terms.',
inputSchema: z.object({
supplierCode: z.string().regex(/^SUP-\d{4}$/).describe('Supplier code'),
}),
outputSchema: z.object({
found: z.boolean(),
name: z.string().optional(),
openPoCount: z.number().optional(),
lateDeliveries12m: z.number().optional(),
paymentTermsDays: z.number().optional(),
}),
// Checked before execute() runs. On failure the tool RETURNS an error
// object to the model instead of throwing, so the run keeps going.
requestContextSchema: z.object({ tenantId: z.string() }),
// One signature only: execute(inputData, context).
execute: async ({ supplierCode }, { requestContext }) => {
// The tenant comes from the server session via RequestContext,
// never from an argument the model was free to invent.
const tenantId = requestContext?.get('tenantId');
const supplier = await db.supplier.findFirst({
where: { code: supplierCode, tenantId },
select: { name: true, openPoCount: true, lateDeliveries12m: true, paymentTermsDays: true },
});
// "Not found" is an answer the agent can explain, not an exception.
return supplier ? { found: true, ...supplier } : { found: false };
},
});RequestContext is how server-side facts reach a tool without passing through the model. The route handler sets tenantId from the session; the tool reads it. The model only ever chooses supplierCode, which the regex constrains to the ERP's own format. Declaring requestContextSchema on the tool adds a check before execute runs. On failure, an agent throws a MastraError before any model call, but a tool returns an error object the model can read, so test both paths rather than assuming one behaviour.
The agent itself is small. The instructions say what it may do and, more importantly, what it may not: it describes supplier risk and never approves anything. Tools are passed as an object, and that object's keys matter more than they look.
// src/mastra/agents/procurement-agent.ts
import { Agent } from '@mastra/core/agent';
import { getSupplierHistory } from '../tools/supplier-tool';
export const procurementAgent = new Agent({
id: 'procurement-agent',
name: 'Procurement Agent',
description: 'Assesses supplier risk for a purchase order',
instructions: `You assess supplier risk for purchase orders in an ERP.
Always call getSupplierHistory before giving an opinion.
You never approve or reject a purchase order. You only describe risk.`,
// provider/model string, resolved by Mastra's model router.
model: 'openai/gpt-5-mini',
// The OBJECT KEY is the toolName the model, the traces and the eval
// checks see. Here that is 'getSupplierHistory', not the tool's id.
tools: { getSupplierHistory },
// One step = one model call. A risk check needs a lookup and an answer.
defaultOptions: { maxSteps: 4 },
hooks: {
// Runs before every tool call from any source, including MCP tools.
// Return { proceed: false, output } here to block a call by policy.
beforeToolCall: ({ toolName, input }) => {
console.info('[procurement-agent] tool call', toolName, input);
},
},
});Calling agent.generate returns text, tool calls, tool results, steps and token usage; agent.stream exposes a textStream plus promises that resolve when the run ends. Passing structuredOutput with a Zod schema returns a validated object on response.object instead of prose, which is how the workflow below consumes this agent. Hooks give one place to log or block tool calls: returning proceed false from beforeToolCall skips the call and hands the agent your output as the tool result.
The tool name a model, a trace and an eval check sees is the key in the tools object, not the id passed to createTool. With tools set to getSupplierHistory, the stream reports getSupplierHistory, and checks.calledTool must use that same string. Pick one convention, keys equal to ids or keys as camelCase, and use it across every agent.
This is where the approval rule moves out of the prompt. The workflow runs a budget check in plain code and the supplier assessment through the agent in parallel, reshapes both results with map, then branches. Low risk, within budget and under 50 million rupiah is auto-approved; anything else suspends for a manager.
// src/mastra/workflows/purchase-order.ts
import { createStep, createWorkflow } from '@mastra/core/workflows';
import { z } from 'zod';
import { getRemainingBudget } from '../../lib/budget';
const AUTO_APPROVE_LIMIT_IDR = 50_000_000;
const PoInput = z.object({
poNumber: z.string(),
supplierCode: z.string(),
costCentre: z.string(),
amountIdr: z.number().positive(),
});
// Plain code: no model involved, so it gives the same answer every run.
const checkBudget = createStep({
id: 'check-budget',
inputSchema: PoInput,
outputSchema: z.object({ withinBudget: z.boolean() }),
execute: async ({ inputData, requestContext }) => {
const remaining = await getRemainingBudget(
requestContext.get('tenantId'),
inputData.costCentre,
);
return { withinBudget: inputData.amountIdr <= remaining };
},
});
const Risk = z.object({
risk: z.enum(['low', 'medium', 'high']),
reason: z.string(),
});
// The one step that asks a model, and it must answer in a schema.
const assessSupplier = createStep({
id: 'assess-supplier',
inputSchema: PoInput,
outputSchema: Risk,
retries: 2, // a provider blip retries this step, not the whole run
execute: async ({ inputData, mastra, requestContext }) => {
const agent = mastra.getAgent('procurementAgent');
const res = await agent.generate(
'Assess supplier ' + inputData.supplierCode + ' for ' + inputData.poNumber +
' worth IDR ' + inputData.amountIdr,
{ structuredOutput: { schema: Risk }, requestContext },
);
return res.object;
},
});
const Routing = z.object({
poNumber: z.string(),
amountIdr: z.number(),
withinBudget: z.boolean(),
risk: z.enum(['low', 'medium', 'high']),
reason: z.string(),
});
const Decision = z.object({
status: z.enum(['approved', 'rejected']),
decidedBy: z.string(),
});
const autoApprove = createStep({
id: 'auto-approve',
inputSchema: Routing,
outputSchema: Decision,
execute: async () => ({ status: 'approved' as const, decidedBy: 'policy:auto' }),
});
const managerApproval = createStep({
id: 'manager-approval',
inputSchema: Routing,
outputSchema: Decision,
suspendSchema: z.object({ poNumber: z.string(), reason: z.string() }),
resumeSchema: z.object({ approved: z.boolean(), managerId: z.string() }),
execute: async ({ inputData, resumeData, suspend }) => {
if (!resumeData) {
// A snapshot goes to storage. The process may restart, or be
// redeployed, before the manager clicks anything.
return await suspend({ poNumber: inputData.poNumber, reason: inputData.reason });
}
return {
status: resumeData.approved ? ('approved' as const) : ('rejected' as const),
decidedBy: resumeData.managerId,
};
},
});
const canAutoApprove = (r: z.infer<typeof Routing>) =>
r.withinBudget && r.risk === 'low' && r.amountIdr <= AUTO_APPROVE_LIMIT_IDR;
export const purchaseOrderWorkflow = createWorkflow({
id: 'purchase-order-approval',
inputSchema: PoInput,
// A branch's output is keyed by the step id that actually ran.
outputSchema: z.object({
'auto-approve': Decision.optional(),
'manager-approval': Decision.optional(),
}),
})
.parallel([checkBudget, assessSupplier])
// Parallel output is keyed by step id; reshape it for the branch.
.map(async ({ inputData, getInitData }) => {
const po = getInitData<z.infer<typeof PoInput>>();
return {
poNumber: po.poNumber,
amountIdr: po.amountIdr,
withinBudget: inputData['check-budget'].withinBudget,
risk: inputData['assess-supplier'].risk,
reason: inputData['assess-supplier'].reason,
};
})
// Conditions are evaluated in order. The model's opinion is one input
// to the rule; the rule itself is code you can read in review.
.branch([
[async ({ inputData }) => canAutoApprove(inputData), autoApprove],
[async ({ inputData }) => !canAutoApprove(inputData), managerApproval],
])
.commit();Three details in this file follow directly from the docs. Parallel and branch outputs are objects keyed by the id of the step that ran, which is why map reads check-budget and assess-supplier, and why the workflow's outputSchema lists both branch steps as optional. Every step in a branch must share the same input and output schemas. Branch conditions are evaluated in order and only the first match runs. The 50 million threshold is an example value for illustration, not a recommendation.
Suspension is what makes this a workflow rather than a long-running request. The manager-approval step calls suspend, the run returns with status suspended, and the snapshot is written to storage. Hours later a different HTTP request recreates the run from its runId and calls resume with data that must match resumeSchema. The result is a discriminated union on status, with success, failed, suspended, tripwire and paused as the possible values, so each route handles each case explicitly.
// src/app/api/po/route.ts (a buyer submits a purchase order)
import { RequestContext } from '@mastra/core/request-context';
import { mastra } from '@/mastra';
import { getSession } from '@/lib/auth';
import { savePoRunId } from '@/lib/po';
export const runtime = 'nodejs';
export async function POST(req: Request) {
const session = await getSession(req);
if (!session) return new Response('Unauthorized', { status: 401 });
const requestContext = new RequestContext<{ tenantId: string }>();
requestContext.set('tenantId', session.tenantId);
const po = await req.json();
const runId = crypto.randomUUID();
const run = await mastra.getWorkflow('purchaseOrderWorkflow').createRun({ runId });
const result = await run.start({ inputData: po, requestContext });
if (result.status === 'suspended') {
// Keep the runId on the PO row. The approve button needs it later.
await savePoRunId(po.poNumber, runId);
return Response.json({ state: 'awaiting-manager', poNumber: po.poNumber });
}
if (result.status === 'success') return Response.json({ state: 'decided', ...result.result });
return Response.json({ state: result.status }, { status: 500 });
}
// src/app/api/po/approve/route.ts (the manager answers, hours later)
export async function POST(req: Request) {
const session = await getSession(req);
if (!session?.roles.includes('po-approver')) return new Response('Forbidden', { status: 403 });
const { runId, approved } = await req.json();
const requestContext = new RequestContext<{ tenantId: string }>();
requestContext.set('tenantId', session.tenantId);
const run = await mastra.getWorkflow('purchaseOrderWorkflow').createRun({ runId });
const result = await run.resume({
step: 'manager-approval',
// managerId comes from the session, not the request body.
resumeData: { approved, managerId: session.userId },
requestContext,
});
return Response.json({ state: result.status });
}Mastra's scorers live in the @mastra/evals package. Quick checks are deterministic micro-scorers that need no model: includes, excludes, matches, calledTool, didNotCall, toolOrder, maxToolCalls, usedNoTools and noToolErrors. runEvals pushes a dataset through an agent or workflow, and anything listed under gates must average 1.0 or the verdict is failed, which is exactly the binary signal a CI job wants.
// src/mastra/procurement-agent.eval.test.ts
import { describe, it, expect } from 'vitest';
import { runEvals } from '@mastra/core/evals';
import { checks } from '@mastra/evals/checks';
import { RequestContext } from '@mastra/core/request-context';
import { mastra } from './index';
const tenant = new RequestContext<{ tenantId: string }>();
tenant.set('tenantId', 'tenant-eval'); // seeded test tenant, not production
describe('procurement agent', () => {
it('looks the supplier up before judging risk', async () => {
const result = await runEvals({
// From the registry, not a direct import: a bare import has no
// Mastra instance, so scores are not persisted and registry
// lookups inside steps and tools fail.
target: mastra.getAgent('procurementAgent'),
data: [
{ input: 'Assess supplier SUP-0193 for PO-2026-0815 worth IDR 42,000,000.', requestContext: tenant },
{ input: 'Is SUP-0007 risky for a rush order of packaging film?', requestContext: tenant },
{ input: 'Assess SUP-9999 for PO-2026-0816.', requestContext: tenant }, // unknown supplier
],
// Gates must average 1.0 or the verdict is 'failed'. No LLM, no cost.
gates: [
// The tools-object KEY, not the createTool id.
checks.calledTool('getSupplierHistory'),
checks.noToolErrors(),
checks.maxToolCalls(3),
],
});
expect(result.verdict).toBe('passed');
expect(result.summary.totalItems).toBe(3);
});
});Two traps are documented and easy to miss. Resolve the target through mastra.getAgent rather than importing the agent, otherwise scores are not persisted and registry lookups inside steps fail. And every calledTool instance shares the id check-called-tool, so two calledTool checks in one run collide in the scores map; use toolOrder when a test cares about more than one tool. For quality judgements that code cannot make, model-graded scorers such as answer relevancy and faithfulness plug into the same array, and live scorers can be attached to an agent with a sampling rate so a fraction of production traffic is scored in the background.
Running mastra dev starts Studio on localhost port 4111, with the REST API documented at the swagger-ui path on the same port. It is the fastest way to see what the agent and the workflow actually do before wiring them into the ERP UI. A practical order for the PO feature:
Studio runs as a separate process from next dev, which is why the storage URL in the index file is absolute: with a relative path, each process resolves the file from its own working directory, and a run started in the Next.js app never appears in Studio. Studio can also be deployed with authentication for a team, but for a single-developer ERP project the local instance is usually enough.
All three are TypeScript-native and all three use Zod for tool schemas, so the choice is about how much framework you want around the model loop. Mastra itself depends on AI SDK provider packages, and its Next.js guide streams through AI SDK UI, so the first two are layers more than rivals.
| Concern | Mastra 1.x | Vercel AI SDK 7 | OpenAI Agents SDK (TypeScript) |
|---|---|---|---|
| Agent loop | Agent class with generate and stream, maxSteps and stopWhen, subagents exposed as tools | ToolLoopAgent class with stopWhen and prepareStep for loop control | Agent plus run, with handoffs between agents and guardrails |
| Deterministic workflows | Built-in engine: then, parallel, branch, foreach, suspend and resume persisted to storage | Documented as patterns in plain TypeScript: sequential, parallel, routing, orchestrator-worker | Orchestration in code or through handoffs; human-in-the-loop and sessions are documented |
| Evals | @mastra/evals scorers, quick checks, runEvals with gates, live sampled scoring | Not part of the agent docs cited here; pair with an external eval tool | Testing guide and built-in tracing; pair with an external eval tool for scoring |
| Local dev UI | Studio on port 4111 with chat, workflow graph, traces and time travel | Not covered in the cited docs; the UI layer is AI SDK UI for your own app | Not covered in the cited docs; traces go to the configured tracing exporter |
| Models | provider slash model strings through the model router, 40-plus providers per the README | Provider packages or model strings, many providers | OpenAI models first, other providers through the documented AI SDK integration |
| Package and licence | @mastra/core 1.73.0, Apache 2.0 outside ee directories, Node.js 22.13 or later | ai 7.0.x on npm, Apache 2.0 | @openai/agents 0.18.0, MIT, zod 4 required as a peer dependency |
A reasonable rule: choose the AI SDK alone when the feature is a chat box with a few tools and the request ends when the stream ends. Choose the OpenAI Agents SDK when the product is built around OpenAI models and multi-agent handoffs. Choose Mastra when the feature needs a process that outlives a request, such as an approval that waits for a person, together with evals and a local UI from the same package. The cost is a larger dependency surface and a fast release cadence, so pin versions and read the changelog before upgrading.
Packages used in this tutorial: @mastra/core for agents, tools, workflows and runEvals; @mastra/libsql for local storage; @mastra/evals for quick checks and scorers; and mastra, the CLI that runs init and dev. For a chat UI in Next.js, the Mastra guide adds @mastra/ai-sdk with the AI SDK useChat hook.
The rule to carry into any agent feature in an ERP: let the agent describe and let code decide. Mastra makes that split cheap in TypeScript, with Zod tools that take tenant data from RequestContext, a workflow whose branch condition is a function a reviewer can read, a suspend step that waits as long as the manager does, and a Vitest gate that fails the build when the agent stops calling its tool. If a prompt is the only place a business rule is written down, that rule belongs in a workflow step.
Sources