AI
AI Agent Evaluation: How to Test Tool Calls and Trajectories
October 202612 min read

Trajectory evaluation grades the path an agent took rather than only its final answer. It checks which tools were called, the arguments passed, the order of calls, whether the goal was reached and how many steps the run cost. This catches failures such as acting on the wrong record that a fluent final message would hide.
An agent can produce a correct-sounding reply after calling the wrong tool, using the wrong customer id or writing before it read the record. A final-answer check only compares the last message, so it passes those runs. Grading the trace moves the check to where the decision was actually made.
promptfoo offers trajectory:tool-used, trajectory:tool-args-match, trajectory:tool-sequence, trajectory:step-count and the model-graded trajectory:goal-success. tool-args-match supports partial and exact modes, and tool-sequence supports in_order and exact modes. All of them read OpenTelemetry trace data, so tracing must be enabled for the eval.
Call the NestJS endpoint with promptfoo's HTTP provider and enable tracing with its built-in OTLP receiver. promptfoo sends a traceparent header, so OpenTelemetry HTTP instrumentation in NestJS joins the same trace. Wrap every tool execution in a span with tool.name and tool.arguments attributes so the trajectory assertions can see it.
Google ADK uses the tool_trajectory_avg_score criterion with EXACT, IN_ORDER or ANY_ORDER match types. Each invocation scores 1.0 for a match or 0.0 for a mismatch, and the case score is the average, with a default threshold of 1.0. The criterion is not supported when the eval uses ADK's user simulation.

Key Takeaway
AI agent evaluation has to grade the trajectory, not only the final answer: which tools the agent chose, the arguments it passed, the order of calls, whether the goal was met and how many steps it cost. Emit tool spans over OpenTelemetry, turn production traces into test cases, and assert on them with promptfoo trajectory checks in CI.
Consider an ERP helpdesk agent that handles credit note requests. A customer service officer types that PT Sinar Jaya wants a credit note for two damaged units on an invoice. The agent replies with a polite, correct-looking summary and a draft number. A final-answer eval scores it as a pass. The trace tells a different story: the agent searched for the customer by a partial name, picked the wrong PT Sinar Jaya out of three, and created the draft against an invoice it never fetched. The answer read well because the model is good at writing; the work underneath was wrong.
That gap is why AI agent evaluation needs a different unit than prompt evaluation. An existing post on this site covers LLM evals in CI for single prompts: a golden dataset, assertions and a pass-rate gate. This one goes one level down, to the multi-step path an agent takes. It covers what to grade in a trajectory, how to build the dataset from production traces, how to make a NestJS agent emit the spans an eval can read, and how to write promptfoo trajectory assertions against them, with OpenAI trace grading and Google ADK evaluation compared at the end. Every assertion name and option below comes from the vendors' current documentation.
A final-answer check compares one string to an expectation. An agent run is a sequence of decisions, and most of the expensive failures happen in the middle of it, where a fluent final message hides them. The table maps common agent failures to what each style of check can see.
| Failure in the run | What a final-answer check sees | What a trajectory check catches |
|---|---|---|
| Right tool, wrong record: the agent looked up the wrong customer | A plausible reply with a customer name in it, usually a pass | tool-args-match fails because the customer id or invoice number differs |
| Write before read: a draft created without fetching the invoice | Nothing, the reply still quotes a draft number | tool-sequence fails because get_invoice does not precede the write |
| Forbidden action: the agent posted the document instead of drafting it | Often nothing, or a reply that sounds even more helpful | A negated tool-used check fails the moment the posting tool appears |
| Looping: the same search repeated eight times before answering | A correct answer, only slower and dearer | step-count with a max fails, and a token ceiling catches the cost |
| Hallucinated argument: an extra flag such as force: true | Nothing at all | tool-args-match in exact mode rejects the unexpected key |
The pattern is that the final message is the least informative artefact of an agent run. It is written last, by the component most skilled at sounding right. A trajectory check moves the assertion to where the decision was made, which is also where the fix will go.
Five properties cover nearly every agent failure worth blocking a merge for. Each maps to a different kind of check, and keeping them separate tells you which part of the agent regressed instead of handing you one blended score.
Grade the deterministic properties with deterministic checks and reserve the model grader for goal success. A model grader asked whether the tool calls were sensible will be inconsistent where a string comparison would have been exact and free.
The best trajectory test cases are runs that already went wrong. Invented edge cases test what you imagined; traces test what users actually typed, including the ambiguous company names and the half-remembered invoice numbers that send agents down the wrong path.
A case derived from the opening scenario looks like this. The assertions describe the path, not the prose, so the test survives any rewording of the agent's reply.
# evals/helpdesk/credit-note-ambiguous-customer.yaml
# Derived from a production trace. Names and numbers are fictional, but the
# shape that broke the agent is kept: three customers share "Sinar Jaya".
- description: 'Credit note, ambiguous customer name, partial damage'
vars:
query: >-
PT Sinar Jaya Abadi wants a credit note for 2 damaged units
on INV-2026-08812
assert:
# The disambiguating lookup must carry the full legal name, not "Sinar Jaya".
- type: trajectory:tool-args-match
value:
name: find_customer
args:
legal_name: 'PT Sinar Jaya Abadi'
# Read the invoice before drafting against it. in_order (the default)
# tolerates extra calls in between, e.g. a stock-return lookup.
- type: trajectory:tool-sequence
value:
steps:
- find_customer
- get_invoice
- create_credit_note_draft
# On the write tool, exact mode rejects invented extras such as force: true.
# ignore drops the key the agent generates fresh on every call.
- type: trajectory:tool-args-match
value:
name: create_credit_note_draft
mode: exact
args:
invoice_no: 'INV-2026-08812'
lines:
- line_no: 1
qty: 2
reason: damaged
ignore:
- idempotency_keypromptfoo reads trajectories from OpenTelemetry spans. Its tracing documentation says it includes a traceparent header when it calls an HTTP target, and the application behind the target can add child spans under that trace. It recognises tool calls from span attributes such as tool.name, and arguments from attributes such as tool.arguments, parsing string values as JSON when possible. So a NestJS-hosted agent needs two things: OpenTelemetry started before the app, with HTTP instrumentation to pick up the incoming traceparent, and a span around every tool execution.
The tracer file goes first in main.ts. The wrapper is the only place tool calls are executed, which also gives you one place to enforce timeouts and audit logging later.
// src/tracing.ts — import this on the FIRST line of main.ts, before NestFactory,
// or the HTTP module is loaded un-instrumented and the traceparent is ignored.
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import {
BatchSpanProcessor,
SimpleSpanProcessor,
} from '@opentelemetry/sdk-trace-base';
import { HttpInstrumentation } from '@opentelemetry/instrumentation-http';
import { resourceFromAttributes } from '@opentelemetry/resources';
const exporter = new OTLPTraceExporter({
// promptfoo's built-in receiver listens on 127.0.0.1:4318 by default.
url: process.env.OTEL_TRACES_URL ?? 'http://127.0.0.1:4318/v1/traces',
});
const sdk = new NodeSDK({
resource: resourceFromAttributes({ 'service.name': 'erp-helpdesk-agent' }),
// Under eval, flush every span immediately so it lands before the response.
spanProcessors: [
process.env.AGENT_EVAL_MODE === '1'
? new SimpleSpanProcessor(exporter)
: new BatchSpanProcessor(exporter),
],
// Extracts the incoming W3C traceparent, so our spans join promptfoo's trace.
instrumentations: [new HttpInstrumentation()],
});
sdk.start();// src/agent/traced-tool.ts — every tool call in the agent loop goes through here.
import { trace, SpanStatusCode } from '@opentelemetry/api';
const tracer = trace.getTracer('erp-helpdesk-agent');
export function tracedTool<T>(
name: string,
args: Record<string, unknown>,
run: () => Promise<T>,
): Promise<T> {
// startActiveSpan parents this span under the request span that
// HttpInstrumentation created from promptfoo's traceparent.
return tracer.startActiveSpan(`execute_tool ${name}`, async (span) => {
// These two keys are what trajectory:tool-used and tool-args-match read.
span.setAttribute('tool.name', name);
span.setAttribute('tool.arguments', JSON.stringify(args));
try {
const result = await run();
// Truncate: an invoice with 400 lines should not become a 2 MB attribute.
span.setAttribute('tool.output', JSON.stringify(result).slice(0, 2000));
return result;
} catch (err) {
span.recordException(err as Error);
span.setStatus({ code: SpanStatusCode.ERROR });
throw err;
} finally {
span.end();
}
});
}
// Usage inside the agent loop, for each tool call the model requests:
// const invoice = await tracedTool('get_invoice', call.args, () =>
// this.erp.getInvoice(call.args.invoice_no as string),
// );Use a simple span processor when exporting to the eval receiver. A batching processor holds spans in memory and flushes them later, so the HTTP response can reach promptfoo before the tool spans do, and an otherwise correct run fails every trajectory assertion. In production the batching processor is still the right choice.
Tool arguments often carry customer names and amounts. promptfoo's receiver accepts a redactAttributes list that replaces matching attribute values before storage, and a retentionDays setting that prunes old traces. Its documentation warns that short patterns over-match, since the key token also matches gen_ai.usage.input_tokens, so name the exact attribute keys you want hidden.
With spans flowing, the config points an HTTP provider at the NestJS endpoint, enables tracing with the built-in OTLP receiver on its default port 4318, and asserts on the path. The assertion family is trajectory:tool-used, trajectory:tool-args-match, trajectory:tool-sequence, trajectory:step-count and the model-graded trajectory:goal-success.
# promptfooconfig.yaml
description: ERP helpdesk agent — trajectory evals
tracing:
enabled: true
failOnReceiverStartFailure: true # no receiver means no trajectories: fail loudly
otlp:
http:
redactAttributes: ['authorization', 'customer_npwp']
storage:
type: sqlite
retentionDays: 14
prompts:
- '{{query}}'
providers:
- id: http
config:
url: http://localhost:3000/agent/run
method: POST
headers:
Content-Type: application/json
body:
message: '{{query}}'
transformResponse: json.reply
defaultTest:
assert:
# Policy for every request: the agent drafts, a human posts.
- type: not-trajectory:tool-used
value: post_credit_note
# Loop guard: more than 12 tool steps is a regression even if the answer is right.
- type: trajectory:step-count
value:
type: tool
max: 12
tests:
- file://evals/helpdesk/*.yamlA few options do most of the work. tool-used accepts a name, a list of names, or a glob pattern with min and max counts. tool-args-match defaults to partial mode, where expected properties are matched recursively as a subset; exact mode rejects any extra argument, and the defaults and ignore lists let exact mode tolerate pagination defaults and volatile ids such as an idempotency key. tool-sequence defaults to in_order, which allows other calls between the expected steps, while exact requires the traced sequence to match step for step. Every promptfoo assertion can be negated with a not- prefix, which is how a forbidden tool is expressed.
Put the rules that hold for every request in defaultTest rather than repeating them per case. No posting tool, no write before the matching read and a step ceiling are policies, not test cases, and declaring them once means a new case cannot forget them.
trajectory:goal-success is model-graded and takes an explicit provider in the documented examples. Give it a goal written as an outcome the grader can verify, such as which customer and invoice the draft must reference, rather than a vague instruction to be helpful. Because it is a model judgement, treat a failure here as a prompt to read the trace, and let the deterministic checks be the ones that block a merge.
Rules the built-in assertions cannot express go in a javascript assertion, which receives the trace as context.trace.spans. The same mechanism handles the cost dimension: sum token usage across spans and fail the run above a ceiling. The ceiling below is an example to calibrate from your own baseline runs, not a recommendation.
# Appended to defaultTest.assert in promptfooconfig.yaml
# Goal success: model-graded, so it reports; the deterministic checks block.
- type: trajectory:goal-success
value: >-
Create a credit note DRAFT for the customer and invoice named in the
request, for the damaged quantity only, and tell the user the draft
number without claiming it has been posted.
provider: openai:gpt-6-luna
metric: goal_success
# Business rule for every case: no draft without having fetched the
# invoice it credits. Per-case tool-sequence checks then pin the order.
- type: javascript
value: |
const names = context.trace.spans
.filter((s) => s.attributes['tool.name'])
.map((s) => s.attributes['tool.name']);
if (!names.includes('create_credit_note_draft')) return true;
return names.includes('get_invoice');
# Cost ceiling from our own GenAI spans. 40k is an example: set it from
# the p95 of your baseline runs, then tighten.
- type: javascript
value: |
const used = context.trace.spans.reduce((total, s) =>
total
+ Number(s.attributes['gen_ai.usage.input_tokens'] ?? 0)
+ Number(s.attributes['gen_ai.usage.output_tokens'] ?? 0), 0);
return used <= 40000;promptfoo also ships a cost assertion, but its documentation notes it needs the provider to return cost information, which a plain HTTP target usually does not. Counting tokens from your own spans works with any backend, and the step-count check gives a second, model-independent signal for loops.
promptfoo is not the only way to grade trajectories. OpenAI and Google both offer trajectory-aware evaluation, each tied to its own agent stack. The differences matter more than the feature lists suggest.
| Option | Where checks live | Trajectory matching | Best fit |
|---|---|---|---|
| promptfoo trajectory assertions | YAML in the repository, run from the CLI in CI | tool-used, tool-args-match in partial or exact mode, tool-sequence in in_order or exact mode, step-count, model-graded goal-success | Any stack that emits OpenTelemetry spans, including a custom NestJS agent |
| OpenAI trace grading | The OpenAI dashboard under Logs, Traces, then Grade all into the evaluation dashboard | Graders applied to traces from Agents SDK apps; run options include model, date range and tool calls | Teams already tracing through the OpenAI Agents SDK who want graded traces without new tooling |
| Google ADK evaluation | test.json or evalset.json files, run with adk eval, pytest or the adk web UI | tool_trajectory_avg_score with EXACT, IN_ORDER or ANY_ORDER; each invocation scores 1.0 or 0.0 and the case averages them, default threshold 1.0 | Agents built on ADK, especially multi-turn flows that use its user simulation |
Two caveats change the choice. OpenAI's graders documentation states that it is deprecating graders as part of the evals and fine-tuning workflows they support, and points to its deprecations page for the timeline, so a dashboard-only eval suite is a migration waiting to happen. ADK's documentation notes that tool_trajectory_avg_score, response_match_score and final_response_match_v2 are not supported with user simulation, so simulated multi-turn runs fall back to rubric-based criteria such as rubric_based_tool_use_quality_v1. Checks that live in your repository, against spans you emit yourself, are the ones that survive a vendor change.
Do not run the eval against live write tools. Point the NestJS agent at a staging ERP or a stubbed tool layer, because a trajectory eval that creates real credit notes is a production incident with a test report attached.
The rule to carry is that an agent is graded by what it did, not by what it said. Instrument every tool call as a span, turn the runs that went wrong into frozen trajectories, check selection, arguments and order deterministically, keep the model grader for goal success, and put a ceiling on steps and tokens. The final answer still matters, but it is the last thing to check, not the only one.
Sources