DevOps
OpenTelemetry GenAI Semantic Conventions: Agent Spans in Tempo
October 202613 min read

They are the OpenTelemetry naming rules for spans, metrics and events produced by generative AI applications, covering model calls, agents, tools and MCP. They now live in the open-telemetry/semantic-conventions-genai repository. Every document there is still in Development status, so attribute names can change between releases.
invoke_agent represents one agent invocation, with INTERNAL kind for an in-process agent and CLIENT kind for a remote hosted agent. invoke_workflow represents a user-facing run that coordinates several agents or GenAI calls, such as an OpenAI Agents run with handoffs. The spec says a standalone agent invocation should not be reported as a workflow.
No. The model call ends when it returns a tool-call request, and the tool runs afterwards. The conventions describe tool spans as siblings under the same invoke_agent span. Nesting them under the chat span inflates the chat span's duration and hides tool latency inside model latency.
Register a TracingProcessor that converts SDK spans to OpenTelemetry spans and exports them over OTLP to a collector such as Alloy that forwards to Tempo. add_trace_processor keeps the default upload to OpenAI's trace dashboard, while set_trace_processors replaces it. The official opentelemetry-instrumentation-genai-openai-agents package is an alternative to writing the processor yourself.
No. Attributes such as gen_ai.input.messages, gen_ai.output.messages and gen_ai.tool.call.arguments are Opt-In, and the conventions say instrumentations should not capture them by default. For production the spec recommends uploading content to external storage and recording a reference on the span. If you do record content on spans in Tempo, its default max_attribute_bytes of 2048 truncates longer values.

Key Takeaway
The OpenTelemetry GenAI semantic conventions model an agent run as an invoke_workflow or invoke_agent span with chat and execute_tool spans as siblings beneath it. A custom OpenAI Agents SDK TracingProcessor can emit that shape to Grafana Tempo, but message content is opt-in, and Tempo truncates attribute values over 2048 bytes by default.
Consider an ERP helpdesk where an invoice-dispute agent hands a ticket from a triage agent to an accounts-payable matcher, which then calls three tools against the purchasing module. A user reports that the agent took forty seconds and gave the wrong answer. The service traces in Tempo show one long HTTP request and a burst of database queries, and nothing else: no agent, no model call, no tool, so there is no way to tell whether the time went to the model, to a looping tool or to a handoff that should never have happened.
The OpenTelemetry GenAI semantic conventions exist to fill that gap. This post walks through what the agent span conventions define today, the span tree an agent run should produce, a complete Python trace processor that maps OpenAI Agents SDK spans onto those conventions and exports them over OTLP, how to handle message capture without leaking ERP data, and the TraceQL queries that make the result useful. Every attribute name and requirement level below comes from the conventions repository itself, which still marks all of it as Development.
The GenAI conventions now live in their own repository, open-telemetry/semantic-conventions-genai, and the old page on opentelemetry.io points there with a notice that it is no longer maintained. Every document in it carries Development status, which means attribute names can still change between releases. Within that, the agent document defines six operations, each identified by gen_ai.operation.name and a span name built from it.
| gen_ai.operation.name | Span name | Kind | When to emit it |
|---|---|---|---|
invoke_workflow | invoke_workflow {gen_ai.workflow.name} | INTERNAL | A user-facing entry point that coordinates several agents or GenAI calls, such as a run with handoffs. Not for a standalone agent. |
invoke_agent | invoke_agent {gen_ai.agent.name} | INTERNAL / CLIENT | One agent invocation. INTERNAL when the agent runs in your process, CLIENT when it is a remote hosted agent. |
create_agent | create_agent {gen_ai.agent.name} | CLIENT | Creating a hosted agent resource on a provider, such as a Bedrock agent. |
plan | plan {gen_ai.agent.name} | INTERNAL | Only when the instrumentation can reliably tell planning apart from ordinary inference. |
chat | chat {gen_ai.request.model} | CLIENT | Each model call. Defined in the model spans document, which the agent document extends. |
execute_tool | execute_tool {gen_ai.tool.name} | INTERNAL | Each tool execution. Agent Skills loads and reads are refinements of this span, not separate spans. |
The requirement levels matter more than the names. On every one of these spans only gen_ai.operation.name is Required, plus gen_ai.provider.name on spans that call a provider. Agent name, description, version and conversation id are Conditionally Required, meaning you set them when you have them. Token counts and request parameters are Recommended. Everything that carries content, from gen_ai.input.messages to gen_ai.tool.call.arguments, is Opt-In, and the model spans document says instrumentations should not capture it by default.
For the invoice-dispute example, the conventions point to a tree like the one below. The root is invoke_workflow because the run involves a handoff, and the spec lists an OpenAI Agents Runner.run with handoffs, sub-agents or agents-as-tools as its example of a workflow. A single agent with no delegation would use invoke_agent as the root instead.
invoke_workflow invoice_dispute INTERNAL gen_ai.conversation.id=TCK-2026-10-0193
├─ invoke_agent Triage INTERNAL gen_ai.agent.name=Triage
│ ├─ chat <model> CLIENT gen_ai.usage.input_tokens=1840
│ └─ agent.handoff INTERNAL (no GenAI operation exists for this)
└─ invoke_agent AP Matcher INTERNAL gen_ai.agent.name=AP Matcher
├─ chat <model> CLIENT finish: tool call
├─ execute_tool get_purchase_order INTERNAL sibling of the chat span, not its child
├─ execute_tool get_goods_receipt INTERNAL
└─ chat <model> CLIENT final answer
# Wrong: execute_tool nested under the chat span that requested it.
# The model call has already returned when the tool runs; nesting it there
# inflates the chat span's duration and hides tool latency inside model latency.The structural rule people most often get wrong is where the tool spans go. A model call returns a tool-call request and ends; the framework then runs the tool, and the next model call starts after it. The plan span description in the conventions makes the intended shape explicit: tool or task spans are typically sibling operations under the same invoke_agent span. Nesting execute_tool under the chat span that asked for it makes the chat span appear to last as long as the tool, which corrupts every latency dashboard built on chat spans.
The handoff has no GenAI operation at all. The well-known values cover chat, embeddings, retrieval, memory operations, create_agent, invoke_agent, invoke_workflow, plan and execute_tool, but nothing for transferring control between agents. The second invoke_agent span already shows that a transfer happened; the handoff span is kept as a plain INTERNAL span so that the moment of transfer stays visible, without pretending it belongs to a convention.
The OpenAI Agents SDK traces every run by default and hands each trace and span to registered processors through four callbacks: on_trace_start, on_trace_end, on_span_start and on_span_end. Each span carries a span_id, a parent_id, timestamps, an optional error, and a typed span_data object. The mapping is mostly one to one, with two exceptions worth deciding up front.
| SDK object | GenAI operation | Notes |
|---|---|---|
Trace | invoke_workflow | The trace name is RunConfig.workflow_name, which defaults to the literal string Agent workflow, so set it. group_id becomes gen_ai.conversation.id. |
AgentSpanData | invoke_agent | AgentSpanData has name, handoffs, tools and output_type. Only the name maps to a convention attribute. |
ResponseSpanData / GenerationSpanData | chat | The default Responses API model produces ResponseSpanData with the Response object and usage; Chat Completions models produce GenerationSpanData with model and usage. |
FunctionSpanData | execute_tool | FunctionSpanData holds the tool name plus input and output, which are empty when sensitive data capture is off. |
HandoffSpanData / GuardrailSpanData | None defined | Emitted as plain spans with vendor-prefixed attributes, so a GenAI-aware backend does not misread them. |
TaskSpanData / TurnSpanData | Skipped | Recent SDK versions add task and turn spans around each run and each model turn. They have no convention, so their children are re-parented onto the nearest mapped ancestor. |
Skipping task and turn spans is a judgement call, not a requirement. Keeping them would add two levels between the agent and its model calls, and TraceQL structural queries such as agent span followed directly by tool span would stop matching. If the per-turn grouping matters to you, emit them as plain INTERNAL spans instead and write your queries with the descendant operator rather than the child operator.
The processor below subclasses TracingProcessor, keeps a dictionary from SDK ids to OpenTelemetry spans, and starts each OpenTelemetry span with an explicit parent context taken from that dictionary. That last part is the point: the SDK tracks its current span in a Python contextvar of its own, so relying on the OpenTelemetry current context would attach spans to whatever HTTP request span happened to be active, or to nothing.
# genai_processor.py
# pip install openai-agents opentelemetry-sdk opentelemetry-exporter-otlp-proto-http
import threading
from typing import Any
from agents.tracing import (
AgentSpanData, FunctionSpanData, GenerationSpanData, GuardrailSpanData,
HandoffSpanData, ResponseSpanData, Span, Trace, TracingProcessor,
)
from opentelemetry import trace as otel_trace
from opentelemetry.trace import SpanKind, Status, StatusCode
class GenAISemconvProcessor(TracingProcessor):
"""Re-emit Agents SDK traces as OpenTelemetry GenAI spans (semconv status: Development)."""
def __init__(self, tracer_provider, capture_content: bool = False):
self._provider = tracer_provider
self._tracer = otel_trace.get_tracer("erp.agents.genai", tracer_provider=tracer_provider)
self._capture = capture_content
self._otel: dict[str, otel_trace.Span] = {} # SDK trace_id / span_id -> OTel span
self._alias: dict[str, str] = {} # skipped SDK span -> nearest mapped ancestor
self._agent: dict[str, str] = {} # mapped key -> agent name, for tool spans
self._lock = threading.Lock() # callbacks arrive from concurrent runs
# -- trace -> invoke_workflow root ---------------------------------------------
def on_trace_start(self, trace: Trace) -> None:
attrs = {"gen_ai.operation.name": "invoke_workflow", "gen_ai.workflow.name": trace.name}
group_id = getattr(trace, "group_id", None)
if group_id: # a real ticket/session id only -- the spec forbids inventing one
attrs["gen_ai.conversation.id"] = group_id
root = self._tracer.start_span(
f"invoke_workflow {trace.name}", kind=SpanKind.INTERNAL, attributes=attrs)
with self._lock:
self._otel[trace.trace_id] = root
def on_trace_end(self, trace: Trace) -> None:
with self._lock:
root = self._otel.pop(trace.trace_id, None)
if root is not None:
root.end()
# -- spans ---------------------------------------------------------------------
def on_span_start(self, span: Span[Any]) -> None:
data = span.span_data
with self._lock:
parent_key = span.parent_id or span.trace_id
parent_key = self._alias.get(parent_key, parent_key)
described = self._describe(data, self._agent.get(parent_key))
if described is None:
# task/turn spans: re-parent their children onto the nearest mapped
# ancestor, so Tempo shows agent > chat, not agent > turn > chat.
self._alias[span.span_id] = parent_key
return
name, kind, attrs = described
parent = self._otel.get(parent_key)
ctx = otel_trace.set_span_in_context(parent) if parent is not None else None
self._otel[span.span_id] = self._tracer.start_span(
name, context=ctx, kind=kind, attributes=attrs) # sampling sees these
self._agent[span.span_id] = (
data.name if isinstance(data, AgentSpanData) else self._agent.get(parent_key, ""))
def _describe(self, data, agent_name):
if isinstance(data, AgentSpanData):
return (f"invoke_agent {data.name}", SpanKind.INTERNAL,
{"gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": data.name})
if isinstance(data, (ResponseSpanData, GenerationSpanData)):
api = "responses" if isinstance(data, ResponseSpanData) else "chat_completions"
return ("chat", SpanKind.CLIENT, # renamed in on_span_end once the model is known
{"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai",
"openai.api.type": api})
if isinstance(data, FunctionSpanData):
attrs = {"gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": data.name,
"gen_ai.tool.type": "function"}
if agent_name:
attrs["gen_ai.agent.name"] = agent_name
return (f"execute_tool {data.name}", SpanKind.INTERNAL, attrs)
if isinstance(data, HandoffSpanData):
return ("agent.handoff", SpanKind.INTERNAL, {}) # no GenAI operation: no gen_ai.*
if isinstance(data, GuardrailSpanData):
return (f"guardrail {data.name}", SpanKind.INTERNAL, {})
return None
def on_span_end(self, span: Span[Any]) -> None:
with self._lock:
self._alias.pop(span.span_id, None)
self._agent.pop(span.span_id, None)
out = self._otel.pop(span.span_id, None)
if out is None:
return
data = span.span_data
if isinstance(data, ResponseSpanData):
if data.response is not None:
# The response echoes the model that answered, so this is response.model;
# the request-side name never reaches this span type.
out.update_name(f"chat {data.response.model}")
out.set_attribute("gen_ai.response.model", data.response.model)
out.set_attribute("gen_ai.response.id", data.response.id)
self._usage(out, data.usage)
elif isinstance(data, GenerationSpanData):
if data.model:
out.update_name(f"chat {data.model}")
out.set_attribute("gen_ai.request.model", data.model)
self._usage(out, data.usage)
elif isinstance(data, FunctionSpanData) and self._capture:
# Opt-In attributes. Both are None when trace_include_sensitive_data=False.
if data.input:
out.set_attribute("gen_ai.tool.call.arguments", data.input)
if data.output is not None:
out.set_attribute("gen_ai.tool.call.result", str(data.output))
elif isinstance(data, HandoffSpanData):
out.set_attribute("openai_agents.handoff.from", data.from_agent or "")
out.set_attribute("openai_agents.handoff.to", data.to_agent or "")
elif isinstance(data, GuardrailSpanData):
out.set_attribute("openai_agents.guardrail.triggered", data.triggered)
if span.error:
# SpanError is a message plus a dict, not an exception class, so there is
# no low-cardinality error name to report: fall back to the spec's _OTHER.
out.set_attribute("error.type", "_OTHER")
out.set_status(Status(StatusCode.ERROR, span.error.get("message", "")))
out.end()
@staticmethod
def _usage(out, usage):
if not usage:
return
out.set_attribute("gen_ai.usage.input_tokens", usage.get("input_tokens", 0))
out.set_attribute("gen_ai.usage.output_tokens", usage.get("output_tokens", 0))
cached = (usage.get("input_tokens_details") or {}).get("cached_tokens", 0)
if cached: # already included in input_tokens, per the spec
out.set_attribute("gen_ai.usage.cache_read.input_tokens", cached)
def shutdown(self) -> None:
self._provider.shutdown()
def force_flush(self) -> None:
self._provider.force_flush()Two details in it follow the conventions closely. Attributes that the spec says should be provided at span creation time, such as the operation name, tool name and agent name, are passed to start_span so that a sampler can use them. The chat span is renamed once the response arrives, because the model name is not known when the span starts. Wiring it in takes a tracer provider, an OTLP exporter pointed at the collector that already forwards to Tempo, and one registration call.
# telemetry.py -- call configure_tracing() once, at process start
from agents import add_trace_processor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from genai_processor import GenAISemconvProcessor
def configure_tracing() -> None:
provider = TracerProvider(resource=Resource.create({
"service.name": "erp-invoice-agent",
"deployment.environment.name": "staging",
}))
# Alloy or the OTel Collector on the same Docker network, forwarding to Tempo.
provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://alloy:4318/v1/traces")))
# add_trace_processor KEEPS the default exporter to OpenAI's trace dashboard.
# set_trace_processors([...]) REPLACES it -- use that when traces must stay on your VPS.
add_trace_processor(GenAISemconvProcessor(provider, capture_content=False))
# agent_run.py
from agents import RunConfig, Runner
result = await Runner.run(
triage_agent,
"Why is the invoice for PO-2026-0412 blocked?",
run_config=RunConfig(
workflow_name="invoice_dispute", # becomes gen_ai.workflow.name
group_id=ticket.number, # becomes gen_ai.conversation.id
trace_include_sensitive_data=False, # tool inputs/outputs never leave the SDK
),
)Long-running workers need one more line. The SDK exports in background batches, and its documentation recommends calling flush_traces after a unit of work in Celery or FastAPI background tasks so that traces are exported immediately. With the processor above, force_flush also flushes the OpenTelemetry batch processor, so the spans reach Tempo before the worker is recycled.
Choose between add_trace_processor and set_trace_processors deliberately. add_trace_processor adds your processor alongside the default exporter, so every trace still goes to OpenAI's hosted trace dashboard as well as to Tempo. set_trace_processors replaces the default list, which is what you want when traces must stay on your own servers. Separately, the SDK documentation states that tracing is unavailable for organisations using OpenAI's APIs under a Zero Data Retention policy, so check that before designing around it.
Most of the value of the conventions comes from a handful of attributes being consistent across every service that emits them. These are the ones where the spec makes a specific demand that is easy to miss.
Resource attributes do the rest. service.name identifies which agent service a trace came from, and an environment attribute keeps staging traces out of production dashboards. Neither is GenAI-specific, which is exactly why agent traces can sit in the same Tempo instance as the rest of the platform and be joined to the HTTP and database spans around them.
Prompts, model outputs and tool results in an ERP agent contain supplier names, invoice amounts and sometimes personal data. The model spans document is explicit that this content is sensitive, often large, and should not be captured by default. It describes three usage patterns, in increasing order of operational maturity.
On the SDK side, RunConfig.trace_include_sensitive_data defaults to true, and the OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA environment variable changes that default. Setting it to false empties the input and output fields before any processor sees them, which is safer than relying on your own processor to drop them. The processor above adds a second gate, capture_content, so that turning on capture is a deliberate change in two places. If you do enable it on a self-hosted stack, Tempo has its own limit to plan for.
# tempo.yaml -- the default silently cuts long attribute values
distributor:
# Default 2048 bytes. A tool result holding a 40-line purchase order, or any
# gen_ai.input.messages value, is truncated before it is stored -- the span
# still arrives, so nothing looks broken until someone reads it.
max_attribute_bytes: 16384
# Watch for it instead of guessing:
# sum by (scope) (rate(tempo_distributor_attributes_truncated_total[5m]))Raising the limit is a trade, not a fix. The Tempo troubleshooting guide documents max_attribute_bytes precisely because large attributes are a cause of querier out-of-memory errors, and the default of 2048 bytes exists to protect the cluster. Raise it only as far as the content you have chosen to capture requires, and prefer the external storage pattern once message capture is more than a debugging aid.
The tempo_distributor_attributes_truncated_total metric carries a scope label, and the distributor also logs a rate-limited example of each truncated attribute. Alert on the metric for the span scope after enabling capture, and you will learn about truncation from a graph instead of from an engineer reading half a tool result.
Once spans follow the conventions, TraceQL queries can be written against meaning rather than against service-specific span names. Span attributes use the span. prefix, intrinsics such as status and duration use span: with a colon, and the structural operators select children, descendants and siblings. These are the queries that answer the questions from the opening example.
// Every failed tool call, with the agent that made it
{ span.gen_ai.operation.name = "execute_tool" && span:status = error }
| select(span.gen_ai.tool.name, span.gen_ai.agent.name)
// Runs where one agent called the same tool more than five times -- a looping agent
{ span.gen_ai.operation.name = "invoke_agent" }
> { span.gen_ai.tool.name = "get_purchase_order" } | count() > 5
// Slow model calls inside the dispute workflow only
{ span.gen_ai.workflow.name = "invoice_dispute" }
>> { span.gen_ai.operation.name = "chat" && span:duration > 8s }
// Every trace for one helpdesk ticket, across retries and follow-up messages
{ span.gen_ai.conversation.id = "TCK-2026-10-0193" }
// Failed spans that sit beside a handoff under the same agent
{ span:name = "agent.handoff" } ~ { span:status = error }The second query is the one that pays for the whole exercise. A looping agent is invisible in service metrics, because each tool call is a fast, successful request, yet it is the most common way an agent burns tokens without producing an answer. The child operator only works because the processor removed the turn spans; with them in place, the same query needs the descendant operator instead.
A custom processor is not the only route. OpenTelemetry now publishes opentelemetry-instrumentation-genai-openai-agents, which registers a tracing processor with the SDK and produces workflow, agent and tool spans. It continues the earlier opentelemetry-instrumentation-openai-agents-v2 package, which now receives only security patches. The two approaches differ in ways that decide the choice.
| Concern | Custom processor | Official package |
|---|---|---|
| Model call spans | Emitted from ResponseSpanData and GenerationSpanData in the same processor. | Not emitted; installing opentelemetry-instrumentation-genai-openai alongside it provides them. |
| Message content | Tool arguments and results only; converting SDK input items to the messages schema is left to you. | Controlled by an environment variable, with an upload completion hook for external storage. |
| Tracking spec changes | Manual. Every rename in a Development-status spec is your edit. | Follows the conventions through package releases, currently in beta. |
| Handoffs, guardrails, naming | Fully under your control, including vendor attributes and span skipping. | Fixed by the package; the workflow span name comes from workflow_name. |
# pip install opentelemetry-instrumentation-genai-openai-agents \
# opentelemetry-instrumentation-genai-openai # the chat spans come from this one
from opentelemetry.instrumentation.genai.openai_agents import OpenAIAgentsInstrumentor
# Default keeps uploading to OpenAI's hosted tracing as well; this routes only to OTel.
OpenAIAgentsInstrumentor().instrument(disable_openai_trace_export=True)
# Message capture is an environment switch, off unless you set it:
# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=...
# or ship content to storage instead of span attributes:
# OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK=upload
# OTEL_INSTRUMENTATION_GENAI_UPLOAD_BASE_PATH=...As of this writing, the opentelemetry-instrumentation-genai-openai-agents package on PyPI is at 1.2b0. By default it keeps the SDK's upload to OpenAI's hosted tracing active alongside OpenTelemetry, and disable_openai_trace_export=True routes traces only through OpenTelemetry.
A reasonable split is to start with the official package for the standard shape and switch to a custom processor only when you need something it does not do, such as per-tool redaction rules or ERP-specific attributes on every agent span. The custom processor in this post is short enough that maintaining it is realistic, but it does mean reading the conventions changelog, because a Development-status spec is allowed to rename the attributes your dashboards depend on.
Agent traces become useful when they follow a shared shape: a workflow or agent root, model calls and tool calls as siblings beneath it, the same attribute names in every service, and content left out unless someone decided to capture it. Pin the conventions version you implemented, set a real conversation id, keep tools out from under chat spans, and check what your backend truncates before you trust what it shows.
Sources