Quick start
Export is off by default. SetOTEL_ENABLED=true and point the exporter at
your collector:
opentelemetry enabled at startup together with the selected
exporters. Exporters are asynchronous: the gateway starts even when the
collector is down and retries in the background.
Without
OTEL_EXPORTER_OTLP_ENDPOINT the SDK sends to
http://localhost:4318 (http/protobuf) or localhost:4317 (grpc).Configuration
OTEL_ENABLED is the only GoModel-specific switch. Everything else follows
the standard
OpenTelemetry environment variables,
so GoModel is configured exactly like any other instrumented service. The
common settings are also available under opentelemetry: in
config.yaml; a variable set in the environment
wins over the YAML value.
A typical production setup:
What is exported
HTTP server spans and metrics
Every request gets aSERVER span named after its route (for example
POST /v1/chat/completions) with the standard http.request.method,
http.route, http.response.status_code, and url.scheme attributes, plus the
http.server.request.duration histogram and the request and response body size
histograms.
Trace context from the caller is honored, so a gateway span appears as a child
of the calling application’s span whenever the client sends traceparent (or
whichever headers OTEL_PROPAGATORS selects).
/health, /health/ready, the Prometheus endpoint (METRICS_ENDPOINT), and
/debug/pprof are excluded: they are polled, and would otherwise dominate the
trace volume.
Provider call spans and GenAI metrics
Each logical call to a model provider is instrumented with:
The response and usage attributes are set on chat completions, Responses API,
and translated
/v1/messages calls, buffered or streamed, once the gateway has
decoded the response. Usage is present when the provider reported it. For
streams, GoModel asks the provider for a usage chunk only while usage tracking
is on (the default). /v1/messages requests that GoModel forwards natively to an Anthropic
provider, and other passthrough calls, carry the request attributes only.
Metrics are histograms in seconds, carrying the same attributes:
gomodel.client.empty_responses counts buffered chat and Responses API calls
that returned 200 with no choices, no output, or no token usage. Those calls
look successful in the histograms above, so alert on this counter instead. Its error.type is
no_choices, no_output (a completed Responses API call with no output
items), or no_usage (content with zero token usage).
A buffered call produces a CLIENT span named <operation> <model>, such as
chat gpt-5, nested under the HTTP server span. Retries and failovers to
another provider are separate calls and therefore separate spans, so a request
that failed over shows exactly which provider failed and which one answered.
A streamed chat completion, Responses API, or translated /v1/messages call
gets a client span that ends when the stream ends, so its duration covers the
whole generation. Other streams, such as passthrough calls, get no client
span: the gateway sees only when they were established, and a span ending at
the headers would misreport latency. Their HTTP server span still covers the
full stream lifetime as seen by the client. A stream that fails to establish,
or ends before delivering its first chunk, always gets a failure span, so
errors are traced.
Prompt and completion capture
Content capture is off by default. Turn it on to see the exchanged messages in your tracing backend, for example as generation input and output in Langfuse:/v1/messages calls then carry the
GenAI message attributes
as JSON:
Text parts longer than 64 KiB are truncated and marked, and each attribute
keeps at most 512 KiB of text: a long conversation loses its oldest messages’
text first. Images, audio, and files are reduced to their part type: binary
data and data URLs are never exported. Requests dropped by the trace sampler
are not captured at all.
The variable also accepts
SPAN_ONLY and SPAN_AND_EVENT, the mode names of
newer OpenTelemetry instrumentations.
Privacy
The exporter is designed so that telemetry can go to a third-party backend without leaking what flows through the gateway:- Prompts, completions, and tool calls are attached to spans only when content capture is on, and never to metrics. Request and response bodies are never exported as-is.
- Credentials and upstream error messages are never exported; failures carry only a status code or an error class.
client.address,network.peer.*,server.address,server.port, anduser_agent.originalare stripped from HTTP spans, and host-derived dimensions are excluded from HTTP metrics so a client cannot inflate metric cardinality through theHostheader.
Reloading
gomodel --reload (SIGHUP) rebuilds the OpenTelemetry pipeline with the
current environment, so exporter, sampling, and propagation changes apply
without a restart. The previous pipeline is flushed before it is discarded.
Local collector example
A minimaldocker-compose.yml that shows traces in Jaeger:
http://localhost:16686.