Skip to main content
GoModel exports OpenTelemetry traces and metrics for inbound HTTP requests (except operational endpoints, see below) and every outbound model provider call. Provider calls follow the GenAI semantic conventions, so a Jaeger, Grafana Tempo, Honeycomb, Datadog, or any other OTLP-compatible backend shows which model was called, through which provider, how long it took, and whether it failed. Credentials and error messages are never exported. Prompts and completions are exported only when you turn on content capture. For Prometheus scraping, see Prometheus Metrics; both can run at the same time.

Quick start

Export is off by default. Set OTEL_ENABLED=true and point the exporter at your collector:
GoModel logs opentelemetry enabled at startup together with the selected exporters. Exporters are asynchronous: the gateway starts even when the collector is down and retries in the background.
Without OTEL_EXPORTER_OTLP_ENDPOINT the SDK sends to http://localhost:4318 (http/protobuf) or localhost:4317 (grpc).

Configuration

OTEL_ENABLED is the only GoModel-specific switch. Everything else follows the standard OpenTelemetry environment variables, so GoModel is configured exactly like any other instrumented service. The common settings are also available under opentelemetry: in config.yaml; a variable set in the environment wins over the YAML value. A typical production setup:
Sampling only affects traces. Metrics are always complete.
Export headers such as authorization travel with every export request. Use an https:// endpoint (or a gRPC endpoint with TLS) whenever headers carry credentials; plaintext is acceptable only for a collector on the same host. GoModel logs a warning at startup when headers are configured and the effective endpoint is plaintext (http://, or gRPC with OTEL_EXPORTER_OTLP_INSECURE=true) and not a loopback address.

What is exported

HTTP server spans and metrics

Every request gets a SERVER span named after its route (for example POST /v1/chat/completions) with the standard http.request.method, http.route, http.response.status_code, and url.scheme attributes, plus the http.server.request.duration histogram and the request and response body size histograms. Trace context from the caller is honored, so a gateway span appears as a child of the calling application’s span whenever the client sends traceparent (or whichever headers OTEL_PROPAGATORS selects). /health, /health/ready, the Prometheus endpoint (METRICS_ENDPOINT), and /debug/pprof are excluded: they are polled, and would otherwise dominate the trace volume.

Provider call spans and GenAI metrics

Each logical call to a model provider is instrumented with: The response and usage attributes are set on chat completions, Responses API, and translated /v1/messages calls, buffered or streamed, once the gateway has decoded the response. Usage is present when the provider reported it. For streams, GoModel asks the provider for a usage chunk only while usage tracking is on (the default). /v1/messages requests that GoModel forwards natively to an Anthropic provider, and other passthrough calls, carry the request attributes only. Metrics are histograms in seconds, carrying the same attributes: gomodel.client.empty_responses counts buffered chat and Responses API calls that returned 200 with no choices, no output, or no token usage. Those calls look successful in the histograms above, so alert on this counter instead. Its error.type is no_choices, no_output (a completed Responses API call with no output items), or no_usage (content with zero token usage). A buffered call produces a CLIENT span named <operation> <model>, such as chat gpt-5, nested under the HTTP server span. Retries and failovers to another provider are separate calls and therefore separate spans, so a request that failed over shows exactly which provider failed and which one answered. A streamed chat completion, Responses API, or translated /v1/messages call gets a client span that ends when the stream ends, so its duration covers the whole generation. Other streams, such as passthrough calls, get no client span: the gateway sees only when they were established, and a span ending at the headers would misreport latency. Their HTTP server span still covers the full stream lifetime as seen by the client. A stream that fails to establish, or ends before delivering its first chunk, always gets a failure span, so errors are traced.

Prompt and completion capture

Content capture is off by default. Turn it on to see the exchanged messages in your tracing backend, for example as generation input and output in Langfuse:
Provider call spans for chat completions, Responses API, and translated /v1/messages calls then carry the GenAI message attributes as JSON: Text parts longer than 64 KiB are truncated and marked, and each attribute keeps at most 512 KiB of text: a long conversation loses its oldest messages’ text first. Images, audio, and files are reduced to their part type: binary data and data URLs are never exported. Requests dropped by the trace sampler are not captured at all. The variable also accepts SPAN_ONLY and SPAN_AND_EVENT, the mode names of newer OpenTelemetry instrumentations.
With capture on, every sampled prompt and completion leaves the gateway with its trace. Enable it only for a backend that may store that data, and consider sampling (OTEL_TRACES_SAMPLER) to limit the volume. The audit log keeps content inside your own storage instead.

Privacy

The exporter is designed so that telemetry can go to a third-party backend without leaking what flows through the gateway:
  • Prompts, completions, and tool calls are attached to spans only when content capture is on, and never to metrics. Request and response bodies are never exported as-is.
  • Credentials and upstream error messages are never exported; failures carry only a status code or an error class.
  • client.address, network.peer.*, server.address, server.port, and user_agent.original are stripped from HTTP spans, and host-derived dimensions are excluded from HTTP metrics so a client cannot inflate metric cardinality through the Host header.

Reloading

gomodel --reload (SIGHUP) rebuilds the OpenTelemetry pipeline with the current environment, so exporter, sampling, and propagation changes apply without a restart. The previous pipeline is flushed before it is discarded.

Local collector example

A minimal docker-compose.yml that shows traces in Jaeger:
Send a request through the gateway and open http://localhost:16686.
Last modified on October 3, 2026