> ## Documentation Index
> Fetch the complete documentation index at: https://gomodel-feat-litellm-db-import.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenTelemetry

> Export GoModel traces and metrics over OTLP with OTEL_ENABLED, configure the collector with standard OTEL_* variables, and see which spans and GenAI metrics are exported.

GoModel exports OpenTelemetry traces and metrics for inbound HTTP requests
(except operational endpoints, see below) and every outbound model provider
call. Provider calls follow the
[GenAI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/),
so a Jaeger, Grafana Tempo, Honeycomb, Datadog, or any other OTLP-compatible
backend shows which model was called, through which provider, how long it took,
and whether it failed.

Credentials and error messages are never exported. Prompts and completions
are exported only when you turn on
[content capture](#prompt-and-completion-capture). For Prometheus scraping, see [Prometheus Metrics](/guides/prometheus-metrics); both
can run at the same time.

## Quick start

Export is **off by default**. Set `OTEL_ENABLED=true` and point the exporter at
your collector:

<CodeGroup>
  ```bash Docker (.env file) theme={null}
  # Add to your .env file, then:
  # OTEL_ENABLED=true
  # OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
  docker run --rm -p 8080:8080 --env-file .env enterpilot/gomodel
  ```

  ```bash Docker (inline -e) theme={null}
  docker run --rm -p 8080:8080 \
    -e OTEL_ENABLED=true \
    -e OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 \
    enterpilot/gomodel
  ```

  ```bash Binary theme={null}
  export OTEL_ENABLED=true
  export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
  ./bin/gomodel
  ```
</CodeGroup>

GoModel logs `opentelemetry enabled` at startup together with the selected
exporters. Exporters are asynchronous: the gateway starts even when the
collector is down and retries in the background.

<Note>
  Without `OTEL_EXPORTER_OTLP_ENDPOINT` the SDK sends to
  `http://localhost:4318` (`http/protobuf`) or `localhost:4317` (`grpc`).
</Note>

## Configuration

`OTEL_ENABLED` is the only GoModel-specific switch. Everything else follows
the standard
[OpenTelemetry environment variables](https://opentelemetry.io/docs/specs/otel/configuration/sdk-environment-variables/),
so GoModel is configured exactly like any other instrumented service. The
common settings are also available under `opentelemetry:` in
[`config.yaml`](/advanced/config-yaml); a variable set in the environment
wins over the YAML value.

| Environment variable | `config.yaml` key | Default | Purpose |
| - | - | - | - |
| `OTEL_ENABLED` | `enabled` | `false` | Turn export on. |
| `OTEL_SERVICE_NAME` | `service_name` | product name (`gomodel`; `gomodel-pro` in Pro) | `service.name` resource attribute. |
| `OTEL_RESOURCE_ATTRIBUTES` | `resource_attributes` | — | Extra resource attributes such as `deployment.environment`. |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | `endpoint` | `http://localhost:4318` | Collector endpoint. `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` and `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` override it per signal. |
| `OTEL_EXPORTER_OTLP_PROTOCOL` | `protocol` | `http/protobuf` | `http/protobuf` or `grpc`; overridable per signal with `OTEL_EXPORTER_OTLP_TRACES_PROTOCOL` and `OTEL_EXPORTER_OTLP_METRICS_PROTOCOL`. |
| `OTEL_EXPORTER_OTLP_HEADERS` | `headers` | — | Headers such as `authorization` for hosted backends. |
| `OTEL_TRACES_EXPORTER` | `traces_exporter` | `otlp` | `otlp` or `none` to switch traces off. |
| `OTEL_METRICS_EXPORTER` | `metrics_exporter` | `otlp` | `otlp` or `none` to switch metrics off. |
| `OTEL_TRACES_SAMPLER` | `sampler` | `parentbased_always_on` | Sampler, e.g. `parentbased_traceidratio` to keep a fraction of new traces. |
| `OTEL_TRACES_SAMPLER_ARG` | `sampler_arg` | — | Sampler argument, e.g. `0.1` for 10%. |
| `OTEL_PROPAGATORS` | `propagators` | `tracecontext,baggage` | Comma-separated: `tracecontext`, `baggage`, `b3`, `b3multi`, `jaeger`, `ottrace`, or `none`. |
| `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` | `capture_message_content` | `false` | Record prompts and completions on provider call spans. See [Prompt and completion capture](#prompt-and-completion-capture). |
| `OTEL_METRIC_EXPORT_INTERVAL` | — | `60000` | Metric export interval in milliseconds. |

A typical production setup:

<CodeGroup>
  ```bash Environment theme={null}
  OTEL_ENABLED=true
  OTEL_SERVICE_NAME=gomodel-production
  OTEL_RESOURCE_ATTRIBUTES=deployment.environment=production
  OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317
  OTEL_EXPORTER_OTLP_PROTOCOL=grpc
  OTEL_TRACES_SAMPLER=parentbased_traceidratio
  OTEL_TRACES_SAMPLER_ARG=0.1
  ```

  ```yaml config.yaml theme={null}
  opentelemetry:
    enabled: true
    service_name: gomodel-production
    resource_attributes:
      deployment.environment: production
    endpoint: http://otel-collector:4317
    protocol: grpc
    headers:
      authorization: "Bearer ${OTEL_BACKEND_TOKEN}"
    sampler: parentbased_traceidratio
    sampler_arg: "0.1"
  ```
</CodeGroup>

Sampling only affects traces. Metrics are always complete.

<Warning>
  Export headers such as `authorization` travel with every export request. Use
  an `https://` endpoint (or a gRPC endpoint with TLS) whenever headers carry
  credentials; plaintext is acceptable only for a collector on the same host.
  GoModel logs a warning at startup when headers are configured and the
  effective endpoint is plaintext (`http://`, or gRPC with
  `OTEL_EXPORTER_OTLP_INSECURE=true`) and not a loopback address.
</Warning>

## What is exported

### HTTP server spans and metrics

Every request gets a `SERVER` span named after its route (for example
`POST /v1/chat/completions`) with the standard `http.request.method`,
`http.route`, `http.response.status_code`, and `url.scheme` attributes, plus the
`http.server.request.duration` histogram and the request and response body size
histograms.

Trace context from the caller is honored, so a gateway span appears as a child
of the calling application's span whenever the client sends `traceparent` (or
whichever headers `OTEL_PROPAGATORS` selects).

`/health`, `/health/ready`, the Prometheus endpoint (`METRICS_ENDPOINT`), and
`/debug/pprof` are excluded: they are polled, and would otherwise dominate the
trace volume.

### Provider call spans and GenAI metrics

Each logical call to a model provider is instrumented with:

| Attribute | Example | Notes |
| - | - | - |
| `gen_ai.operation.name` | `chat` | Also `embeddings`, `generate_content`, and others depending on the endpoint. |
| `gen_ai.request.model` | `gpt-5` | The model sent upstream, after alias and virtual model resolution. |
| `gen_ai.provider.name` | `openai` | Semantic provider name: `openai`, `anthropic`, `aws.bedrock`, `gcp.vertex_ai`, `gcp.gemini`, `azure.ai.openai`, `x_ai`, … or `unknown`. |
| `gomodel.provider.name` | `openai-eu` | The exact provider name from your GoModel configuration. |
| `gen_ai.request.stream` | `true` | Present only on streaming calls. |
| `error.type` | `429`, `timeout`, `network_error`, `empty_stream` | Present only on failures; `empty_stream` marks a stream that ended before its first chunk. |
| `http.response.status_code` | `200` | Upstream status. |
| `gen_ai.response.model` | `gpt-5-2025-08-07` | The model the provider reports, often a dated version. |
| `gen_ai.response.id` | `chatcmpl-…` | The provider's response ID. |
| `gen_ai.response.finish_reasons` | `["stop"]` | One per choice. Responses API calls report `stop` when completed, or the incomplete reason such as `max_output_tokens`. |
| `gen_ai.usage.input_tokens` | `1200` | Prompt tokens, as GoModel's [usage tracking](/advanced/configuration#token-usage-tracking) records them. |
| `gen_ai.usage.output_tokens` | `85` | Completion tokens, including reasoning tokens. |

The response and usage attributes are set on chat completions, Responses API,
and translated `/v1/messages` calls, buffered or streamed, once the gateway has
decoded the response. Usage is present when the provider reported it. For
streams, GoModel asks the provider for a usage chunk only while usage tracking
is on (the default). `/v1/messages` requests that GoModel forwards natively to an Anthropic
provider, and other passthrough calls, carry the request attributes only.

Metrics are histograms in seconds, carrying the same attributes:

| Metric | Recorded for |
| - | - |
| `gen_ai.client.operation.duration` | Every buffered call, and every streaming call that fails before the stream is established or ends before its first chunk. |
| `gen_ai.client.operation.time_to_first_chunk` | Every successful streaming call, measured until the first response bytes arrive. |

`gomodel.client.empty_responses` counts buffered chat and Responses API calls
that returned `200` with no choices, no output, or no token usage. Those calls
look successful in the histograms above, so alert on this counter instead. Its `error.type` is
`no_choices`, `no_output` (a completed Responses API call with no output
items), or `no_usage` (content with zero token usage).

A buffered call produces a `CLIENT` span named `<operation> <model>`, such as
`chat gpt-5`, nested under the HTTP server span. Retries and failovers to
another provider are separate calls and therefore separate spans, so a request
that failed over shows exactly which provider failed and which one answered.

A streamed chat completion, Responses API, or translated `/v1/messages` call
gets a client span that ends when the stream ends, so its duration covers the
whole generation. Other streams, such as passthrough calls, get no client
span: the gateway sees only when they were established, and a span ending at
the headers would misreport latency. Their HTTP server span still covers the
full stream lifetime as seen by the client. A stream that fails to establish,
or ends before delivering its first chunk, always gets a failure span, so
errors are traced.

### Prompt and completion capture

Content capture is **off by default**. Turn it on to see the exchanged
messages in your tracing backend, for example as generation input and output
in [Langfuse](/guides/langfuse):

<CodeGroup>
  ```bash Environment theme={null}
  OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true
  ```

  ```yaml config.yaml theme={null}
  opentelemetry:
    capture_message_content: true
  ```
</CodeGroup>

Provider call spans for chat completions, Responses API, and translated
`/v1/messages` calls then carry the
[GenAI message attributes](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/)
as JSON:

| Attribute | Content |
| - | - |
| `gen_ai.input.messages` | The request messages, including system messages, tool calls, and tool results. |
| `gen_ai.output.messages` | The response messages with their finish reason. |
| `gen_ai.system_instructions` | The Responses API `instructions`. |

Text parts longer than 64 KiB are truncated and marked, and each attribute
keeps at most 512 KiB of text: a long conversation loses its oldest messages'
text first. Images, audio, and files are reduced to their part type: binary
data and data URLs are never exported. Requests dropped by the trace sampler
are not captured at all.
The variable also accepts `SPAN_ONLY` and `SPAN_AND_EVENT`, the mode names of
newer OpenTelemetry instrumentations.

<Warning>
  With capture on, every sampled prompt and completion leaves the gateway with
  its trace. Enable it only for a backend that may store that data, and
  consider sampling (`OTEL_TRACES_SAMPLER`) to limit the volume. The
  [audit log](/advanced/configuration#audit-logging) keeps content inside your own storage
  instead.
</Warning>

## Privacy

The exporter is designed so that telemetry can go to a third-party backend
without leaking what flows through the gateway:

* Prompts, completions, and tool calls are attached to spans only when
  [content capture](#prompt-and-completion-capture) is on, and never to
  metrics. Request and response bodies are never exported as-is.
* Credentials and upstream error messages are never exported; failures carry
  only a status code or an error class.
* `client.address`, `network.peer.*`, `server.address`, `server.port`, and
  `user_agent.original` are stripped from HTTP spans, and host-derived
  dimensions are excluded from HTTP metrics so a client cannot inflate metric
  cardinality through the `Host` header.

## Reloading

`gomodel --reload` (SIGHUP) rebuilds the OpenTelemetry pipeline with the
current environment, so exporter, sampling, and propagation changes apply
without a restart. The previous pipeline is flushed before it is discarded.

## Local collector example

A minimal `docker-compose.yml` that shows traces in Jaeger:

```yaml theme={null}
services:
  gomodel:
    image: enterpilot/gomodel
    env_file: .env
    environment:
      OTEL_ENABLED: "true"
      OTEL_EXPORTER_OTLP_ENDPOINT: http://jaeger:4318
      OTEL_METRICS_EXPORTER: none
    ports: ["8080:8080"]
  jaeger:
    image: jaegertracing/jaeger:2.7.0
    ports: ["16686:16686"]
```

Send a request through the gateway and open `http://localhost:16686`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.