Observability for LLM endpoints: what actually matters
Standard application monitoring covers latency, error rate, and uptime, but LLM endpoints introduce failure modes that generic dashboards miss entirely: silent quality degradation, token-level cost spikes, and prompt-injection attempts that never trigger an HTTP error. This post outlines the observability layer we consider necessary before calling an LLM deployment production-ready.
Token-level metrics, not just request-level ones
A request that returns 200 OK can still be a problem if it consumed ten times the expected number of output tokens or hit a maximum-length cutoff mid-response. Track input tokens, output tokens, and time-to-first-token separately, since each maps to a different part of the system that could be misbehaving.
Time-to-first-token is particularly important for streaming interfaces, since a slow first token creates a perceived latency problem even if total generation time is acceptable. Alert on this metric independently from total request duration.
Quality drift is invisible to standard monitoring
A model can continue returning valid, well-formed responses that are simply wrong or off-topic, and nothing in a standard APM tool will flag this. Sample a percentage of production responses for automated or periodic human review against a rubric specific to the task.
For RAG systems, log which retrieved chunks were used for each response, so a drop in answer quality can be traced back to a retrieval problem, an embedding index that has gone stale, or a change in the underlying model's behavior.
Cost observability at the request level
Aggregate monthly cost dashboards catch budget overruns after the fact. Request-level cost logging, tagged by customer, feature, or API key, lets a team identify the specific workflow driving cost before it becomes a large invoice line item.
This is especially relevant when different requests route to different models by design, for example a cheap model for simple classification and a larger model for open-ended generation. Cost logs should reflect which model actually served each request.
- Log model name, token counts, and estimated cost per request
- Tag requests by feature or customer for cost attribution
- Alert on sudden shifts in average tokens per request, which often signal a prompt or retrieval regression
Tracing across retrieval, generation, and tools
A single user-facing request in a RAG or agentic system may involve an embedding call, a vector search, a reranking step, one or more LLM calls, and possibly a tool invocation. Distributed tracing across these steps, with each step's latency and output logged, is the difference between a five-minute root-cause investigation and a multi-hour one.
Standard tracing tools work well here as long as each step is instrumented consistently. The additional discipline required is capturing enough of the intermediate output, such as retrieved chunk identifiers or tool call arguments, to reconstruct why a particular response was generated.
Alerting thresholds specific to LLM behavior
Generic error-rate alerting misses cases where the model returns a technically successful response containing a refusal, an empty string, or a repeated token loop. Add explicit checks for these patterns and alert on their frequency.
Set alerting thresholds based on a baseline measured over at least a few weeks of production traffic rather than arbitrary numbers, since normal token counts and latency vary significantly by use case and by time of day.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.