AI & LLM monitoring: See what your agents are really doing
Trace every agent, tool call, and model request. Score quality on live traffic and attribute cost to the token, in the same platform that runs the rest of your stack. OpenTelemetry-native, deploy anywhere, priced per GB instead of per span.

Purpose-built for agents. Priced for reality.
Most LLM tools are a silo you bolt onto your real observability stack, metered per span, locked to one cloud. OpenObserve is one platform for your AI and everything under it.
Unified, not bolt-on
LLM traces land next to your logs, metrics, traces, and RUM. When an agent is slow, see the pod, database, or vector store behind it. No second tool, no swivel-chair.
Deploy anywhere
Cloud in four regions, self-hosted, bring-your-own-cloud, or bring-your-own-bucket. Federated search spans regions and clouds while keeping egress under control.
Radically efficient
A Rust engine on columnar Parquet storage, billed per GB. Agentic apps fan out into hundreds of spans per request. You shouldn't pay for each one.
Instrument what you already run
Frameworks, model providers, gateways, and no-code builders, all through standard OpenTelemetry. If it emits gen_ai spans, it works today.
18 of 91 shown
OpenAI (Python)
Token usage, latency, and cost per call.
llm-providersAnthropic (Python)
Trace Claude calls with full token accounting.
llm-providersLangChain
Trace every chain, tool, and LLM step.
ai-frameworksGoogle Gemini
Prompts, completions, and latency.
ai-frameworksAmazon Bedrock
Trace foundation-model invocations.
llm-providersMistral
Span-level traces for every completion.
llm-providersOllama
Local model traces - nothing leaves the host.
llm-providersDeepSeek
Prompt, completion, and cost telemetry.
llm-providersOpenAI (JS/TS)
JavaScript SDK calls, fully instrumented.
llm-providersOpenAI Assistants
Trace assistant runs and tool calls.
llm-providersAnthropic (JS/TS)
TypeScript-first Claude tracing via OTel.
llm-providersCohere
Token cost and latency per generation.
llm-providersGroq
Sub-second inference, fully traced.
llm-providersHugging Face
Trace inference endpoints and pipelines.
llm-providersvLLM
Self-hosted serving, span by span.
llm-providersTogether AI
Token and latency metrics per call.
llm-providersFireworks AI
Fast inference with full trace context.
llm-providersxAI Grok
Capture every Grok completion.
llm-providersFrom a runaway agent to the trace behind it
Trace & map
Know exactly what your agents touch. Every agent request is a distributed trace. OpenObserve maps each agent to the models, tools, services, and datastores it calls, with request counts and error health on every edge, so a runaway loop or a failing tool is obvious at a glance.
Agent Graph
Renders the full call tree across sub-agents, tools, and models.
Health at a glance
Border colors flag healthy, degraded, and critical paths by error rate.
Compare releases
Filter by environment, agent, and version to compare releases.

Debug sessions
Debug a bad answer in one view. Open any session and replay the whole conversation: every turn, every tool call, the model behind it, and where the cost and latency actually went. Stop grepping logs to reconstruct what an agent did.
Session ribbons
Break down cost, duration, and tokens per turn.
Tool, cost, and latency hotspots
Surface the expensive, slow steps instantly.
One hop to the trace
Jump straight from a turn to its full distributed trace.

Evaluate in production
Measure quality on real traffic, continuously. Online evaluations score live spans, traces, and whole sessions the moment they arrive. Use LLM-as-judge with your own provider, or call your remote scorer. Set a healthy threshold per metric and watch quality on a live dashboard instead of a one-off notebook.
Built-in scorers
Relevance, hallucination, toxicity, bias, and more.
Score at any scope
Span, trace, or session scope, on a sample or everything.
Versioned score configs
A Quality view that flags what needs attention.

Close the loop
From production trace to eval set. Route real traces into review queues, let humans score them alongside the automatic evaluators, and distill the good and bad ones into datasets you can test future versions against. Discovery surfaces the failures worth reviewing in the first place.
Annotation queues
Reviewer scores layered over system scores.
Distill to a dataset
One click to turn a reviewed trace into a dataset.
Agent Behavior
Catches loops and groups failures by kind.

AI is a first-class layer in one unified stack
Your AI and LLM traffic is just another source, flowing through the same correlation engine as your frontend, APIs, databases, and infrastructure. Traces, metrics, logs, LLM observability, evals, and AI SRE live in one platform, queried together with SQL and PromQL.
Unified observability stack
Every signal from your estate flows into one correlated view.



Your AI data, where you want it
Prompts and completions carry your most sensitive data. Run OpenObserve however your security and cost model demands, and query across all of it.
Managed Cloud
Fully hosted, zero ops. Four regions today, with data residency where you need it.
Self-hosted
A single Rust binary from laptop to petabyte scale. Open source, no license wall.
Bring your own Bucket
Point storage at your own S3 or object storage bucket, or let us operate a managed deployment inside your cloud account. Your perimeter, your data; nothing locked in.
Go deeper on AI observability
Guides, walkthroughs, and conversations on tracing agents, evaluating LLMs, and running it all on OpenTelemetry.


