A SaaS observability platform from Datadog, Inc. It collects metrics, logs and traces from servers, containers and cloud services into one place, correlates them, and surfaces performance degradation and early signs of failure. Since its debut in 2010 it has been a staple of infrastructure monitoring, but its expansion into generative AI has been striking in recent years. LLM Observability tracks response quality and cost for generative AI applications, while Agent Observability visualizes an AI agent’s input, tool calls and output as a conversation graph. On top of that, a set of AI agents called Bits AI automates failure investigation and incident response, and an official MCP server lets external AI agents such as Claude Code query operational data directly.
Key Features
- LLM Observability: Automatically instruments major LLM and agent frameworks including OpenAI, Anthropic, Gemini, Amazon Bedrock, LangChain and CrewAI, tracking response quality, token cost and latency. You can trace “why did this answer come back” down to the individual trace
- Agent Observability (AI agent monitoring): Renders an agent’s input, tool calls, sub-agent calls and final output as a conversation graph. You can follow where the reasoning went wrong or which tool failed, which suits debugging multi-agent setups
- Bits AI: A set of AI agents including SRE Agent, Dev Agent, Security Analyst and Bits Chat. It automates operational work such as asking in natural language “which services had a rising error rate in the last hour”, narrowing down root cause candidates, and proposing fixes
- Official MCP server: Lets AI agents such as Claude Code query Datadog metrics, logs, traces and incidents directly. It is a remote connection supporting OAuth 2.0 or API key plus application key authentication, with separate endpoints per site (US, EU and others)
- Infrastructure monitoring and APM: The foundation that visualizes servers, containers and distributed traces across the stack. With over 1,000 integrations, most major clouds and middleware can be brought under monitoring with configuration alone
- Governance features: Sensitive Data Scanner for detecting and redacting sensitive data, plus role-based access control. The controls you need when running AI agents in production are provided by the product itself
Pricing
Pricing is modular and usage-based, and each product bills on a different axis (hosts, GB, events, credits). The main items are excerpted below.
| Product | Annual billing | On-demand | Notes |
|---|---|---|---|
| Infrastructure Monitoring Free | $0 | ─ | Up to 5 hosts, 1-day metric retention |
| Infrastructure Monitoring Pro | $15/host/month | $18/host/month | 1,000+ integrations, 15-month retention |
| Infrastructure Monitoring Enterprise | $23/host/month | $27/host/month | ML-based alerting, Governance Console |
| APM (with Infrastructure) | $31/host/month | $48/host/month | Distributed tracing, service discovery |
| APM Enterprise (same) | $40/host/month | $60/host/month | Code-level analysis via Continuous Profiler |
| Log Management (ingest) | $0.10/GB/month | Same | Per ingested or scanned GB |
| Log Management (standard indexing) | $1.70/million events/month | $2.55/million events/month | 15-day retention by default |
| AI Credits | $500/500 credits/month | $1.30/credit | Consumed by Bits Chat, Bits Investigation, Bits Code and others |
Unit pricing for LLM Observability is not listed separately on the official pricing page and varies by configuration, so checking with Datadog directly is the reliable route. Many products are not listed here either, so refer to the official pricing page for an actual estimate.
Pricing is current as of August 2026. Please check the official site for the latest information.
Pros & Cons
✅ Pros
- Metrics, logs, traces and security can be correlated on a single platform, so you can reach root cause without moving between tools
- Over 1,000 integrations bring major clouds and middleware under monitoring almost immediately
- LLM Observability and Agent Observability make the quality, cost and decision path of generative AI applications and AI agents visible
- The official MCP server lets AI agents query operational data in natural language
- A free tier (up to 5 hosts) makes it possible to start small and expand
⚠️ Cons
- Because pricing is modular and usage-based, total cost becomes harder to predict the more products you combine
- With several billing axes (hosts, log volume, APM, AI credits), unit prices can feel high for small teams
- The sheer number of features means initial setup and cost optimization tend to require expertise
- Endpoints differ per site (US, EU and others), so connection settings such as the MCP server require awareness of which site you are on
Comparison with Similar Services
| Criteria | Datadog | New Relic | Dynatrace | Grafana Cloud / OSS stack |
|---|---|---|---|---|
| Main strength | Breadth of integration and depth of AI features | Application-centric observability and AI anomaly detection | Automated root cause analysis via Davis AI | Low-cost operation on an OSS base |
| Billing model | Usage-based per product | Users plus data volume | Per host, per hour | Usage-based; self-hosted is free |
| Generative AI monitoring | LLM Observability, Agent Observability | Supported | Supported | Built by combining OSS instrumentation |
| AI agent integration | Official MCP server | ─ | ─ | ─ |
| Target scale | Mid to large, cloud native | Small to mid-size | Large enterprise | Small teams or cost-focused organizations |
If integrating log analytics with security operations (SIEM) is the priority, Splunk also belongs in the comparison.
Who Is It For
- Development teams putting generative AI applications and AI agents into production who want to continuously observe response quality, cost and failure patterns
- SREs and infrastructure engineers who want infrastructure, APM and logs in one place across a microservices cloud environment
- Teams that want to shorten the first response to incidents and are willing to invest in AI-driven root cause investigation
- Developers who want AI agents such as Claude Code to read operational data and take over investigation and operational work
- Organizations that would rather integrate observability and security monitoring than split them across separate tools
Summary
Datadog has grown from an infrastructure monitoring staple into a platform that both monitors AI and monitors with AI. LLM Observability and Agent Observability are the tools for answering the questions you inevitably hit when running generative AI applications in production — why did this output happen, and where is the cost growing — while the official MCP server opens operational data up to the AI agent side. On the other hand, the usage-based design makes cost harder to forecast the wider you cast the net. Starting from the free tier or a narrow monitoring scope and adding only the products you need is the practical path.