AI Deck

Datadog — A SaaS observability platform that now sees inside AI agents and LLM applications

A SaaS observability platform from Datadog, Inc. It collects metrics, logs and traces from servers, containers and cloud services into one place, correlates them, and surfaces performance degradation and early signs of failure. Since its debut in 2010 it has been a staple of infrastructure monitoring, but its expansion into generative AI has been striking in recent years. LLM Observability tracks response quality and cost for generative AI applications, while Agent Observability visualizes an AI agent’s input, tool calls and output as a conversation graph. On top of that, a set of AI agents called Bits AI automates failure investigation and incident response, and an official MCP server lets external AI agents such as Claude Code query operational data directly.

Key Features

  • LLM Observability: Automatically instruments major LLM and agent frameworks including OpenAI, Anthropic, Gemini, Amazon Bedrock, LangChain and CrewAI, tracking response quality, token cost and latency. You can trace “why did this answer come back” down to the individual trace
  • Agent Observability (AI agent monitoring): Renders an agent’s input, tool calls, sub-agent calls and final output as a conversation graph. You can follow where the reasoning went wrong or which tool failed, which suits debugging multi-agent setups
  • Bits AI: A set of AI agents including SRE Agent, Dev Agent, Security Analyst and Bits Chat. It automates operational work such as asking in natural language “which services had a rising error rate in the last hour”, narrowing down root cause candidates, and proposing fixes
  • Official MCP server: Lets AI agents such as Claude Code query Datadog metrics, logs, traces and incidents directly. It is a remote connection supporting OAuth 2.0 or API key plus application key authentication, with separate endpoints per site (US, EU and others)
  • Infrastructure monitoring and APM: The foundation that visualizes servers, containers and distributed traces across the stack. With over 1,000 integrations, most major clouds and middleware can be brought under monitoring with configuration alone
  • Governance features: Sensitive Data Scanner for detecting and redacting sensitive data, plus role-based access control. The controls you need when running AI agents in production are provided by the product itself

Pricing

Pricing is modular and usage-based, and each product bills on a different axis (hosts, GB, events, credits). The main items are excerpted below.

ProductAnnual billingOn-demandNotes
Infrastructure Monitoring Free$0Up to 5 hosts, 1-day metric retention
Infrastructure Monitoring Pro$15/host/month$18/host/month1,000+ integrations, 15-month retention
Infrastructure Monitoring Enterprise$23/host/month$27/host/monthML-based alerting, Governance Console
APM (with Infrastructure)$31/host/month$48/host/monthDistributed tracing, service discovery
APM Enterprise (same)$40/host/month$60/host/monthCode-level analysis via Continuous Profiler
Log Management (ingest)$0.10/GB/monthSamePer ingested or scanned GB
Log Management (standard indexing)$1.70/million events/month$2.55/million events/month15-day retention by default
AI Credits$500/500 credits/month$1.30/creditConsumed by Bits Chat, Bits Investigation, Bits Code and others

Unit pricing for LLM Observability is not listed separately on the official pricing page and varies by configuration, so checking with Datadog directly is the reliable route. Many products are not listed here either, so refer to the official pricing page for an actual estimate.

Pricing is current as of August 2026. Please check the official site for the latest information.

Pros & Cons

Pros

  • Metrics, logs, traces and security can be correlated on a single platform, so you can reach root cause without moving between tools
  • Over 1,000 integrations bring major clouds and middleware under monitoring almost immediately
  • LLM Observability and Agent Observability make the quality, cost and decision path of generative AI applications and AI agents visible
  • The official MCP server lets AI agents query operational data in natural language
  • A free tier (up to 5 hosts) makes it possible to start small and expand

⚠️ Cons

  • Because pricing is modular and usage-based, total cost becomes harder to predict the more products you combine
  • With several billing axes (hosts, log volume, APM, AI credits), unit prices can feel high for small teams
  • The sheer number of features means initial setup and cost optimization tend to require expertise
  • Endpoints differ per site (US, EU and others), so connection settings such as the MCP server require awareness of which site you are on

Comparison with Similar Services

CriteriaDatadogNew RelicDynatraceGrafana Cloud / OSS stack
Main strengthBreadth of integration and depth of AI featuresApplication-centric observability and AI anomaly detectionAutomated root cause analysis via Davis AILow-cost operation on an OSS base
Billing modelUsage-based per productUsers plus data volumePer host, per hourUsage-based; self-hosted is free
Generative AI monitoringLLM Observability, Agent ObservabilitySupportedSupportedBuilt by combining OSS instrumentation
AI agent integrationOfficial MCP server
Target scaleMid to large, cloud nativeSmall to mid-sizeLarge enterpriseSmall teams or cost-focused organizations

If integrating log analytics with security operations (SIEM) is the priority, Splunk also belongs in the comparison.

Who Is It For

  • Development teams putting generative AI applications and AI agents into production who want to continuously observe response quality, cost and failure patterns
  • SREs and infrastructure engineers who want infrastructure, APM and logs in one place across a microservices cloud environment
  • Teams that want to shorten the first response to incidents and are willing to invest in AI-driven root cause investigation
  • Developers who want AI agents such as Claude Code to read operational data and take over investigation and operational work
  • Organizations that would rather integrate observability and security monitoring than split them across separate tools

Summary

Datadog has grown from an infrastructure monitoring staple into a platform that both monitors AI and monitors with AI. LLM Observability and Agent Observability are the tools for answering the questions you inevitably hit when running generative AI applications in production — why did this output happen, and where is the cost growing — while the official MCP server opens operational data up to the AI agent side. On the other hand, the usage-based design makes cost harder to forecast the wider you cast the net. Starting from the free tier or a narrow monitoring scope and adding only the products you need is the practical path.

← Blog