Observability on a budget: Datadog vs. Grafana vs. Honeycomb
How Datadog, Grafana Cloud, and Honeycomb charge for metrics, logs, and traces, why bills grow, and how a startup keeps observability costs under control.
By Lance King · · 10 min read
This guide is for founders and early engineers who need to see what their production systems are doing without signing up for a bill that grows faster than revenue. Observability means collecting three kinds of telemetry: metrics (numbers over time, like request rate), logs (lines of text your code writes), and traces (the path of a single request through your services). Datadog, Grafana, and Honeycomb all handle that job well. The real decision is which pricing model matches the shape of your data, and how much lock-in you accept along the way.
The short answer
Choose Datadog if you want one polished product covering infrastructure, APM, and logs, you run a modest number of hosts, and you will actively manage custom metrics and log indexing.
Choose Grafana Cloud if you want generous free limits, open-source formats (Prometheus, Loki, Tempo), and the option to self-host the same stack later.
Choose Honeycomb if your hardest problems are “why is this one request slow” questions and you want to query high-cardinality data without paying per unique tag value.
Choose self-hosted Grafana or SigNoz if you have real ops capacity and data volumes large enough that per-GB pricing hurts more than running servers does.
Whatever you choose, instrument with OpenTelemetry so switching later is a config change, not a rewrite.
What actually matters
- The unit you pay for. Datadog charges largely per host plus per GB and per indexed event. Grafana Cloud charges per active metric series and per GB of logs and traces. Honeycomb charges per event. The same workload can be cheap on one and expensive on another.
- Cardinality sensitivity. Cardinality is the number of unique combinations of tag (or label) values. Tag a metric with
user_idand you may create millions of distinct series. Some pricing models punish this, others don’t. - Log volume. Logs are usually the fastest-growing line item, because every new feature adds log lines and nobody removes them.
- Free tier headroom. How far you can get before the first invoice, and what retention you give up for it.
- Portability. Whether your instrumentation and dashboards survive a vendor change.
- Operational cost of self-hosting. Free software still needs storage, upgrades, and someone on call for it.
Pricing as of October 2026
All figures below come from each vendor’s public pricing page as of October 2026. Annual-billing prices are shown where the vendor lists both; on-demand is higher.
Datadog. The free Infrastructure tier covers up to 5 hosts with 1-day metric retention. Infrastructure Pro is $15 per host per month (annual) or $18 on-demand, with 15-month metric retention and 100 custom metrics allotted per host. Enterprise is $23 per host per month annual ($27 on-demand) with 200 custom metrics per host. APM starts at $31 per host per month annual ($48 on-demand), including 150 GB of span ingestion and 1 million indexed spans per APM host. Logs are $0.10 per GB ingested, plus $1.70 per million log events per month to index them for search (annual billing), with cheaper Flex storage at $0.05 per million events stored for long-term retention.
Grafana Cloud. The free tier includes 10,000 active metric series, 50 GB each of logs, traces, and profiles per month, 14-day retention, and 3 active users. Pro has a $19 monthly platform fee, then metrics from $6.50 per 1,000 series beyond the included 10,000. Logs and traces are priced in three parts: $0.05 per GB to process, $0.40 per GB to write, and $0.10 per GB to retain. Pro retention is 13 months for metrics and 30 days for logs, traces, and profiles. Enterprise starts at a $25,000 yearly commit.
Honeycomb. The free plan covers up to 20 million events and 100 million metric data points per month. Pro starts at $150 per month for 50 million events and scales up to 750 million events per month. Honeycomb raised its Pro rate on July 1, 2026 from $1.30 to $3.00 per million events; existing customers can stay on legacy plans through December 31, 2026. Retention is 60 days for most customers. Each span in a trace counts as one event, so a trace with 150 spans is 150 events.
Comparison table
Prices in this section are as of October 2026.
| Criterion | Datadog | Grafana Cloud | Honeycomb |
|---|---|---|---|
| Main billing unit | Per host, plus per GB ingested and per million indexed log events | Per active metric series, per GB of logs and traces | Per event (each span is an event) |
| Free tier | 5 hosts, 1-day metric retention | 10k series, 50 GB logs, 50 GB traces, 14-day retention | 20M events per month |
| Entry paid price | $15 per host per month (Infra Pro, annual) | $19 per month platform fee plus usage | From $150 per month |
| High-cardinality tags | Each tag combination is a separate custom metric | Each label combination is a separate active series | Not billed by number of unique values |
| Default retention (paid) | 15 months metrics; log retention depends on index tier | 13 months metrics, 30 days logs and traces | 60 days for most customers |
| OpenTelemetry | Supported via Datadog Agent, OTel Collector, or direct OTLP | Native OTLP ingestion | OpenTelemetry supported on all plans |
| Self-host option | No | Yes, the same open-source stack (AGPLv3) | No |
How bills actually grow
Startups rarely get surprised by the per-host line. The surprises come from three places.
Custom metrics and cardinality. Datadog counts a custom metric as each unique combination of metric name and tag values, including the host tag. Suppose you track checkout.latency tagged with 50 endpoints, 5 status codes, and 20 hosts. That one metric can produce up to 5,000 custom metrics, far past the 100 per host that Pro includes. Extra custom metrics are billed per 100, at a rate that depends on your contract. Grafana Cloud has the same multiplication under a different name: its own docs show a single CPU metric across 6 modes, 10 hosts, and 4 CPUs becoming 240 active series. Honeycomb is the outlier here. Its docs state it does not bill time series metrics by the number of unique series, and events are billed by count, not by how many fields or distinct values they contain. That is the main reason teams pick it for debugging high-cardinality problems.
Log volume. Logs scale with traffic and with how chatty your code is. On Datadog you pay once to ingest and again to index, so a debug-level logger left on in production costs you twice. On Grafana Cloud you pay to process, write, and retain each GB. A good habit: treat every log line as a recurring cost, because it is.
Traces at full fidelity. Sending every span from every request is the easiest way to blow through event and span allowances. On Honeycomb, a busy service with deep traces burns events fast. Sampling, covered below, is the fix.
OpenTelemetry is your exit door
OpenTelemetry (often shortened to OTel) is a Cloud Native Computing Foundation project that gives you vendor-neutral APIs, SDKs, and a Collector for traces, metrics, and logs. It is explicitly not a backend: it generates and ships data, and you point it at whichever backend you like.
All three vendors accept OTel data. Datadog supports it through its own Agent, through a standard OpenTelemetry Collector exporting over OTLP (the OpenTelemetry wire protocol), or by direct OTLP ingestion. Grafana Cloud ingests metrics, logs, and traces over OTLP natively. Honeycomb lists OpenTelemetry support on every plan.
The practical advice: instrument your code with OTel SDKs, run an OTel Collector between your apps and the vendor, and keep vendor-specific agents and libraries to a minimum. Then switching backends, or sending data to two of them during a trial, is a Collector config change. Dashboards and alerts still have to be rebuilt, so that is where your remaining lock-in lives. Write the choice down as a technical decision record so the next engineer knows why.
Self-hosting Grafana vs. Grafana Cloud
The Grafana stack (Grafana for dashboards, Mimir for metrics, Loki for logs, Tempo for traces) is open source under the AGPLv3 license. AGPL is fine for running internally, but have your lawyer look at it if you plan to modify and offer it as a service to others.
Self-hosting removes per-series and per-GB fees, but you take on object storage, scaling, upgrades, and the pager when your monitoring system is the thing that’s down. For most startups, Grafana Cloud’s free tier (10,000 series and 50 GB of logs) covers the early stage at no cost, and the bill becomes meaningful only once you have the traffic, and usually the team, to judge self-hosting properly. This is a classic build vs. buy call: self-host when the bill clearly exceeds the cost of an engineer’s time to run it, not before.
Alternatives worth knowing
Prices in this section are as of October 2026.
- Sentry focuses on errors and performance. The free Developer plan covers 1 user, 5,000 errors, and 5 million spans per month; Team is $26 per month and Business $80 per month (annual billing) with 50,000 errors. Many startups pair Sentry for errors with a cheaper metrics backend.
- Axiom is built around cheap log and event ingestion. Its free Personal plan includes 500 GB of ingest per month with 30-day retention; the paid Axiom Cloud plan is a $25 monthly platform fee with 1 TB of ingest included.
- SigNoz is an open-source, OpenTelemetry-native platform covering logs, metrics, and traces. The self-hosted Community Edition is free (MIT-licensed outside its enterprise directories). SigNoz Cloud starts at $49 per month including $49 of usage, then $0.30 per GB for logs and traces (15-day retention) and $0.10 per million metric samples.
Which one for you
Solo founder, one or two services. Grafana Cloud free plus Sentry free will cover you for a long time. Honeycomb’s free 20 million events is also enough for a small app if you sample. Avoid paying per host before you have customers.
Seed to Series A startup, a handful of services. This is where the choice matters. If your team wants one tool and will enforce discipline on tags and log indexing, Datadog is productive and well integrated. If you want cost predictability and portable formats, Grafana Cloud Pro. If your incidents are mostly “some users are slow and we don’t know why,” Honeycomb answers those questions faster than metric dashboards do.
Growing company with heavy log volume. Look hard at log-specific pricing. Route only what you search to an indexed tier, archive the rest cheaply, and compare Axiom or a self-hosted Loki for the bulk.
Regulated company. Check data residency and retention controls before price. Region and compliance options vary by vendor and by plan, and they change, so get them in writing from each vendor’s current documentation. Self-hosting gives you full control of where data lives, at the cost of running it yourself.
Practical cost controls
- Sample traces. Head sampling decides at the start of a request (for example, keep 5% of traces). Tail sampling decides after the trace completes, so you can keep every trace with an error or high latency and drop the boring ones. Tail sampling runs in the OTel Collector and needs memory to hold traces while it decides.
- Set log levels per environment. Production should default to
infoorwarn. Makedebugsomething you turn on for a service temporarily. - Index less than you ingest. On Datadog, use exclusion filters and Flex storage for logs you keep for audit but rarely search.
- Ban unbounded tags. No user IDs, request IDs, or raw URLs as metric tags. Put those on traces or events, where they belong.
- Shorten retention where nobody looks. Most debugging happens within days. Pay for long retention only on data you actually query later.
- Set budget alerts. Every vendor shows usage. Check it weekly while you’re growing.
Mistakes to avoid
- Instrumenting with a vendor’s proprietary SDK everywhere. It makes the first week easier and every future migration harder.
- Treating the free tier as a long-term plan. Know what happens at the limit: dropped data, a blocked account, or automatic charges.
- Adding a tag “just in case.” On per-series pricing, one tag can multiply your metric bill.
- Self-hosting to save money before you have the volume. The first incident where your monitoring is down while production is down will cost more than the subscription.
- Comparing list prices without your own data. Export a week of real telemetry volume (hosts, series, GB of logs, spans) and price it on each vendor’s calculator.
A quick checklist
- Count hosts, active metric series, GB of logs per month, and spans per month today.
- Instrument with OpenTelemetry and run a Collector.
- Decide which tags are allowed on metrics and write it down.
- Set production log level to
infoor higher. - Turn on trace sampling, with errors always kept.
- Pick retention per signal based on how far back you really look.
- Set a monthly budget alert with each vendor.
- Record the decision and its assumptions so you can revisit it at the next funding stage.
Sources
- Datadog pricing
- Datadog log management pricing
- Datadog docs: Custom metrics billing
- Datadog docs: OpenTelemetry in Datadog
- Grafana Cloud pricing
- Grafana Cloud docs: Metrics invoice and active series
- Grafana Cloud docs: Send data using OTLP
- Grafana Labs licensing
- Grafana Mimir on GitHub (license)
- Honeycomb pricing
- Honeycomb docs: 2026 Pro plan changes
- Honeycomb docs: How Honeycomb calculates usage
- Honeycomb docs: Security and data protection
- OpenTelemetry: What is OpenTelemetry?
- OpenTelemetry: Sampling
- Sentry pricing
- Axiom pricing
- SigNoz pricing
- SigNoz on GitHub (README and license)