Datadog monitoring and alerting
Datadog is what you buy when you would rather not operate a monitoring stack. Metrics, traces and logs arrive correlated out of the box, which shortens incidents considerably — for a price that needs watching.
Where we use Datadog
Datadog is part of the stack on this service. Each page covers how we work, what you get and what it costs to start.
Datadog in practice
Datadog’s proposition is that metrics, traces and logs arrive already correlated. During an incident that is worth a great deal: you see a latency spike, click into the traces behind it, and read the logs for the specific slow request — without moving between three tools and correlating timestamps by hand.
The counterweight is cost, and it grows in ways that are easy to miss: per host, per custom metric, per gigabyte of logs indexed. Configured deliberately — sampling, log filtering, retention tiers — it is reasonable. Configured by default, it produces the invoice that makes people consider moving to Prometheus.
What we build with Datadog
Infrastructure, application and log data in one place, correlated by trace so a slow request can be followed end to end.
Request paths across services, which is how you find the one call responsible for a latency spike.
Monitors tied into the on-call rotation, tuned to reduce noise rather than maximise coverage.
Is Datadog right for you?
Ask usA good fit when
- Teams that want full observability working this week
- Distributed systems where tracing across services is the need
- Organisations preferring to buy rather than operate monitoring
- Environments with many technologies needing integrations
Probably not when
- Cost-sensitive projects at meaningful scale
- Teams with existing Prometheus and Grafana expertise
- Very simple applications where the tooling exceeds the need
What we run alongside Datadog
The rest of the setup, and why each piece is there. We keep this list short on purpose — every dependency is something someone has to maintain.
- APM tracing
- Where most of the value is — code-level visibility into slow endpoints.
- Log pipelines
- Filtering and sampling before indexing, which is where log cost is controlled.
- Monitors and SLOs
- Alerting tied to user-facing objectives rather than raw thresholds.
- Terraform provider
- Monitors and dashboards as code, so they are reviewed and not lost.
Why Datadog
Let’s talkCorrelated by default
Jumping from a metric spike to the traces and logs behind it is one click, not three tools.
Integrations already exist
Hundreds of supported technologies, so instrumentation is mostly configuration rather than development.
Strong APM
Code-level visibility into slow endpoints and database calls without building the pipeline yourself.
What we get called in to fix
Get a second opinionRunaway bills
Unbounded custom metrics and every log indexed. Almost always reducible substantially without losing signal.
Noisy monitors
Alerts inherited from defaults that nobody tuned, until the team mutes the channel.
Half-instrumented services
Some services traced and others not, which breaks exactly the correlation you paid for.
Dashboards clicked together
No definitions in code, so nothing survives a reorganisation.
Datadog or the alternative
The comparisons we are actually asked to make, answered the way we would answer them on a call.
Datadog to buy it working now. The open pair for cost control and ownership, at the price of running it.
Comparable products. Usually decided by pricing model and which integrations matter to you.
Datadog for correlation with traces. A dedicated stack when log volume makes indexing cost dominate.
Got an idea? Let’s make it real.
Tell us the short version
This could be the first step towards a new and successful collaboration. A one-line idea and a finished spec are both fine — tell us the problem, the deadline you’re working to and what’s in your way.
Keep looking
Frequently asked questions
It can be, and it grows with hosts, custom metrics and log volume. We configure retention and sampling deliberately.
Datadog when you want it working this week with no stack to run; the open-source pair when cost control and ownership matter more.
Often — log sampling, index filters and pruning unused custom metrics are the usual places the money goes.
Usually — log sampling and index filters, pruning unused custom metrics, and right-sizing host counts are the recurring wins.
Only if the cost genuinely justifies the operational burden. Running your own observability stack is not free either.