Availability 7 min read

Uptime monitoring with traces and HTTP checks

Find out within minutes when a service stops answering or starts failing requests. Alert on the traces your services already send, then add OpenTelemetry Collector HTTP checks for the endpoints and certificates that real traffic cannot cover.

By JeremyFunk

01 · Definition

Uptime monitoring probes an endpoint from outside on a schedule

Uptime monitoring sends a scheduled request to a public endpoint from outside your infrastructure and records whether it answered. A check asserts an HTTP status and usually a response time limit. When enough consecutive checks fail, it notifies someone.

The same tools watch the days left until the TLS certificate expires and whether the DNS record resolves. Hosted products run each check from several regions, so a network problem near one prober does not page you.

Availability as a number

The result is reported as availability: the time the endpoint answered, divided by the total time. Each added nine cuts the downtime you are allowed by a factor of ten.

AvailabilityDowntime allowed in 30 days
99% 7 h 12 min
99.9% 43 min 12 s
99.95% 21 min 36 s
99.99% 4 min 19 s

Detection time comes out of that allowance. A probe that runs every 60 seconds and needs two failed checks before it alerts takes up to two minutes to notice an outage. Against a 99.99% target, that is close to half the month's allowance spent before anyone is paged.

02 · Coverage

A passing uptime check covers one URL once a minute

A probe proves that a single request to a single URL succeeded. Most teams point it at /health, which returns 200 for as long as the process is running.

Take a service that handles 10 requests per second. A deploy makes 2% of them fail. Over a day the service fails 17,280 real requests, while the probe on /health sends 1,440 requests and every one of them passes.

17,280real requests failed that day
1,440probe checks sent to /health
0probe checks failed

Pointing the probe at the broken route does not change much. Each check has a 2% chance of failing, so two consecutive failures come up about once in 2,500 pairs of checks. At one check a minute, the alert fires roughly every 42 hours while the incident fails 12 requests a minute.

A healthy status code from an unhealthy service

A health endpoint that returns 200 while the database is unreachable passes every check. Asserting on the response body narrows that gap, and the probe still tests only the path you scripted.

03 · Real traffic

Traces turn every real request into an availability check

A service instrumented with OpenTelemetry records a span for every request it handles, with a duration and a status. That is the same pass or fail result a probe produces, recorded for every route and every real user.

Signals to alert on

Error rate

The share of requests that ended in an error. It catches the partial failure a probe samples past.

Throughput

The number of requests in the window. When a service stops answering, its spans stop, and a drop to zero is the outage alert.

Apdex

The share of requests that were fast enough and succeeded. It catches the outage where every response is still a 200 and takes four seconds. See What is Apdex?

Because the alert is computed from spans, the failing traces are already stored when it fires. A failed probe gives you a URL and a timestamp.

Failures compared: outside probe and trace alerts

FailureProbe on /healthAlerts on traces
2% of checkout requests return 500 after a deploy Passes Error rate fires
A dependency slows every request to 4 s Passes, unless it has a latency limit Apdex and p95 fire
The process crashes and the load balancer returns 502 Fails Throughput drops to zero and fires
The TLS certificate expires Fails, and can warn weeks ahead Throughput drops after expiry and an alert fires
A DNS change points the domain nowhere Fails Throughput drops and an alert fires
An outage at 3am on a service with no night traffic Fails Nothing to measure
Users in one region cannot reach you Fails only from a prober in that region A partial dip, often under the threshold

Alerts Alerts in some cases Stays silent

A probe misses the common incidents: partial errors and slowdowns inside the application.

Traces have a narrower gap. They cannot see a failure that happens before the request reaches your code, and they have nothing to measure when no requests arrive.

04 · Outside probes

Failures only an outside uptime probe can see

Our position is that uptime monitoring is a narrow tool that gets bought by default. It is often the first monitoring a team sets up, because it takes a URL and five minutes. It then stays in place long after the service has outgrown what one request a minute can describe.

The cost is a false sense of security. A status board that shows 100% for the month proves that /health answered 43,200 times. It is easy to read that as proof the service was healthy, and the 2% of failed requests from the earlier example never appear on it. A team that relies on the board alone hears about that incident from its users.

A probe is the right addition when one of these applies:

01

Traffic is low or stops overnight

With no requests there are no spans. A probe generates the traffic the alert needs.

02

Certificates and DNS

An expiry date is known weeks ahead, and only something that reads the certificate can warn you before it passes.

03

Edge, CDN and load balancer failures

A request rejected in front of your service never produces a span.

04

Reachability by region

A routing or resolver problem in one region takes probers in several regions to see, each with its own DNS resolution.

05

Evidence for an SLA

Customers and contracts often expect an availability figure measured by a third party, or a public status page.

06

Endpoints you do not instrument

A vendor API or a legacy service with no telemetry can still be probed.

05 · In Maple

Availability alerts from traces in Maple

Maple alerts on the traces your services already send, and it accepts check results from the OpenTelemetry Collector as metrics. It does not run hosted probes of its own.

The uptime monitoring docs walk through each rule below with screenshots.

  1. 01

    Alert on traffic stopping

    A Throughput drop rule scoped to a service fires when its request count falls below a threshold. A window with no requests counts as zero, so a full outage fires it.

  2. 02

    Alert on failing requests

    A High error rate rule fires above 5% over five minutes, and opens one incident per failing service.

  3. 03

    Alert on slow requests

    A Low Apdex score rule scores errors and slow requests together against a 500ms target.

  4. 04

    Replay the rule against last week

    The rule form replays a rule over a past range and shades the periods where it would have held an incident open. A threshold that would have fired on a normal Tuesday needs to move.

The rule preview chart over the last hour: throughput near 200 requests per window, a drop to zero for about 15 minutes, and a shaded band marked Would have fired 1 time, longest 25 minutes.
A Throughput drop rule replayed over the last hour. Requests stopped for about 15 minutes, and the shaded band marks the period the rule would have held an incident open.

A throughput rule needs traffic to drop from. For a service that sits idle for hours, use an HTTP check instead.

06 · HTTP checks

HTTP uptime checks with the OpenTelemetry Collector

For the endpoints that need a probe, the Collector's http_check receiver requests a list of URLs on an interval and reports each result as metrics: the status class of the response, the request duration, connection errors, and the seconds left on the TLS certificate. Maple stores them like any other metric, so you can chart them and alert on them.

The Collector config and the settings for each rule are in the uptime monitoring docs.

Rules to build on the check metrics

Endpoint down

No check of a URL returned a 2xx in the last two minutes. One incident per URL, about three minutes into an outage.

Checks stopped

No check results arrived at all, which means the Collector itself is down.

Certificate

Fewer than 14 days are left before a certificate expires.

The rule preview chart over the last 30 minutes, grouped by URL. One URL drops from 1 to 0 for about five minutes inside a shaded band marked Would have fired 1 time, longest 8 minutes.
The endpoint rule replayed over 30 minutes of checks. One endpoint returned 503 for five checks, and the rule would have fired once.

One Collector is one vantage point

This setup probes from one place, through one DNS resolver. It cannot tell you that users in another region are cut off, and a network problem next to the Collector looks the same as an outage. It has no status page.

If you need probes from several regions, a public status page, or an availability report for customers, run a dedicated uptime product next to Maple. Those products operate prober fleets in many regions, and a single Collector cannot stand in for that.

Set it up in Maple

Collector config and alert rules, step by step

The docs page has the Collector config, the settings for each rule above, and screenshots of the rule forms.

Sources and further reading

FAQ

よくある質問

Does Maple have uptime monitoring?
Maple does not run hosted uptime probes. It alerts on error rate, throughput and Apdex computed from the traces your services send, and it stores the results of the OpenTelemetry Collector's http_check receiver as metrics you can chart and alert on. For probes from several regions or a public status page, run a dedicated uptime product next to Maple.
Do I need an uptime monitor if I already have APM or tracing?
For a service with steady traffic, alerts on error rate and throughput cover most of what an uptime monitor is bought for, and they catch partial failures a probe samples past. You still need an outside probe for endpoints with little traffic, for certificate and DNS problems, for failures at the CDN or load balancer, and when availability has to be measured by a third party.
What is the difference between uptime and availability?
Uptime is the share of time a system was running and reachable. Availability is the share of time, or of requests, in which it did its job for users. A process can be up while every checkout request fails, so availability measured on real requests is the stricter of the two numbers.
How often should an uptime check run?
Every 60 seconds is the usual default, including for the OpenTelemetry Collector's http_check receiver. Time to detect an outage is the interval multiplied by the number of failed checks required before alerting, so a 60 second interval with two required failures takes up to two minutes. Shorter intervals detect faster and add load on the endpoint.
Can OpenTelemetry do uptime checks?
Yes. The Collector Contrib distribution includes the http_check receiver, which requests a list of URLs on an interval and reports status, duration, errors and TLS certificate expiry as metrics. It checks from wherever the Collector runs, so one Collector gives you one vantage point.

今日、最初のトレースを。

SDKを追加し、OTLPエクスポーターをMapleに向ければトレースが届きます。

maple.dev:OpenTelemetryの上のオブザーバビリティ