Short answer: A mail whose subject starts with DatasourceNoData is not a disk measurement. The Grafana rule asked Prometheus for the latest free-space sample, the scrape had just hit scrape_timeout, and Prometheus had marked that series stale, so the instant query returned no series. Grafana sends the NoData state on that evaluation. The rule’s for wait applies to the threshold, not to NoData. Query last_over_time of the free-space ratio over a few minutes, and set noDataState to KeepLast. A separate up == 0 rule is what reports an exporter that stays down.
What the mail is reporting
Grafana’s default title groups by folder and alertname. On an empty result it sets alertname to DatasourceNoData and keeps the rule title in rulename. The summary annotation is text stored on the rule, so the body still says the filesystem is almost full. That sentence was written when the rule was written. Prometheus did not return a free-space number in that evaluation.
What a timed-out scrape does to the series
With scrape_interval at 30s and scrape_timeout left at its 10s default, a scrape that reaches the deadline stores up 0, scrape_samples_scraped 0, and scrape_duration_seconds at about 10. Prometheus then writes staleness markers for the series that target had been exporting. An instant selector evaluated on a marker returns no series. The previous sample is still in the TSDB. A marker is not a sample, so last_over_time over a range that contains it returns that sample.
The exporter is not what used the 10 seconds. In the minute this fired, a blackbox probe counted as a 10s scrape ran inside the exporter in tens of milliseconds once the HTTP request arrived. Requests whose client had already closed were logged as context canceled. Exporters reached by other paths stalled in the same seconds. The Prometheus host’s CPU, disk, and IKE SA were quiet, and its TCP retransmit counter rose. The scrapes were waiting on packets. Which hop dropped them is not identified, and the alert fix does not need that hop.
The working query
- Keep the threshold: available bytes divided by size, under 0.20,
for10m. - Replace the instant ratio with
min(last_over_time((avail / size)[5m:])), same matchers as before. An empty subquery step uses the evaluation interval. - Set
noDataState: KeepLast. A gap longer than the range then holds the last state instead of mailingDatasourceNoData. - Add
up == 0for that scrape job,for10m, with its own summary. That mail means the exporter is not being read.
At the timestamp where the instant ratio was empty, last_over_time over 5 minutes returned the previous ratio. A disk that is actually under 20% still has to stay there for the pending period. Use a range longer than one scrape interval and shorter than the time you will accept a stale gauge.
This is for a gauge whose last good sample is still the right input, such as disk, memory, or inodes. A probe_success or up rule is the opposite: a missing sample is the event.
Approaches that look simpler and fail
| Approach | Why it fails when one scrape times out |
|---|---|
Trust the for: 10m already on the rule |
That wait applies before a threshold breach becomes Alerting. NoData is sent on the empty evaluation. |
Raise scrape_timeout above 10s |
A response can arrive after the new deadline too. The rule is still one instant sample wide. |
Set noDataState to OK |
An outage longer than the range looks healthy, including a full disk you can no longer see. |
Set noDataState to Alerting |
Every gap pages as if the threshold had fired. |
| Read the summary annotation as the measured value | The annotation is constant. DatasourceNoData attaches it unchanged. |
A longer scrape timeout fits a target that is slow and still answers. It does not define an empty result.
When you can drop last_over_time
Drop it only if a missed scrape should page, or if the number already comes from a recording rule that used a range. Those are choices about what an empty result means, not a hidden Prometheus flag. Keep last_over_time while one timed-out scrape can stale the gauge and NoData sends mail.
Why this stays published
We hit this when a root-disk rule mailed that the filesystem was almost full while it was about 60% free. A Prometheus scrape timeout stales the series, Grafana mails DatasourceNoData immediately, and last_over_time of the ratio is what keeps that gap from becoming a disk alert.
本文由 HoHo 與 AI 協作整理,最後更新於 2026 年 10 月 8 日。
