{"id":64,"date":"2026-10-08T14:05:18","date_gmt":"2026-10-08T06:05:18","guid":{"rendered":"https:\/\/hoho.live\/luda\/index.php\/2026\/10\/08\/how-last_over_time-stops-a-scrape-gap-from-firing-a-grafana-disk-alert\/"},"modified":"2026-10-08T14:05:18","modified_gmt":"2026-10-08T06:05:18","slug":"how-last_over_time-stops-a-scrape-gap-from-firing-a-grafana-disk-alert","status":"publish","type":"post","link":"https:\/\/hoho.live\/luda\/index.php\/2026\/10\/08\/how-last_over_time-stops-a-scrape-gap-from-firing-a-grafana-disk-alert\/","title":{"rendered":"How last_over_time stops a scrape gap from firing a Grafana disk alert"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Short answer:<\/strong> A mail whose subject starts with <strong>DatasourceNoData<\/strong> is not a disk measurement. The <strong>Grafana<\/strong> rule asked <strong>Prometheus<\/strong> for the latest free-space sample, the scrape had just hit <strong>scrape_timeout<\/strong>, and Prometheus had marked that series stale, so the instant query returned no series. Grafana sends the NoData state on that evaluation. The rule&#8217;s <code>for<\/code> wait applies to the threshold, not to NoData. Query <code>last_over_time<\/code> of the free-space ratio over a few minutes, and set <code>noDataState<\/code> to <code>KeepLast<\/code>. A separate <code>up == 0<\/code> rule is what reports an exporter that stays down.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>What the mail is reporting<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">Grafana&#8217;s default title groups by folder and <code>alertname<\/code>. On an empty result it sets <code>alertname<\/code> to <code>DatasourceNoData<\/code> and keeps the rule title in <code>rulename<\/code>. The summary annotation is text stored on the rule, so the body still says the filesystem is almost full. That sentence was written when the rule was written. Prometheus did not return a free-space number in that evaluation.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>What a timed-out scrape does to the series<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">With <code>scrape_interval<\/code> at 30s and <code>scrape_timeout<\/code> left at its 10s default, a scrape that reaches the deadline stores <code>up<\/code> 0, <code>scrape_samples_scraped<\/code> 0, and <code>scrape_duration_seconds<\/code> at about 10. Prometheus then writes <strong>staleness markers<\/strong> for the series that target had been exporting. An instant selector evaluated on a marker returns no series. The previous sample is still in the TSDB. A marker is not a sample, so <code>last_over_time<\/code> over a range that contains it returns that sample.<\/p>\n\n\n<p class=\"wp-block-paragraph\">The exporter is not what used the 10 seconds. In the minute this fired, a blackbox probe counted as a 10s scrape ran inside the exporter in tens of milliseconds once the HTTP request arrived. Requests whose client had already closed were logged as context canceled. Exporters reached by other paths stalled in the same seconds. The Prometheus host&#8217;s CPU, disk, and IKE SA were quiet, and its TCP retransmit counter rose. The scrapes were waiting on packets. Which hop dropped them is not identified, and the alert fix does not need that hop.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>The working query<\/strong><\/h4>\n\n<ol class=\"wp-block-list\">\n<li>Keep the threshold: available bytes divided by size, under 0.20, <code>for<\/code> 10m.<\/li>\n<li>Replace the instant ratio with <code>min(last_over_time((avail \/ size)[5m:]))<\/code>, same matchers as before. An empty subquery step uses the evaluation interval.<\/li>\n<li>Set <code>noDataState: KeepLast<\/code>. A gap longer than the range then holds the last state instead of mailing <code>DatasourceNoData<\/code>.<\/li>\n<li>Add <code>up == 0<\/code> for that scrape job, <code>for<\/code> 10m, with its own summary. That mail means the exporter is not being read.<\/li>\n<\/ol>\n\n\n<p class=\"wp-block-paragraph\">At the timestamp where the instant ratio was empty, <code>last_over_time<\/code> over 5 minutes returned the previous ratio. A disk that is actually under 20% still has to stay there for the pending period. Use a range longer than one scrape interval and shorter than the time you will accept a stale gauge.<\/p>\n\n\n<p class=\"wp-block-paragraph\">This is for a gauge whose last good sample is still the right input, such as disk, memory, or inodes. A <code>probe_success<\/code> or <code>up<\/code> rule is the opposite: a missing sample is the event.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>Approaches that look simpler and fail<\/strong><\/h4>\n\n<figure class=\"wp-block-table\">\n<table>\n<thead>\n<tr>\n<th>Approach<\/th>\n<th>Why it fails when one scrape times out<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Trust the <code>for: 10m<\/code> already on the rule<\/td>\n<td>That wait applies before a threshold breach becomes Alerting. NoData is sent on the empty evaluation.<\/td>\n<\/tr>\n<tr>\n<td>Raise <code>scrape_timeout<\/code> above 10s<\/td>\n<td>A response can arrive after the new deadline too. The rule is still one instant sample wide.<\/td>\n<\/tr>\n<tr>\n<td>Set <code>noDataState<\/code> to OK<\/td>\n<td>An outage longer than the range looks healthy, including a full disk you can no longer see.<\/td>\n<\/tr>\n<tr>\n<td>Set <code>noDataState<\/code> to Alerting<\/td>\n<td>Every gap pages as if the threshold had fired.<\/td>\n<\/tr>\n<tr>\n<td>Read the summary annotation as the measured value<\/td>\n<td>The annotation is constant. DatasourceNoData attaches it unchanged.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n\n\n<p class=\"wp-block-paragraph\">A longer scrape timeout fits a target that is slow and still answers. It does not define an empty result.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>When you can drop last_over_time<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">Drop it only if a missed scrape should page, or if the number already comes from a recording rule that used a range. Those are choices about what an empty result means, not a hidden Prometheus flag. Keep <code>last_over_time<\/code> while one timed-out scrape can stale the gauge and NoData sends mail.<\/p>\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<h4 class=\"wp-block-heading\"><strong>Why this stays published<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">We hit this when a root-disk rule mailed that the filesystem was almost full while it was about 60% free. <strong>A Prometheus scrape timeout stales the series, Grafana mails DatasourceNoData immediately, and last_over_time of the ratio is what keeps that gap from becoming a disk alert.<\/strong><\/p>\n\n\n<p class=\"wp-block-paragraph\">\u672c\u6587\u7531 HoHo \u8207 AI \u5354\u4f5c\u6574\u7406\uff0c\u6700\u5f8c\u66f4\u65b0\u65bc 2026 \u5e74 10 \u6708 8 \u65e5\u3002<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Short answer: A mail whose subject starts with DatasourceNoData is not a disk measurement. The Grafana rule asked Prometheus for the latest free-space sample, the scrape had just hit scrape_timeout, and Prometheus had marked that series stale, so the instant query returned no series. Grafana sends the NoData state&hellip;<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[26,7,25],"tags":[],"class_list":["post-64","post","type-post","status-publish","format-standard","hentry","category-grafana","category-homelab","category-prometheus"],"_links":{"self":[{"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/posts\/64","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/comments?post=64"}],"version-history":[{"count":0,"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/posts\/64\/revisions"}],"wp:attachment":[{"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/media?parent=64"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/categories?post=64"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/hoho.live\/luda\/index.php\/wp-json\/wp\/v2\/tags?post=64"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}