Skip to content

Add native maintenance timing and visibility-map statistics - #2104

Draft
Alena0704 wants to merge 16 commits into
apache:mainfrom
Alena0704:vacuum_stats_03_native_timing_vm
Draft

Alena0704 wants to merge 16 commits into
apache:mainfrom
Alena0704:vacuum_stats_03_native_timing_vm

Conversation

@Alena0704

Copy link
Copy Markdown
Collaborator

What does this PR do?

Expose maintenance duration and visibility-map stability through native
cumulative statistics and catalog views.

  • Track VACUUM and ANALYZE elapsed time, index vacuum time, and database
    vacuum totals.
  • Measure cost-based delays when track_cost_delay_timing is enabled and
    expose delay time in progress reporting.
  • Count completed vacuums that entered failsafe mode and vacuums interrupted
    by errors.
  • Track cleared all-visible and all-frozen VM bits.
  • Add pg_class.relallfrozen and maintain it through vacuum, analyze and
    relation lifecycle operations.
  • Measure active AO vacuum phases, excluding gaps between phases.

Expose the statistics through SQL getters and local, per-instance and cluster
summary views. Catalog exposure and tests are kept in a separate commit.

Test Plan

  • Added SQL and isolation coverage for VM and AO behavior.

Impact

Changes the pg_class layout and bumps the catalog version; development
clusters require initdb. Updates the cumulative-statistics file format.

Type of Change

  • Bug fix (non-breaking change)
  • New feature (non-breaking change)
  • Breaking change (fix or feature with breaking changes)
  • Documentation update

Breaking Changes

Test Plan

  • Unit tests added/updated
  • Integration tests added/updated
  • Passed make installcheck
  • Passed make -C src/test installcheck-cbdb-parallel

Checklist

michaelpq and others added 16 commits October 8, 2026 15:00
…alyze

This commit adds four fields to the statistics of relations, aggregating
the amount of time spent for each operation on a relation:
- total_vacuum_time, for manual vacuum.
- total_autovacuum_time, for vacuum done by the autovacuum daemon.
- total_analyze_time, for manual analyze.
- total_autoanalyze_time, for analyze done by the autovacuum daemon.

This gives users the option to derive the average time spent for these
operations with the help of the related "count" fields.

Bump PGSTAT_FILE_FORMAT_ID for the additions in PgStat_StatTabEntry.
Catalog functions, views and their version bump follow separately.

Author: Sami Imseih
Reviewed-by: Bertrand Drouvot, Michael Paquier
Discussion: https://postgr.es/m/CAA5RZ0uVOGBYmPEeGF2d1B_67tgNjKx_bKDuL+oUftuoz+=Y1g@mail.gmail.com

(cherry picked from commit 30a6ed0)


Cloudberry adaptation: time AO vacuum from its first phase and pass the
start timestamp to the collector at final cleanup, so the changed API is
usable by heap and AO immediately.

Co-authored-by: Alena Rybakina <alenka.rybakina@gmail.com>
This function is used in both vacuum and analyze code paths, and a
follow-up commit will require distinguishing between the two.  This
commit forces callers to specify whether they are in a vacuum or
analyze path, but it does not use that information for anything
yet.

Author: Nathan Bossart <nathandbossart@gmail.com>
Co-authored-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal

Backported from PostgreSQL 18 (commit e5b0b0c)
as a prerequisite of the vacuum statistics series.  Cloudberry adaptation:
the append-optimized and AOCS sample-row acquisition and compaction call
vacuum_delay_point() too; the former pass is_analyze = true, the latter
false.
This commit adds the amount of time spent sleeping due to
cost-based delay to the pg_stat_progress_vacuum and
pg_stat_progress_analyze system views.  A new configuration
parameter named track_cost_delay_timing, which is off by default,
controls whether this information is gathered.  For vacuum, the
reported value includes the sleep time of any associated parallel
workers.  However, parallel workers only report their sleep time
once per second to avoid overloading the leader process.

Bumps catversion.

Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Co-authored-by: Nathan Bossart <nathandbossart@gmail.com>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Robert Haas <robertmhaas@gmail.com>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Reviewed-by: Masahiro Ikeda <ikedamsh@oss.nttdata.com>
Reviewed-by: Dilip Kumar <dilipbalaut@gmail.com>
Reviewed-by: Sergei Kornilov <sk@zsrv.org>
Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal

Backported from PostgreSQL 18 (commit bb8dff9)
as a prerequisite of the vacuum statistics series.  Cloudberry adaptation:
the vacuum progress parameters of this branch end at num_dead_tuples, so
the delay time is PROGRESS_VACUUM_DELAY_TIME 7 (param8 of the view).  A
parallel vacuum worker only accumulates its sleep time: reporting it to the
leader needs pgstat_progress_parallel_incr_param() of PostgreSQL 17, and
Cloudberry does not run parallel vacuum at all (see dead_items_alloc()).
The new GUC is listed in unsync_guc_names_array, next to track_io_timing,
as every GUC has to be in one of the two lists here.



SQL progress-view changes are separated into the native SQL reporting
commit; measurement, the GUC and progress counters are introduced here.
Port the process-local VacuumDelayTime accumulator from the REL_2_STABLE
commit. Accumulate measured cost-based VACUUM sleep time only while
track_cost_delay_timing is enabled, excluding ANALYZE. Consumers take
before/after differences to attribute delays to their own operations.

The preceding PostgreSQL 18 progress-view backport already provides the
GUC, its documentation, sample setting, unsynchronized registration and
sleep timing. Reuse that measurement and preserve both VACUUM and ANALYZE
progress reporting. Keep the accumulator in milliseconds as double to
match this branch's cumulative timing consumers, rather than the source
branch's int64 microseconds.

Extract the accumulator from the index/database timing commit and fold
in the ANALYZE exclusion previously included in the combined heap/index/AO
measurements commit. The final branch contents are unchanged.

Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal
Backported from PostgreSQL 18 (commit bb8dff9).
(cherry picked from commit 01aaacb)

Co-authored-by: Nathan Bossart <nathandbossart@gmail.com>
Co-authored-by: Alena Rybakina <alenka.rybakina@gmail.com>
Accumulate elapsed time and cost-based delays over active AO vacuum phases,
excluding the gaps between phases. Discard state left over from an
interrupted operation. Prepare per-index timing measurements for the
subsequent cumulative-reporting change.

Adapted-from: 6f38fb2
Record index-pass elapsed and cost-delay time, including parallel workers
with the leader's autovacuum classification. Add cumulative table delay
time and database totals without adding index time to database totals a
second time. Include removed tuple counts in verbose index reports.

Catalog getters, views, documentation and SQL-dependent tests follow in
the separate native SQL reporting commit.
Accumulate relation and database failsafe counts for completed heap
vacuums. Pass false for AO, which has no heap wraparound failsafe mode.
Keep this counter distinct from scans made aggressive by the freeze age.

Catalog getters, views and tests follow in the native SQL reporting commit.
Count ERROR reports in the heap vacuum callback using local counters for database-local and shared relations. Transfer them to pending database statistics at transaction end. Test cancellation of a running vacuum.
Add visible_page_marks_cleared and frozen_page_marks_cleared counters to
pg_stat_all_tables tracking the number of times the all-visible and
all-frozen bits are cleared in the visibility map. These bits are cleared by
backend processes during regular DML operations. Hence, the counters are placed
in table statistic entry.

A high visible_page_marks_cleared rate relative to DML volume indicates
that modifications are scattered across previously-clean pages rather
than concentrated on already-dirty ones, causing index-only scans to
fall back to heap fetches.  A high frozen_page_marks_cleared rate indicates
that vacuum's freezing work is being frequently undone by concurrent
DML.

Authors: Alena Rybakina <lena.ribackina@yandex.ru>,
         Andrei Lepikhov <lepihov@gmail.com>,
         Andrei Zubkov <a.zubkov@postgrespro.ru>
Reviewed-by: Dilip Kumar <dilipbalaut@gmail.com>,
         Masahiko Sawada <sawada.mshk@gmail.com>,
         Ilia Evdokimov <ilya.evdokimov@tantorlabs.com>,
         Jian He <jian.universality@gmail.com>,
         Kirill Reshke <reshkekirill@gmail.com>,
         Alexander Korotkov <aekorotkov@gmail.com>,
         Jim Nasby <jnasby@upgrade.com>,
         Sami Imseih <samimseih@gmail.com>,
         Karina Litskevich <litskevichkarina@gmail.com>,
         Andrey Borodin <x4mmm@yandex-team.ru>

Backported-from: https://www.postgresql.org/message-id/attachment/205562/v45-0009-Track-table-VM-stability.patch

Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Connect AO index-pass timing measurements to cumulative statistics and
report interrupted vacuums. Add cumulative delay reporting for AO tables.
Failsafe and visibility-map counters remain inapplicable to AO.

The cluster isolation test is included later with its gp_stat_vacuum
dependency. Active-phase elapsed reporting is connected in a later commit.
Describe track_cost_delay_timing alongside the instrumentation it enables.
Give the accumulated instrumentation fields their own statistics-file format
identifier independent of the pluggable-statistics series.
Add relallfrozen, an estimate of the number of pages marked all-frozen
in the visibility map.

pg_class already has relallvisible, an estimate of the number of pages
in the relation marked all-visible in the visibility map. This is used
primarily for planning.

relallfrozen, together with relallvisible, is useful for estimating the
outstanding number of all-visible but not all-frozen pages in the
relation for the purposes of scheduling manual VACUUMs and tuning vacuum
freeze parameters.

A future commit will use relallfrozen to trigger more frequent vacuums
on insert-focused workloads with significant volume of frozen data.

Bump catalog version

Author: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Nathan Bossart <nathandbossart@gmail.com>
Reviewed-by: Robert Treat <rob@xzilla.net>
Reviewed-by: Corey Huinker <corey.huinker@gmail.com>
Reviewed-by: Greg Sabino Mullane <htamfids@gmail.com>
Discussion: https://postgr.es/m/flat/CAAKRu_aj-P7YyBz_cPNwztz6ohP%2BvWis%3Diz3YcomkB3NpYA--w%40mail.gmail.com
(cherry picked from commit 99f8f3f)

Cloudberry adaptation: carry both visibility-map counts through QE-to-QD reports, replicated-table normalization, ANALYZE and CLUSTER. AO and indexes keep relallfrozen zero. Avoid VM I/O on the dispatcher during index catalog updates. Include local and distributed catalog tests. PostgreSQL 18 statistics-import APIs and their tests do not exist on this PostgreSQL 16 branch.

Use relallfrozen in the VM isolation test. Hold a repeatable-read snapshot and explicitly populate the statistics cache before concurrent DELETE; flush in the writer to avoid depending on the flush interval.
Add an elapsed-time reporting entry point and connect the previously
introduced AO active-phase timers. Exclude gaps between phases from
cumulative vacuum elapsed and cost-delay time, and remove the obsolete
first-phase timing baseline. SQL exposure and tests follow separately.
Add native SQL getters and local/cluster views for maintenance elapsed and
cost-delay time, failsafe runs, interrupted vacuums and cleared VM bits.
Expose progress delay and test timing, VM snapshots and relallfrozen for
heap and append-optimized relations without a statistics extension.
ANALYZE reads relallvisible and relallfrozen together instead of dispatching two queries. Preserve replicated-table normalization and local visibility-map counting. Check both coordinator catalog counts against the segment totals in the distributed and replicated regression cases.
Move cluster timing, index-summary cardinality, and track_cost_delay_timing scope checks from the extension regression into the core vacuum_stats test. Cover distributed and replicated tables, and explicitly flush database statistics before checking their timing aggregate.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants