Skip to content

Add native VACUUM work statistics - #2105

Draft
Alena0704 wants to merge 9 commits into
apache:mainfrom
Alena0704:vacuum_stats_04_native_work
Draft

Alena0704 wants to merge 9 commits into
apache:mainfrom
Alena0704:vacuum_stats_04_native_work

Conversation

@Alena0704

Copy link
Copy Markdown
Collaborator

What does this PR do?

Make VACUUM work available through native cumulative statistics for tables,
indexes and databases.

Collect tuple removal, remaining and missed dead tuples, scanned and removed
pages, freezing work, newly all-visible pages, and freeze-age-driven vacuums.
For append-optimized tables, collect segment counts, completed compactions,
moved tuples and remaining hidden rows.

Report index work as per-call deltas and correct SP-GiST accounting for
previously empty pages. Database totals exclude separate index reports to
avoid double counting.

Store dead_tuples and total_file_segs as last-run samples; other work counters
accumulate. Use ordinary statistics snapshots, resets, DROP handling and
clean-shutdown persistence.

Measurement, cumulative storage, and catalog functions/views are introduced
in separate commits.

Impact

Makes VACUUM work statistics available for heap, AO/AOCO tables and indexes
without an extension. Native cumulative counters are updated when
track_counts is enabled and exposed through local and cluster-wide views.

Adds counter maintenance during VACUUM and increases relation and database
statistics entry sizes. Bumps the catalog and cumulative-statistics file
versions; development clusters initialized with the previous catalog version
require initdb.

Type of Change

  • Bug fix (non-breaking change)
  • New feature (non-breaking change)
  • Breaking change (fix or feature with breaking changes)
  • Documentation update

Test Plan

  • Unit tests added/updated
  • Integration tests added/updated
  • Passed make installcheck
  • Passed make -C src/test installcheck-cbdb-parallel

Checklist

Assemble backend-local tuple and page measurements and take per-call
snapshots of the existing index results. Access methods accumulate results
across passes, so retain each pass's delta without counting earlier work
again. Share the measurement helper between serial and parallel vacuum.

Keep SP-GiST's newly deleted page count cumulative within one vacuum,
without counting pages that were already empty on entry. Handle AMs that
reset their newly deleted page count between passes.

Adapt to PostgreSQL 16's existing recently_dead_tuples, missed_dead_tuples,
pages_scanned, pages_removed and missed_dead_pages counters. Assemble these
measurements independently of optional hooks; storage and SQL follow in
separate reporting commits.

Adapted-from: f2592a3
Introduce dead_pages in LVRelState and PgStat_VacuumStats. Count heap pages
that retain recently dead tuples in both pruning and no-prune scans.
Keep this measurement distinct from pages with dead tuples missed because
of buffer pins. For index cleanup, derive the count from deleted pages
that are not yet reusable.

Cumulative statistics storage and SQL exposure follow separately.

Adapted-from: 37a4753
Include pages_frozen and tuples_frozen in the backend-local VACUUM
measurements. PostgreSQL 16 already counts frozen_pages and tuples_frozen
and prints them in VACUUM VERBOSE, so reuse those counters without adding
a second increment at the freezing site.

Cumulative statistics storage and SQL exposure follow separately.

Adapted-from: 93ae811
Add the backend-local pages_all_visible measurement. Count new visibility
map marks under the heap page lock, including empty pages and second-pass
cleanup. Test the current VM bit, rather than a potentially stale skipping
decision, and do not recount existing marks or all-frozen-only upgrades.

Cumulative statistics storage and SQL exposure follow separately.

Adapted-from: 3aefc13
Remember the result of vacuum_get_cutoffs before DISABLE_PAGE_SKIPPING can
force aggressive mode. Include this classification in backend-local work
measurements, keeping freeze-age vacuums distinct from page-skipping
options and wraparound failsafe activation.

Cumulative statistics storage and SQL exposure follow separately.

Adapted-from: cefc6ba
Count completed compactions, moved tuples, scanned pages and the final
segment count without an extension hook. Discard state left over from an
interrupted operation. Keep per-index work measurements for each pass.
Native cumulative storage and SQL reporting follow separately.

Adapted-from: 6f38fb2
Connect the previously introduced heap, index and AO measurements to
ordinary relation and database statistics independently of any extension
hook. Include them in native snapshots, resets, transactional drops and
clean-shutdown persistence. Store dead_tuples and total_file_segs as
last-run samples; other work counters accumulate. Exclude index reports
and last-run samples from database totals.

Advance the statistics-file format for the stored fields. Catalog
functions, views and tests follow in a separate SQL reporting commit.
Publish cumulative tuple, page and AO compaction counters through native
SQL getters and local, per-segment and cluster summary views. Keep last-run
snapshots separate from cumulative work and normalize replicated relations.
Include heap/index/AO tests for counters, resets, snapshots and persistence.
Reset work for one relation or the current database without clearing timing, invocation counts, other native counters, or extension metrics. Add a privileged gp_stat_reset_vacuum_stats wrapper for all primary segments and cover scope, permissions, pending work, and restart behavior.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant