Repository navigation
Conversation
This field is used to track if a stats kind can use a custom format representation on disk when reading or writing its stats case. On HEAD, this exists for replication slots stats, that need a mapping between an internal index ID and the slot names. named_on_disk is currently used nowhere and the callbacks to_serialized_name and from_serialized_name are in charge of checking if the serialization of the stats data should apply, so let's remove it. Reviewed-by: Andres Freund Discussion: https://postgr.es/m/ZmKVlSX_T5YvIOsd@paquier.xyz (cherry picked from commit b19db55)
Shared statistics with a fixed number of objects are read from the stats file in pgstat_read_statsfile() using members of PgStat_ShmemControl and following an order based on their PgStat_Kind value. Instead of being explicit, this commit changes the stats read to iterate over the pgstat_kind_infos array to find the memory locations to read into, based on a new shared_ctl_off in PgStat_KindInfo that can be used to define the position of this stats kind in shared memory. This makes the read logic simpler, and eases the introduction of future improvements aimed at making this area more pluggable for external modules. Original idea suggested by Andres Freund. Author: Tristan Partin Reviewed-by: Andres Freund, Michael Paquier Discussion: https://postgr.es/m/D12SQ7OYCD85.20BUVF3DWU5K7@neon.tech (cherry picked from commit 9004abf)
This changes pgstat.c so as the three types of entries that can exist in a pgstats file are not hardcoded anymore, replacing them with descriptively-named macros, when reading and writing stats files: - 'N' for named entries, like replication slot stats. - 'S' for entries identified by a hash. - 'E' for the end-of-file This has come up while working on making this area of the code more pluggable. The format of the stats file is unchanged, hence there is no need to bump PGSTAT_FILE_FORMAT_ID. Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Zmqm9j5EO0I4W8dx@paquier.xyz (cherry picked from commit 9fd0252)
This is similar to 9004abf, but this time for the write part of the stats file. The code is changed so as, rather than referring to individual members of PgStat_Snapshot in an order based on their PgStat_Kind value, a loop based on pgstat_kind_infos is used to retrieve the contents to write from the snapshot structure, for a size of PgStat_KindInfo's shared_data_len. This requires the addition to PgStat_KindInfo of an offset to track the location of each fixed-numbered stats in PgStat_Snapshot. This change is useful to make this area of the code more easily pluggable, and reduces the knowledge of specific fixed-numbered kinds in pgstat.c. Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Zot5bxoPYdS7yaoy@paquier.xyz (cherry picked from commit b68b29b)
This new callback gives fixed-numbered stats the possibility to take actions based on the area of shared memory allocated for them. This removes from pgstat_shmem.c any knowledge specific to the types of fixed-numbered stats, and the initializations happen in their own files. Like b68b29b, this change is useful to make this area of the code more pluggable, so as custom fixed-numbered stats can take actions after their shared memory area is initialized. Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Zot5bxoPYdS7yaoy@paquier.xyz (cherry picked from commit 21471f1)
This new entry type is used for all the fixed-numbered statistics, making possible support for custom pluggable stats. In short, we need to be able to detect more easily if a stats kind exists or not when reading back its data from the pgstats file without a dependency on the order of the entries read. The kind ID of the stats is added to the data written. The data is written in the same fashion as previously, with the fixed-numbered stats first and the dshash entries next. The read part becomes more flexible, loading fixed-numbered stats into shared memory based on the new entry type found. Bump PGSTAT_FILE_FORMAT_ID. Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Zot5bxoPYdS7yaoy@paquier.xyz (cherry picked from commit 9e4664d)
A follow-up patch is planned to make cumulative statistics pluggable, and using a type is useful in the internal routines used by pgstats as PgStat_Kind may have a value that was not originally in the enum removed here, once made pluggable. While on it, this commit switches pgstat_is_kind_valid() to use PgStat_Kind rather than an int, to be more consistent with its existing callers. Some loops based on the stats kind IDs are switched to use PgStat_Kind rather than int, for consistency with the new time. Author: Michael Paquier Reviewed-by: Dmitry Dolgov, Bertrand Drouvot Discussion: https://postgr.es/m/Zmqm9j5EO0I4W8dx@paquier.xyz Cloudberry: PGSTAT_KIND_RESQUEUE is kept as the sixth variable-numbered kind, so the IDs of the fixed-numbered kinds are shifted by one. (cherry picked from commit 3188a45)
This is useful to know which part of a stats file is corrupted when reading it, adding to the server logs a WARNING with details about what could not be read before giving up with the remaining data in the file. Author: Michael Paquier Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Zp8o6_cl0KSgsnvS@paquier.xyz (cherry picked from commit ca1ba50)
This commit adds support in the backend for $subject, allowing out-of-core extensions to plug their own custom kinds of cumulative statistics. This feature has come up a few times into the lists, and the first, original, suggestion came from Andres Freund, about pg_stat_statements to use the cumulative statistics APIs in shared memory rather than its own less efficient internals. The advantage of this implementation is that this can be extended to any kind of statistics. The stats kinds are divided into two parts: - The in-core "builtin" stats kinds, with designated initializers, able to use IDs up to 128. - The "custom" stats kinds, able to use a range of IDs from 128 to 256 (128 slots available as of this patch), with information saved in TopMemoryContext. This can be made larger, if necessary. There are two types of cumulative statistics in the backend: - For fixed-numbered objects (like WAL, archiver, etc.). These are attached to the snapshot and pgstats shmem control structures for efficiency, and built-in stats kinds still do that to avoid any redirection penalty. The data of custom kinds is stored in a first array in snapshot structure and a second array in the shmem control structure, both indexed by their ID, acting as an equivalent of the builtin stats. - For variable-numbered objects (like tables, functions, etc.). These are stored in a dshash using the stats kind ID in the hash lookup key. Internally, the handling of the builtin stats is unchanged, and both fixed and variabled-numbered objects are supported. Structure definitions for builtin stats kinds are renamed to reflect better the differences with custom kinds. Like custom RMGRs, custom cumulative statistics can only be loaded with shared_preload_libraries at startup, and must allocate a unique ID shared across all the PostgreSQL extension ecosystem with the following wiki page to avoid conflicts: https://wiki.postgresql.org/wiki/CustomCumulativeStats This makes the detection of the stats kinds and their handling when reading and writing stats much easier than, say, allocating IDs for stats kinds from a shared memory counter, that may change the ID used by a stats kind across restarts. When under development, extensions can use PGSTAT_KIND_EXPERIMENTAL. Two examples that can be used as templates for fixed-numbered and variable-numbered stats kinds will be added in some follow-up commits, with tests to provide coverage. Some documentation is added to explain how to use this plugin facility. Author: Michael Paquier Reviewed-by: Dmitry Dolgov, Bertrand Drouvot Discussion: https://postgr.es/m/Zmqm9j5EO0I4W8dx@paquier.xyz (cherry picked from commit 7949d95)
This is useful for extensions to get snapshot and shmem data for custom cumulative statistics when these have a fixed number of objects, so as these do not need to know about the snapshot internals, aka pgStatLocal. An upcoming commit introducing an example template for custom cumulative stats with fixed-numbered objects will make use of these. I have noticed that this is useful for extension developers while hacking my own example, actually. Author: Michael Paquier Reviewed-by: Dmitry Dolgov, Bertrand Drouvot Discussion: https://postgr.es/m/Zmqm9j5EO0I4W8dx@paquier.xyz (cherry picked from commit 2eff9e6)
There were two spots in pgstat_read_statsfile() where is was possible to finish with a null-pointer-dereference crash for custom pgstats kinds: - When reading stats for a fixed-numbered stats entry. - When reading a variable stats entry with name serialization. For both cases, these issues were reachable by starting a server after changing shared_preload_libraries so as the stats written previously could not be loaded. The code is changed so as the stats are ignored in this case, like the other code paths doing similar sanity checks. Two WARNINGs are added to be able to debug these issues. A test is added for the case of fixed-numbered stats with the module injection_points. Oversights in 7949d95, spotted while looking at a different report. Discussion: https://postgr.es/m/Ztj0Jftsn4xXuXtl@paquier.xyz Cloudberry: the injection_points test module does not exist here (it appeared in PostgreSQL 17), so its part of the change is left out. (cherry picked from commit 341e9a0)
This new field controls if entries of a stats kind should be written or not to the on-disk pgstats file when shutting down an instance. This affects both fixed and variable-numbered kinds. This is useful for custom statistics by itself, and a patch is under discussion to add a new builtin stats kind where the write of the stats is not necessary. All the built-in stats kinds, as well as the two custom stats kinds in the test module injection_points, set this flag to "true" for now, so as stats entries are written to the on-disk pgstats file. Author: Bertrand Drouvot Reviewed-by: Nazir Bilal Yavuz Discussion: https://postgr.es/m/Zz7T47nHwYgeYwOe@ip-10-97-1-34.eu-west-3.compute.internal Cloudberry: the injection_points test module does not exist here (it appeared in PostgreSQL 17), so its part of the change is left out. (cherry picked from commit c06e71d)
This includes all the definitions for the various PGSTAT_KIND_* values, the range allowed for custom stats kinds and some macros related all that. One use-case behind this split is the possibility to use this information for frontend tools, without having to rely on pgstat.h and a backend footprint. Author: Michael Paquier Reviewed-by: Bertrand Drouvot Discussion: https://postgr.es/m/Z24fyb3ipXKR38oS@paquier.xyz Cloudberry: there are no backend statistics here (PostgreSQL 18), so the sixth variable-numbered kind is PGSTAT_KIND_RESQUEUE instead of PGSTAT_KIND_BACKEND. (cherry picked from commit d35ea27)
This commit changes stats kinds to have the following bounds, making their handling in core cheaper by default: - PGSTAT_KIND_CUSTOM_MIN 128 -> 24 - PGSTAT_KIND_MAX 256 -> 32 The original numbers were rather high, and showed an impact on performance in pgstat_report_stat() for the case of simple queries with its early-exit path if there are no pending statistics to flush. This logic will be improved more in a follow-up commit to bring the performance of pgstat_report_stat() on par with v17 and older versions. Lowering the bounds is a change worth doing on its own, independently of the other improvement. These new numbers should be enough to leave some room for the following years for built-in and custom stats kinds, with stable ID numbers. At least that should be enough to start with this facility for extension developers. It can be always increased in the tree depending on the requirements wanted. Per discussion with Andres Freund and Bertrand Drouvot. Discussion: https://postgr.es/m/eb224uegsga2hgq7dfq3ps5cduhpqej7ir2hjxzzozjthrekx5@dysei6buqthe Backpatch-through: 18 Cloudberry: the injection_points test module does not exist here, so its part of the change is left out. (cherry picked from commit ac000fca743eff923d1feb4bc722d905901ae540)
Keep resource queue statistics across clean restarts by enabling write_to_file for the Cloudberry resqueue kind, and cover persistence with a restart TAP test. Advance the statistics file format identifier for the pluggable statistics layout.
Add an optional reset_data_cb for variable-numbered statistics kinds. Call it under the existing entry lock before resetting the timestamp. Kinds without the callback retain the existing zero-fill behavior.
…alyze This commit adds four fields to the statistics of relations, aggregating the amount of time spent for each operation on a relation: - total_vacuum_time, for manual vacuum. - total_autovacuum_time, for vacuum done by the autovacuum daemon. - total_analyze_time, for manual analyze. - total_autoanalyze_time, for analyze done by the autovacuum daemon. This gives users the option to derive the average time spent for these operations with the help of the related "count" fields. Bump PGSTAT_FILE_FORMAT_ID for the additions in PgStat_StatTabEntry. Catalog functions, views and their version bump follow separately. Author: Sami Imseih Reviewed-by: Bertrand Drouvot, Michael Paquier Discussion: https://postgr.es/m/CAA5RZ0uVOGBYmPEeGF2d1B_67tgNjKx_bKDuL+oUftuoz+=Y1g@mail.gmail.com (cherry picked from commit 30a6ed0) Cloudberry adaptation: time AO vacuum from its first phase and pass the start timestamp to the collector at final cleanup, so the changed API is usable by heap and AO immediately. Co-authored-by: Alena Rybakina <alenka.rybakina@gmail.com>
This function is used in both vacuum and analyze code paths, and a follow-up commit will require distinguishing between the two. This commit forces callers to specify whether they are in a vacuum or analyze path, but it does not use that information for anything yet. Author: Nathan Bossart <nathandbossart@gmail.com> Co-authored-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com> Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal Backported from PostgreSQL 18 (commit e5b0b0c) as a prerequisite of the vacuum statistics series. Cloudberry adaptation: the append-optimized and AOCS sample-row acquisition and compaction call vacuum_delay_point() too; the former pass is_analyze = true, the latter false.
This commit adds the amount of time spent sleeping due to cost-based delay to the pg_stat_progress_vacuum and pg_stat_progress_analyze system views. A new configuration parameter named track_cost_delay_timing, which is off by default, controls whether this information is gathered. For vacuum, the reported value includes the sleep time of any associated parallel workers. However, parallel workers only report their sleep time once per second to avoid overloading the leader process. Bumps catversion. Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com> Co-authored-by: Nathan Bossart <nathandbossart@gmail.com> Reviewed-by: Sami Imseih <samimseih@gmail.com> Reviewed-by: Robert Haas <robertmhaas@gmail.com> Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com> Reviewed-by: Masahiro Ikeda <ikedamsh@oss.nttdata.com> Reviewed-by: Dilip Kumar <dilipbalaut@gmail.com> Reviewed-by: Sergei Kornilov <sk@zsrv.org> Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal Backported from PostgreSQL 18 (commit bb8dff9) as a prerequisite of the vacuum statistics series. Cloudberry adaptation: the vacuum progress parameters of this branch end at num_dead_tuples, so the delay time is PROGRESS_VACUUM_DELAY_TIME 7 (param8 of the view). A parallel vacuum worker only accumulates its sleep time: reporting it to the leader needs pgstat_progress_parallel_incr_param() of PostgreSQL 17, and Cloudberry does not run parallel vacuum at all (see dead_items_alloc()). The new GUC is listed in unsync_guc_names_array, next to track_io_timing, as every GUC has to be in one of the two lists here. SQL progress-view changes are separated into the native SQL reporting commit; measurement, the GUC and progress counters are introduced here.
Port the process-local VacuumDelayTime accumulator from the REL_2_STABLE commit. Accumulate measured cost-based VACUUM sleep time only while track_cost_delay_timing is enabled, excluding ANALYZE. Consumers take before/after differences to attribute delays to their own operations. The preceding PostgreSQL 18 progress-view backport already provides the GUC, its documentation, sample setting, unsynchronized registration and sleep timing. Reuse that measurement and preserve both VACUUM and ANALYZE progress reporting. Keep the accumulator in milliseconds as double to match this branch's cumulative timing consumers, rather than the source branch's int64 microseconds. Extract the accumulator from the index/database timing commit and fold in the ANALYZE exclusion previously included in the combined heap/index/AO measurements commit. The final branch contents are unchanged. Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com> Discussion: https://postgr.es/m/ZmaXmWDL829fzAVX%40ip-10-97-1-34.eu-west-3.compute.internal Backported from PostgreSQL 18 (commit bb8dff9). (cherry picked from commit 01aaacb) Co-authored-by: Nathan Bossart <nathandbossart@gmail.com> Co-authored-by: Alena Rybakina <alenka.rybakina@gmail.com>
Accumulate elapsed time and cost-based delays over active AO vacuum phases, excluding the gaps between phases. Discard state left over from an interrupted operation. Prepare per-index timing measurements for the subsequent cumulative-reporting change. Adapted-from: 6f38fb2
Record index-pass elapsed and cost-delay time, including parallel workers with the leader's autovacuum classification. Add cumulative table delay time and database totals without adding index time to database totals a second time. Include removed tuple counts in verbose index reports. Catalog getters, views, documentation and SQL-dependent tests follow in the separate native SQL reporting commit.
Accumulate relation and database failsafe counts for completed heap vacuums. Pass false for AO, which has no heap wraparound failsafe mode. Keep this counter distinct from scans made aggressive by the freeze age. Catalog getters, views and tests follow in the native SQL reporting commit.
Assemble backend-local tuple and page measurements and take per-call snapshots of the existing index results. Access methods accumulate results across passes, so retain each pass's delta without counting earlier work again. Share the measurement helper between serial and parallel vacuum. Keep SP-GiST's newly deleted page count cumulative within one vacuum, without counting pages that were already empty on entry. Handle AMs that reset their newly deleted page count between passes. Adapt to PostgreSQL 16's existing recently_dead_tuples, missed_dead_tuples, pages_scanned, pages_removed and missed_dead_pages counters. Assemble these measurements independently of optional hooks; storage and SQL follow in separate reporting commits. Adapted-from: f2592a3
Count ERROR reports in the heap vacuum callback using local counters for database-local and shared relations. Transfer them to pending database statistics at transaction end. Test cancellation of a running vacuum.
Introduce dead_pages in LVRelState and PgStat_VacuumStats. Count heap pages that retain recently dead tuples in both pruning and no-prune scans. Keep this measurement distinct from pages with dead tuples missed because of buffer pins. For index cleanup, derive the count from deleted pages that are not yet reusable. Cumulative statistics storage and SQL exposure follow separately. Adapted-from: 37a4753
Add visible_page_marks_cleared and frozen_page_marks_cleared counters to
pg_stat_all_tables tracking the number of times the all-visible and
all-frozen bits are cleared in the visibility map. These bits are cleared by
backend processes during regular DML operations. Hence, the counters are placed
in table statistic entry.
A high visible_page_marks_cleared rate relative to DML volume indicates
that modifications are scattered across previously-clean pages rather
than concentrated on already-dirty ones, causing index-only scans to
fall back to heap fetches. A high frozen_page_marks_cleared rate indicates
that vacuum's freezing work is being frequently undone by concurrent
DML.
Authors: Alena Rybakina <lena.ribackina@yandex.ru>,
Andrei Lepikhov <lepihov@gmail.com>,
Andrei Zubkov <a.zubkov@postgrespro.ru>
Reviewed-by: Dilip Kumar <dilipbalaut@gmail.com>,
Masahiko Sawada <sawada.mshk@gmail.com>,
Ilia Evdokimov <ilya.evdokimov@tantorlabs.com>,
Jian He <jian.universality@gmail.com>,
Kirill Reshke <reshkekirill@gmail.com>,
Alexander Korotkov <aekorotkov@gmail.com>,
Jim Nasby <jnasby@upgrade.com>,
Sami Imseih <samimseih@gmail.com>,
Karina Litskevich <litskevichkarina@gmail.com>,
Andrey Borodin <x4mmm@yandex-team.ru>
Backported-from: https://www.postgresql.org/message-id/attachment/205562/v45-0009-Track-table-VM-stability.patch
Co-authored-by: Andrei Lepikhov <lepihov@gmail.com>
Co-authored-by: Andrei Zubkov <zubkov@moonset.ru>
Include pages_frozen and tuples_frozen in the backend-local VACUUM measurements. PostgreSQL 16 already counts frozen_pages and tuples_frozen and prints them in VACUUM VERBOSE, so reuse those counters without adding a second increment at the freezing site. Cumulative statistics storage and SQL exposure follow separately. Adapted-from: 93ae811
Connect AO index-pass timing measurements to cumulative statistics and report interrupted vacuums. Add cumulative delay reporting for AO tables. Failsafe and visibility-map counters remain inapplicable to AO. The cluster isolation test is included later with its gp_stat_vacuum dependency. Active-phase elapsed reporting is connected in a later commit.
Snapshot buffer activity and block I/O timing around heap, index and AO vacuum work. Assemble backend-local resource records and accumulate AO resources over active phases. Keep relation-local block counts, and subtract index resources from table-only totals to avoid double counting. Introduce optional instrumentation and its consumer hook. Allocate it only when a consumer is installed. Persistent storage and SQL access through the extension follow separately. Native timing and work counters are introduced in independent branches.
Add relallfrozen, an estimate of the number of pages marked all-frozen in the visibility map. pg_class already has relallvisible, an estimate of the number of pages in the relation marked all-visible in the visibility map. This is used primarily for planning. relallfrozen, together with relallvisible, is useful for estimating the outstanding number of all-visible but not all-frozen pages in the relation for the purposes of scheduling manual VACUUMs and tuning vacuum freeze parameters. A future commit will use relallfrozen to trigger more frequent vacuums on insert-focused workloads with significant volume of frozen data. Bump catalog version Author: Melanie Plageman <melanieplageman@gmail.com> Reviewed-by: Nathan Bossart <nathandbossart@gmail.com> Reviewed-by: Robert Treat <rob@xzilla.net> Reviewed-by: Corey Huinker <corey.huinker@gmail.com> Reviewed-by: Greg Sabino Mullane <htamfids@gmail.com> Discussion: https://postgr.es/m/flat/CAAKRu_aj-P7YyBz_cPNwztz6ohP%2BvWis%3Diz3YcomkB3NpYA--w%40mail.gmail.com (cherry picked from commit 99f8f3f) Cloudberry adaptation: carry both visibility-map counts through QE-to-QD reports, replicated-table normalization, ANALYZE and CLUSTER. AO and indexes keep relallfrozen zero. Avoid VM I/O on the dispatcher during index catalog updates. Include local and distributed catalog tests. PostgreSQL 18 statistics-import APIs and their tests do not exist on this PostgreSQL 16 branch. Use relallfrozen in the VM isolation test. Hold a repeatable-read snapshot and explicitly populate the statistics cache before concurrent DELETE; flush in the writer to avoid depending on the flush interval.
Count completed compactions, moved tuples, scanned pages and the final segment count without an extension hook. Discard state left over from an interrupted operation. Keep per-index work measurements for each pass. Native cumulative storage and SQL reporting follow separately. Adapted-from: 6f38fb2
Add WAL record, full-page-image and byte deltas to the backend-local VACUUM resource measurements. Snapshot pgWalUsage around the same heap, index and AO phases as buffer usage. Keep index WAL separate from table resources and accumulate only active AO phases. Reuse PostgreSQL's existing pgWalUsage instrumentation. Delivery, persistence and SQL exposure of these counters are added by the separate extension reporting commit.
Add an elapsed-time reporting entry point and connect the previously introduced AO active-phase timers. Exclude gaps between phases from cumulative vacuum elapsed and cost-delay time, and remove the obsolete first-phase timing baseline. SQL exposure and tests follow separately.
Add native SQL getters and local/cluster views for maintenance elapsed and cost-delay time, failsafe runs, interrupted vacuums and cleared VM bits. Expose progress delay and test timing, VM snapshots and relallfrozen for heap and append-optimized relations without a statistics extension.
Connect the previously introduced heap, index and AO measurements to ordinary relation and database statistics independently of any extension hook. Include them in native snapshots, resets, transactional drops and clean-shutdown persistence. Store dead_tuples and total_file_segs as last-run samples; other work counters accumulate. Exclude index reports and last-run samples from database totals. Advance the statistics-file format for the stored fields. Catalog functions, views and tests follow in a separate SQL reporting commit.
Publish cumulative tuple, page and AO compaction counters through native SQL getters and local, per-segment and cluster summary views. Keep last-run snapshots separate from cumulative work and normalize replicated relations. Include heap/index/AO tests for counters, resets, snapshots and persistence.
Combine timing/VM and work counters, including their shared AO phase state, statistics layouts, SQL views and tests.
Combine native statistics with optional buffers/WAL measurements. Keep extension work payloads with their consumer in the extension change.
Join pluggable statistics with native VACUUM counters and resource instrumentation for the extension series.
Connect heap and index reporting hooks to pluggable per-relation and per-database storage. Expose buffer and WAL usage with basic vacuum work in local SQL views. Include collection control, reset permissions, transactional DROP/OID-reuse cleanup, persistence and their tests.
Distributed CREATE can encounter the same existing statistics entry on every QE. Use LOG_SERVER_ONLY on executors while retaining WARNING on the coordinator and utility connections. The counter reset is unchanged. Validated OID reuse on a coordinator and three primary segments: one client warning, executor diagnostics only in server logs, and counters reset on every instance.
Document that the original reset functions affect only the connected node, including when called on QD, and expose that scope in SQL function comments. Correct the interaction with pg_stat_reset and remove the obsolete fixed entry-size estimate. Validation: documentation markup and git diff --check.
Print buffer, WAL and I/O measurements for heap work, index passes, and each AO vacuum phase. Exclude index resources from the enclosing table or phase report, include local buffer writes, and keep ordinary VACUUM quiet. Remove the extension hook declaration, definition and collection gates from the instrumentation prerequisite. The extension patch introduces that interface together with its consumer. Drop unused AO aggregate reporting from the core-only patch. Add core TAP coverage without an extension and document VERBOSE output.
ANALYZE reads relallvisible and relallfrozen together instead of dispatching two queries. Preserve replicated-table normalization and local visibility-map counting. Check both coordinator catalog counts against the segment totals in the distributed and replicated regression cases.
Reset work for one relation or the current database without clearing timing, invocation counts, other native counters, or extension metrics. Add a privileged gp_stat_reset_vacuum_stats wrapper for all primary segments and cover scope, permissions, pending work, and restart behavior.
Remove duplicate work fields from the hook payload and extension storage. Read native work alongside separately collected resources, including when resource collection is disabled. Keep extension resets limited to their own metrics and document the independent collection and reset scopes.
Introduce gp_stats_vacuum_tables, indexes and database views, cluster summaries, and gp_ reset wrappers with the initial extension. Carry over replicated-table normalization, synchronized collection control coverage, transactional DROP tests, and cluster installation and failover documentation.
Move cluster timing, index-summary cardinality, and track_cost_delay_timing scope checks from the extension regression into the core vacuum_stats test. Cover distributed and replicated tables, and explicitly flush database statistics before checking their timing aggregate.
Carry native cluster timing checks into the shared extension prerequisites.
Use the core vacuum_stats regression for native timing summaries and local timing GUC scope. Keep extension tests focused on resource reporting, cluster aggregation, independent resets, and transactional cleanup.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Add ext_vacuum_statistics to accumulate and expose VACUUM resource usage for
heap tables, indexes and databases.
Connect the core instrumentation to a custom cumulative statistics kind.
Expose buffer activity, I/O timing, WAL usage and basic vacuum work through
local views in the ext_vacuum_statistics schema.
Provide collection control, privileged reset functions, clean-shutdown
persistence and transactional cleanup of statistics for dropped objects.
Preserve counters when DROP is rolled back and prevent stale statistics
from being reused with an object's OID.
Impact
Makes cumulative VACUUM resource usage available through SQL for each table, index and database. Comparing buffer activity, WAL generation and I/O time with native work counters helps identify expensive relations and assess the cost of vacuuming them.
Resource collection and resets are independent of native work statistics, allowing separate measurement intervals without duplicating counters or clearing core statistics.
Type of Change
Breaking Changes
Test Plan
make installcheckmake -C src/test installcheck-cbdb-parallelChecklist