Repository navigation
New scaling controller and logic - #6897
nadav-govari wants to merge 1 commit into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/model/shard_table.rs#L112-L118
Exclude closed shards from throughput totals
After a scale-down, closed shards retain nonzero short- and long-term rate readings while their traffic moves to the remaining open shards. Including those readings here while counting only open shards in num_open_shards double-counts the source throughput on the next reconciliation, which can immediately scale back up, reset the cooldown, and leave the source overprovisioned.
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/ingest/ingest_controller.rs#L948-L950
Filter live ingesters by readiness
When the pool contains an Initializing, Retiring, or Failed ingester, keys() still includes it, unlike the existing placement logic in eligible_ingesters, which requires status.is_ready(). Consequently shard statistics and scale-down candidates treat shards on a non-serving ingester as usable capacity, potentially suppressing needed scale-up or sending close requests to an unavailable node during status transitions.
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/ingest/scaling_controller.rs#L352-L353
Split the new controller below the file-size limit
The newly added scaling_controller.rs is 780 lines long, exceeding the repository's mandatory 500-line cap for new files. Split the tests or controller responsibilities into separate modules so the new file complies with the documented maintainability constraint.
AGENTS.md reference: AGENTS.md:L133-L136
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/ingest/scaling_controller.rs#L166-L176
Wait for throughput reports before scaling down
After a control-plane restart, shards loaded from the metastore have default zero rates, while the cluster feature flag can already indicate that every indexer supports v2. Because reconciliation is not synchronized with receipt of ReportIndexerState data, the first loop can interpret missing measurements as idle traffic and close an actively used source down to min_shards; a delayed or temporarily failing reporter is enough to trigger this. Track report receipt/freshness and skip scale-down until current measurements have arrived.
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/ingest/scaling_controller.rs#L87-L90
Remove deleted sources from the cooldown map
Every source that scales is inserted into last_shard_count_changes, but deleting its source or index never removes that entry. Since SourceUid contains the index incarnation, recreated or auto-created indexes do not reuse the key, so a long-running cluster with source churn accumulates these entries indefinitely; retain only currently modeled sources or clear the entry during deletion.
https://github.com/quickwit-oss/quickwit/blob/62fd855d1832fff55aef286fb5277faf9654ef61/quickwit-control-plane/src/control_plane.rs#L230-L234
Document the new scaling protocol and rollout behavior
This change replaces gossip with gRPC reports once a cluster-wide feature condition is met and changes when and how shard scaling occurs, but no architecture or protocol documentation is updated. Document the v2 reporting transition, gating behavior, and scaling semantics so operators and future maintainers can reason about mixed-version rollouts.
AGENTS.md reference: AGENTS.md:L21-L24
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Description
Core logic of the change: The new scaling logic.
The new scaler only considers scaling decisions on the control plane loop, once per second, rather than on every report from an indexer. It does both scale up followed by rebalancing. All of this requires all indexers to report that theyre on the new scaling logic.
Other than that, consolidates shared shard scale up/down logic between the two scaling controller on the ingest controller. It also uses the existing ShardStats so there's some refactoring there too.
How was this PR tested?
Unit tests.