Skip to content

feat(agents): run queue, one theme per run and the theme toggle - #12112

Merged
MarkusNeusinger merged 8 commits into
mainfrom
feat/agents-run-queue
Oct 9, 2026
Merged

MarkusNeusinger merged 8 commits into
mainfrom
feat/agents-run-queue

Conversation

@MarkusNeusinger

Copy link
Copy Markdown
Owner

Summary

  • A FIFO run queue in front of every agent chat turn, as decided by the owner after spikes S and S2: one pipeline run in flight per instance (AGENT_RUN_CONCURRENCY, 1), at most AGENT_RUNS_PER_MINUTE starts (1, sliding window), a maximum wait of AGENT_QUEUE_MAX_WAIT_S (600 s) with the capacity that wait allows (10), 503 capacity or error{capacity} beyond it, status{step:"queued", position, waiting} on the stream, waiting and in_flight on GET /v1/status, queued time outside the request deadline, one queued or running turn per user, a premium lane that nothing sets yet, and budget refusals before a queue place.
  • Serial renders by construction: SerialRenderer wraps whichever backend renders (sandbox, local, fake, adk web), AGENT_RENDER_CONCURRENCY defaults to 1, one theme per slot, pipeline renders ahead of theme toggles.
  • One theme per run (light unless the user asks for dark): the job carries its themes, the gates judge the job's themes, the reviewer sees and names only the rendered theme, artifacts and the exported plot.py run line follow, and the padded fallback names the theme.
  • A theme toggle without LLM calls: POST /v1/sessions/{sid}/versions/{version}/render {theme} and its BFF mirror POST /debug/agent/sessions/{sid}/versions/{version}/render render the other theme of a finished version from its stored code and data through the host gates; a toggle and a turn refuse each other, one toggle per user in flight, at most two tries per theme, a 120 s wait for the render slot.
  • The BFF caps a relayed turn at AGENT_TURN_MAX_S (590 s, below anyplot-api's 600 s request timeout) so every stream still ends with error and done.
  • Two adversarial reviews (security/abuse, correctness/ADK) with 12 findings, all applied except the capacity-formula change, which is documented instead (see open decisions).

Plan

docs/concepts/agent-network.md (bounds table, SSE protocol, render section and open decisions updated in this PR). Follow-ups, not in this PR: the real sandbox backend and the renderer service, the agents image and Cloud Build, the frontend.

Open decisions for the owner

  • Through the BFF a turn can wait about 385 s, not the full 600 s (the 590 s cap minus the 190 s budget and the 15 s heartbeat). Either raise anyplot-api's --timeout to about 900 s and AGENT_TURN_MAX_S to about 890 s, or lower AGENT_QUEUE_MAX_WAIT_S to about 385 s.
  • The capacity formula (rate x max wait / 60) assumes runs end within a minute; with 90 s runs, 6 of 10 admitted entries are served and 4 expire. Admitting by estimated start time would change the owner's formula, so it is documented, not changed.
  • "One run or toggle per user" is stricter than "refused while that session has a run"; toggles last seconds, so this bounds waiting toggles to one per user.

Test plan

  • uv run ruff check . and uv run ruff format --check . clean
  • uv run --extra typecheck --extra agents mypy api core agents: no issues in 95 files
  • ENVIRONMENT=test uv run pytest tests/unit tests/integration: 7,807 passed, 1 skipped
  • uv run python -m tools.changelog check --base origin/main and uv lock --check pass
  • Queue unit tests (24): FIFO, concurrency, premium order, rate window, capacity, expiry, withdraw, heartbeat; serial renderer tests (9)
  • End-to-end /v1 through the queue with two users on a fake clock: the second sees queued position 1 before its run starts; full queue answers 503; 409 for a second turn or toggle of the same user; the 180 s timer is armed only when the run leaves the queue; a client gone while queued leaves the queue
  • Toggle tests: render without a model call, padded canvas, failed render retried once then answered from the record, 409 while a run is active, 404/422/other user; BFF relay incl. the turn cap ending a stream with capacity and done
  • Owner: decide the three open decisions above

MarkusNeusinger and others added 6 commits October 9, 2026 22:57
Owner decision of 2026-10-09, after spikes S and S2 showed one 4 GiB
instance serves one sandbox at a time safely:

- A two-lane FIFO run queue (agents/anyplot/run_queue.py) in front of
  every /messages turn: AGENT_RUN_CONCURRENCY (1) runs in flight,
  AGENT_RUNS_PER_MINUTE (1) starts in a sliding 60 s window, at most
  AGENT_QUEUE_MAX_WAIT_S (600) of waiting and rate x wait / 60 (10)
  entries. A full queue is 503 capacity before the stream, an expired
  entry error{code:"capacity"} inside it. The stream sends ready, then
  status{step:"queued", position, waiting} at once, on change and every
  15 s, then the run; the deadline and its abort start with the run. The
  run registry covers queued entries (409 run_active), cancel and purge
  withdraw a waiting entry, a gone client withdraws its own, and the stale
  sweep uses the queue's clock and maximum wait. /v1/status reports
  waiting and in_flight. A premium lane exists but nothing sets it.
- One theme per run: PipelineArgs.theme (light by default), the render
  job, the host gates (now judging the job's themes, not THEMES), the
  reviewer's images, the artifacts, the padded line and the canvas line
  all name that theme. The root's prompt and tool description say when
  to pass dark, and its session block names the latest version's theme.
- POST /v1/sessions/{sid}/versions/{version}/render {theme}: renders
  another theme of a finished version from its stored run form through
  the render backend and gates R1-R3 with the padding fallback, with no
  adapter, reviewer or queue; answers ok, needs_attention
  (canvas_padded) or failed (render, error) with the version's artifacts,
  409 while the session has a run, 404 for an unknown or swept version.
- AGENT_RENDER_CONCURRENCY defaults to 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
- POST /debug/agent/sessions/{sid}/versions/{version}/render {theme}
  mirrors the agents service's theme toggle: version 0-999 (0 is the
  latest), theme light or dark, an unknown field is 422, and upstream
  errors keep the agent error shape ({detail, ref}).
- The anyplot/1 relay passes position and waiting on status events, so
  status{step:"queued"} reaches the browser.
- Queued time no longer spends the turn budget: every queued status, and
  the first event after the wait, restarts AGENT_REQUEST_TIMEOUT_S, as
  the agents service starts its own deadline only when the run leaves
  the queue. A queue that falls silent still ends with error upstream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
- Design doc: the bounds paragraph becomes a bounds table with the
  queue, rate, maximum wait, queue length, per-user and render rows; the
  request flow, the /v1 table, the render section (one theme, serial
  renders, the padded line naming the theme), the SSE protocol (queued
  status, the BFF's budget restart), the Serving flags (--concurrency=20
  and --timeout=900 for the queue's open streams), the risks and two open
  decisions (the anyplot-api 600 s timeout against a 780 s queued turn,
  and whether "Create plot" carries the site theme).
- agents/README.md: the new modules and settings, the adk web bypass of
  the queue, and AGENT_RUNS_PER_MINUTE=60 for local iteration.
- docs/reference/api.md: the toggle route and its answers, the queued
  status, 503 capacity, waiting and in_flight on /status.
- changelog.d/agents-run-queue.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…start before expiry

Review fixes for the run queue branch:

- SerialRenderer (render/serial.py) wraps every backend in Services.backend:
  one theme per slot under AGENT_RENDER_CONCURRENCY, a freed slot goes to a
  waiting pipeline render before a waiting toggle, and a bounded wait raises
  RenderBusy. The local and sandbox backends drop their own semaphore.
- The theme toggle registers in the run registry, so it and a turn refuse
  each other with 409 run_active and a user has one run or toggle in flight;
  it waits for a slot at most request deadline minus render timeout (503
  capacity after that); a theme that failed the host gates is recorded and
  rendered at most twice, while an `error` stays unrecorded.
- The queue starts an entry whose turn comes at the moment its wait runs out,
  never one a late pump finds overdue, and documents that the owner's capacity
  formula assumes runs within the 60 s window.
- A user over the daily budget gets the budget refusal before the queue, and a
  queued turn counts as use of its dataset for the idle sweep.
- Reviewer defects name the rendered theme; plot.py's run line names it too.
- Tests: serial renderer, toggle in flight, busy slot, budget precheck,
  dataset touch, the deadline armed only when the run starts, start at the
  maximum wait, long runs past the capacity promise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
AGENT_TURN_MAX_S (590 s) bounds a chat turn, queue wait included, so Cloud
Run's 600 s timeout never cuts a stream without error and done. A turn still
queued when its run could no longer finish inside the cap ends with
error{capacity}, and closing the upstream takes it out of the queue before it
spends a token. A queued status restarts the budget with the queue's 15 s
heartbeat on top, because the run may start that long before its first event.
The status route's docstring lists every upstream field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…s promise

The design doc's bounds table, render section, risks and open decisions, the
agents README, the API reference and the changelog fragment describe the
render slots in front of every backend, the toggle's registry entry and slot
wait, the budget check before the queue, the BFF turn cap and what it means
for the 600 s maximum wait, and the capacity formula's assumption about run
length.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Copilot AI balanced review requested due to automatic review settings October 9, 2026 21:41

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

A refinement can silently reset a dark plot to light when the pipeline call omits its theme.

3 open findings
What changed in this PR

Adds bounded agent-run scheduling, serial rendering, and single-theme runs with on-demand theme rendering.

Changes:

  • Adds FIFO run queue, rate limits, status events, and turn deadlines.
  • Serializes rendering and renders one theme per pipeline run.
  • Adds theme-toggle APIs, persistence, documentation, and tests.
File Description
.env.example Documents the turn cap.
agents/​README.md Documents queue and theme behavior.
agents/​anyplot/​agent.py Exposes the latest theme to the agent.
agents/​anyplot/​code/​export.py Adds theme-aware run instructions.
agents/​anyplot/​pipeline.py Runs and reviews one theme.
agents/​anyplot/​prompts/​reviewer.md Adapts review instructions for one render.
agents/​anyplot/​prompts/​root.md Guides theme selection and toggling.
agents/​anyplot/​render/​__init__.py Moves concurrency to the wrapper.
agents/​anyplot/​render/​backends/​local.py Removes backend-local concurrency.
agents/​anyplot/​render/​backends/​sandbox.py Removes backend-local concurrency.
agents/​anyplot/​render/​contract.py Centralizes the theme contract.
agents/​anyplot/​render/​gates.py Evaluates only requested themes.
agents/​anyplot/​render/​serial.py Adds prioritized render slots.
agents/​anyplot/​render/​store.py Supports adding theme images.
agents/​anyplot/​run_queue.py Implements queued run scheduling.
agents/​anyplot/​schemas.py Adds theme and artifact contracts.
agents/​anyplot/​services.py Stores per-theme outcomes and wraps backends.
agents/​anyplot/​settings.py Adds queue and concurrency settings.
agents/​anyplot/​sub_agents/​reviewer.py Attaches only rendered-theme images.
agents/​anyplot/​theme_render.py Implements model-free theme rendering.
agents/​anyplot/​tools/​session.py Extends pipeline tool arguments.
agents/​main.py Integrates queue and toggle routes.
agents/​stream.py Adds queued SSE statuses.
api/​routers/​agent.py Relays queue events and toggle requests.
changelog.d/​agents-run-queue.md Records the feature.
core/​config.py Adds the BFF turn limit.
docs/​concepts/​agent-network.md Updates the architecture design.
docs/​reference/​api.md Documents new API behavior.
tests/​unit/​agents/​code/​test_export.py Tests theme-aware exports.
tests/​unit/​agents/​runtime/​test_pipeline.py Tests single-theme results.
tests/​unit/​agents/​runtime/​test_registry.py Verifies the expanded tool schema.
tests/​unit/​agents/​runtime/​test_render.py Tests theme-scoped gates and storage.
tests/​unit/​agents/​runtime/​test_run_queue.py Tests queue semantics.
tests/​unit/​agents/​runtime/​test_run_queue_flow.py Tests queue integration.
tests/​unit/​agents/​runtime/​test_serial_render.py Tests render serialization.
tests/​unit/​agents/​runtime/​test_service_flow.py Tests theme runs and toggles.
tests/​unit/​agents/​runtime/​test_stream.py Tests queued stream events.
tests/​unit/​agents/​test_settings.py Tests queue configuration.
tests/​unit/​api/​test_agent_router.py Tests BFF queue and toggle behavior.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread agents/anyplot/schemas.py Outdated
Comment thread agents/anyplot/pipeline.py
Comment thread docs/reference/api.md Outdated
… change

PipelineArgs.theme is optional: a refinement on base='previous' without a theme
renders the latest version's theme instead of falling back to light; a new plot
stays light. The tool description and the root prompt say so, the pipeline
docstring names the theme field, and the API reference heading no longer reads
as a theme statement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
@codecov

codecov Bot commented Oct 9, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@MarkusNeusinger
MarkusNeusinger merged commit f40d298 into main Oct 9, 2026
16 checks passed
@MarkusNeusinger
MarkusNeusinger deleted the feat/agents-run-queue branch October 9, 2026 22:45
MarkusNeusinger added a commit that referenced this pull request Oct 10, 2026
## Summary
- anyplot-api's Cloud Run request timeout rises from 600 to 900 seconds
(`api/cloudbuild.yaml`) and the BFF's turn cap `AGENT_TURN_MAX_S` from
590 to 890 seconds, so a chat turn at the back of a full run queue (600
seconds of waiting plus the 180-second run) ends with its own `error`
and `done` instead of being cut at about 385 seconds.
- The design document records the owner's decisions of 2026-10-10 on the
three open points from #12112: raise the timeout (this PR), keep the
queue's capacity formula, keep one run or theme toggle in flight per
user.
- Docs: `docs/reference/api.md` and `.env.example` follow the new
values.

## Plan
`docs/concepts/agent-network.md` (bounds table, SSE section, open
decisions, risks).

## Test plan
- [x] `uv run ruff check .`, `uv run ruff format --check .`, mypy (95
files), `tools.changelog check`, `uv lock --check` pass
- [x] `tests/unit/api/test_agent_router.py` and `tests/unit/core` pass
(the turn-cap test overrides the setting, so the default change touches
no assertion)
- [ ] After merge: the deploy-api Cloud Build applies `--timeout=900` to
the `anyplot-api` service (watch the build and `gcloud run services
describe anyplot-api
--format='value(spec.template.spec.timeoutSeconds)'`)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
MarkusNeusinger added a commit that referenced this pull request Oct 10, 2026
…12115)

## Summary
- Adds `anyplot-renderer` (`agents/renderer/`): a small FastAPI service
with `POST /render`, `POST /render/{job_id}/cancel` and `GET /status`
behind Cloud Run IAM and the ID-token claims check, that runs each theme
of a job in its own Cloud Run sandbox (`sandbox do`, never `--write` or
egress, explicit `PATH`), one at a time, with the limits spikes S and S2
measured: run directories on a 512 MiB in-memory volume the service
requires on Cloud Run (`503 volume_missing` without it), a watchdog that
kills a run past 64 MiB, 10,000 files or a 512 MiB `MemAvailable` floor,
a kill that counts only once the launcher has exited, one retry when the
launcher fails before the harness starts, bounded output, a cleaned
probe, rlimits, a stop on client disconnect or cancel, JSON `500
internal` without the message, and a self-restart when a launcher
outlives its kill. It only runs code; the host gates stay in the agents
service. The image installs the plotting libraries without ADK;
`agents/renderer/cloudbuild.yaml` builds, deploys a `candidate`, smokes
and promotes it like the API, and creates the service on the first build
only when `describe` reports it missing.
- Adds the `remote` render backend, now the default
(`AGENT_RENDERER=remote`, `AGENT_RENDER_URL`), so a local `adk web` and
the deployed agents service render in the same sandboxes: an ID token
from the metadata server when deployed, or from `AGENT_RENDER_TOKEN` or
the developer's Application Default Credentials in development; retries
over a cold start (1, 3, 8, 15 s within 30 s) for connection errors and
front-end 429/5xx only, never for the renderer's own answers; a cancel
call for abandoned renders; limit and measured value in the repair
feedback.
- Hardens the caller check in both services:
`X-Serverless-Authorization` wins whenever it is present, because Cloud
Run then checks only that header and passes `Authorization` through
unverified. The harness prints a start and an end line and sets soft and
hard limits alike. CI gains a `renderer-image` job (`ci-image.yml`) that
builds the image and renders a seaborn plot through the harness with no
network.

## Plan
Design: `docs/concepts/agent-network.md` ("Render", "Serving and
infrastructure"; status line and service table updated). The renderer
split for phase 1 was decided on 2026-10-10 (one renderer for the
deployed agents and for `adk web`). Follows #12111 and #12112.

## Test plan
- [x] `uv run ruff check . && ruff format --check .`, `mypy api core
agents` (101 files), `pytest tests/unit tests/integration` (7935 passed;
executor tests run the real harness against a fake launcher: kill path,
cancel, disconnect, stuck launcher, retry once and twice, byte, file and
memory watchdog, bounded output, refused output files; service tests
cover the caller check, the slot, low memory, the volume check and
`/status`; the `remote` backend is tested against the in-process app and
once end to end through the real executor and host gates), changelog
check, `uv lock --check`
- [x] Deployed to Cloud Run as a throwaway service
(`anyplot-renderer-spike`, a role-less service account, Cloud Build
`5b86cc7a` from this head with `_SMOKE_RENDER=false`, 3 min 51 s): the
first-build path created the service; `gcloud beta run deploy
--sandbox-launcher` and the repeated `--add-volume` flags are accepted
by the Cloud Build gcloud image (587.0.0); an anonymous call gets 403;
`/status` with the owner's gcloud token reports `sandbox: true`,
`volume: true` (the in-memory volume reports its 512 MiB limit), 3,524
MiB available; a real matplotlib render answered in 2.9 s with a 108 KB
PNG; spike X (122 cases) renders through it from a laptop with
`--gcloud-token`
- [x] Configured build on the throwaway with `_SMOKE_RENDER=true` (Cloud
Build `a89a46b2` from `09f50b62d`, 3 min 47 s): the `--no-traffic`
candidate path, the smoke as the service's own account (impersonated; a
build cannot mint an ID token for itself, which builds `da6f53f5` and
`a9d3ac7f` showed and the two fix commits address), `/status` with the
volume in place, a real render in 3.2 s, promote. The `renderer-image`
CI job ran and passed on every push.
- [x] `/tmp` fill through the promoted revision (the review's open
question): code writing 40 MiB files into the sandbox's private `/tmp`
ends after 4.6 s with `OSError: [Errno 28] No space left on device`; the
instance stays up (3,451 MiB available before, 3,328 after), the memory
watchdog never fires. Decided in the design doc: 4 GiB with one sandbox
at a time, no `RENDERER_BIND_TMP`, no 8 GiB instance. The throwaway
service, its service account and its images are deleted.
- [ ] Not verifiable before merge: the Application Default Credentials
token path against Cloud Run IAM; `.github/workflows/` has no pre-merge
loop
- [ ] After merge nothing deploys: the real `anyplot-renderer` service
account, the first build and the `deploy-renderer` trigger are owner
tasks (design doc, task 14)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
MarkusNeusinger added a commit that referenced this pull request Oct 10, 2026
## Summary
- Adds the admin-only "Use with my data" chat page at
`/debug/agent?spec=&library=&language=`
(`app/src/pages/AgentChatPage.tsx`): data panel with a 200 KB counter,
20-row preview, a column dropdown for every spec role (series families
get one dropdown per member), the chat thread, a progress timeline that
follows the `anyplot/1` stream, a queue-position counter in the chat's
corner, `.stop()` through the cancel route, and library pills. Built
only with `VITE_ENABLE_AGENT_CHAT=true` (off in `app/cloudbuild.yaml`,
so production ships without the chunk) and opens only in local
development or for a browser carrying the admin hint.
- Adds the result card (`app/src/sections/agent-chat/ResultCard.tsx`)
with the plot page's overlay actions (copy image, download PNG, open
full size), a light/dark switch through the theme toggle route, the
adapted `plot.py` with copy and download, `data.csv` download, the
change list and residual notes, a refine composer and the quick-feedback
slot; plus the `.adapt()` overlay button on the plot page, gated by
build flag, admin hint and the BFF's eligibility route so public
visitors never call the debug API.
- Two small contract changes on the service side so the page can address
versions and roles reliably: the `plot` event carries `version` (agents
translator and BFF allowlist), and the dataset answer lists the spec's
`roles` with a 422 that keeps the binding check's lines. Analytics
events (enum properties only) are documented in
`docs/reference/plausible.md`; a loopback mock of the BFF
(`app/scripts/agent-bff-mock.mjs`) drives the whole flow without the
API, the agents service or a model.

## Plan
Design: `docs/concepts/agent-network.md` (sections "Frontend", "Result
card", "Analytics"); this is the frontend PR of the phase-1 roadmap,
built on the merged runtime core (#12111) and run queue (#12112).

## Test plan
- [x] `cd app && yarn lint && yarn fm:check && yarn type-check && yarn
test` (81 files, 750 tests) and `yarn build` with the flag on and off;
the default build contains no agent client
- [x] `uv run --extra dev ruff check . && ruff format --check .`, `mypy
api core agents`, `pytest tests/unit tests/integration` (7816 passed),
changelog check, `uv lock --check`
- [x] Driven in the browser against the mock at 1280 px and 390 px,
light and dark: parse, bindings incl. series families and a refused set,
queued counter, progress timeline, result card actions (clipboard PNG,
downloads named `plot.py`/`data.csv`, open full size, copy code), theme
switch, repair round, refusal, error, stop then a new turn without a
409, ineligible pair; no horizontal scroll, 16 px gutter at 390 px
- [ ] After merge: `deploy-app` and `deploy-api` Cloud Builds succeed;
production stays unchanged (`_VITE_ENABLE_AGENT_CHAT` false,
`AGENT_ENABLED` false)
- [ ] Once the agents service is deployed: open
`/debug/agent?spec=scatter-basic&library=matplotlib` as admin and run
one "Create plot" end to end

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants