Repository navigation
feat(agents): anyplot-renderer service and the remote render backend - #12115
Merged
Merged
Conversation
The renderer runs adapted plot code in Cloud Run sandboxes as a service of its own (owner decision of 2026-10-10), so a local `adk web` and the deployed agents service share one renderer. - agents/renderer/: FastAPI POST /render and GET /status behind the IAM caller check (copied from agents/main.py; RENDERER_AUDIENCES and RENDERER_ALLOWED_CALLERS), one render slot with a bounded wait, a MemAvailable floor, and the sandbox executor built on spikes S and S2: no --write or --allow-egress, an explicit PATH, run directories under /tmp/runs with a byte and entry watchdog, killpg + `sandbox delete --force` confirmed by the launcher's exit (stuck launchers block renders), one retry with a new sandbox name when the launcher fails before the harness starts, bounded stdout/stderr, O_NOFOLLOW + fstat output reads. wire.py is the JSON contract; no ADK, agents.anyplot or core import, and no __init__.py so `adk web agents` does not list it as an agent. - Dockerfile (plotting venv without ADK, baked font cache, harness at /opt/anyplot/harness.py, runs as root for the launcher) and cloudbuild.yaml (candidate deploy that bootstraps the service on the first build, IAM 403 smoke and an optional real render smoke, promote). - agents/anyplot/render/backends/remote.py: the remote backend, now the AGENT_RENDERER default, with an ID-token source (metadata server when deployed; AGENT_RENDER_TOKEN or the user's ADC id_token in development), one retry on connection errors and Cloud Run front-end 429/5xx, and every other failure mapped to RendererUnavailable. The in-process sandbox backend stays as an unused stub. - harness.py: HARNESS start/end lines, soft = hard rlimits, opt-in RLIMIT_AS and RLIMIT_NPROC, the baked font cache copied into MPLCONFIGDIR. - Tests for the executor (fake launcher, real harness), the service, the contract and the remote backend in process; docs and changelog fragment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Apply the adversarial review of the renderer service and the remote backend:
- The deploy smoke mints its token for the service URL (Cloud Run accepts
no tag URL as an audience); the deploy adds status.url to
RENDERER_AUDIENCES itself and creates the service only on NOT_FOUND.
- Plot output can no longer crash /render: a forged harness end line, a
probe nested too deep or holding an oversized number, and host OSErrors
(503 io) are handled; any other error answers a JSON 500 internal, which
the remote backend never retries.
- The remote backend rides out a cold start (1, 3, 8, 15 s within 30 s)
and posts /render/{job_id}/cancel when its render is cancelled or times
out; the service races each render against http.disconnect.
- The watchdog also kills a run below a MemAvailable floor (memory), counts
an unmeasurable tree as over budget (file_budget), and the service
requires the run volume on Cloud Run and exits on a stuck launcher.
- The caller check reads X-Serverless-Authorization when present, in the
renderer and in agents/main.py.
- Renderer reasons reach R1 with the measured value and the limit; a
launcher failure is RendererUnavailable. Failures are logged.
- CI builds the renderer image and renders through its harness.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
Duplicate render execution and a short-lived watchdog-budget bypass must be resolved before approval.
6 open findings
Enforce final resource budget check after process exit · New Make render job IDs idempotent across retries · New Do not report a single attempt as two failures · New Narrow disk isolation heading to the run directory · New Correct containment claim for private sandbox tmp · New Update Dockerfile count from two to three · New
What changed in this PR
Adds a dedicated Cloud Run sandbox renderer and makes it the default agent rendering backend.
Changes:
- Adds the authenticated renderer service, sandbox executor, image, and deployment pipeline.
- Adds remote rendering with ID-token authentication, retries, cancellation, and limit feedback.
- Expands CI, tests, documentation, and caller-header hardening.
| File | Description |
|---|---|
agents/renderer/wire.py |
Defines renderer wire contracts. |
agents/renderer/settings.py |
Adds renderer configuration and limits. |
agents/renderer/main.py |
Implements renderer API and scheduling. |
agents/renderer/executor.py |
Runs and monitors sandbox processes. |
agents/renderer/auth.py |
Validates IAM-forwarded caller claims. |
agents/renderer/Dockerfile |
Builds the renderer runtime image. |
agents/renderer/cloudbuild.yaml |
Deploys, smokes, and promotes revisions. |
agents/main.py |
Hardens caller-header selection. |
agents/README.md |
Documents renderer development and deployment. |
agents/anyplot/settings.py |
Makes remote rendering configurable and default. |
agents/anyplot/render/__init__.py |
Constructs the remote backend. |
agents/anyplot/render/contract.py |
Carries renderer limit details. |
agents/anyplot/render/gates.py |
Adds limit-specific repair feedback. |
agents/anyplot/render/harness.py |
Adds limits, cache seeding, and usage reporting. |
agents/anyplot/render/runtimes/python.py |
Configures the matplotlib cache seed. |
agents/anyplot/render/backends/remote.py |
Implements remote calls, tokens, retries, and cancellation. |
agents/anyplot/render/backends/sandbox.py |
Reframes the in-process backend as a stub. |
agents/anyplot/render/backends/local.py |
Targets the renderer image locally. |
agents/anyplot/render/backends/__init__.py |
Updates backend documentation. |
.github/workflows/ci-image.yml |
Adds renderer image smoke testing. |
tests/unit/agents/test_settings.py |
Tests remote-renderer settings. |
tests/unit/agents/runtime/test_stream.py |
Tests serverless authorization precedence. |
tests/unit/agents/runtime/test_render.py |
Tests remote backend construction. |
tests/unit/agents/renderer/test_wire.py |
Tests contracts and package isolation. |
tests/unit/agents/renderer/test_service.py |
Tests renderer routes and cancellation. |
tests/unit/agents/renderer/test_remote_backend.py |
Tests remote rendering and authentication. |
tests/unit/agents/renderer/test_executor.py |
Tests sandbox execution and limits. |
tests/unit/agents/renderer/helpers.py |
Provides renderer test doubles. |
tests/unit/agents/renderer/conftest.py |
Provides hermetic executor fixtures. |
tests/unit/agents/renderer/__init__.py |
Defines the renderer test package. |
docs/workflows/overview.md |
Documents renderer image CI. |
docs/reference/repository.md |
Adds renderer repository structure. |
docs/index.md |
Updates agent-network status. |
changelog.d/agents-renderer.md |
Records the new service and backend. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…r cannot slip past the watchdog Copilot's findings on #12115: - `job_id` and the requested themes are now an idempotency key in the renderer. The remote backend and the deploy smoke replay a POST /render after an ambiguous failure, and the first request may have arrived: a replay with the same payload joins the render in flight or gets the stored answer of a finished one (the last 16 successes), another payload under the same key is refused with 409 job_conflict, and only the last waiter to leave cancels a render. The key includes the themes because the backend sends one request per theme under the job's id. - The watchdog samples between sleeps, so a run that wrote past the byte or entry budget and exited before the next sample was accepted. The run directory is checked once more after the code exits; such a run is refused like a live one, with nothing collected and its sandbox deleted. - The backend's "failed twice" message now says once or twice by `attempts`. - The executor heading, the fragment and the workflow overview describe the run volume and the sandbox's private /tmp as they are. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
3 of 5 tasks
…ata server
`gcloud auth print-identity-token --include-email` refuses Cloud Build's own
identity ("Invalid account type for --include-email. Requires an impersonate
service account", build da6f53f5 on 2026-10-10), so the smoke could never
pass with `_SMOKE_RENDER=true`. The token now comes from the metadata
server with `format=full`, which carries the `email` claim the caller check
reads. The same build verified the rest of the configured path: the
`--no-traffic` candidate deploy and the repeated `--add-volume` flags.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…vice's own account Cloud Build's metadata server has no `identity` endpoint (build a9d3ac7f answered 404), and gcloud refuses `--include-email` for the build's own identity, so a build cannot mint an ID token for itself. The smoke now impersonates the renderer's service account (`gcloud auth print-identity-token --impersonate-service-account`), which needs roles/iam.serviceAccountTokenCreator on that account for the build's account, roles/run.invoker on the service for the service account, and the service account in _ALLOWED_CALLERS. The design doc's IAM list and owner task 14 carry the three grants. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…ox stays Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
MarkusNeusinger
added a commit
that referenced
this pull request
Oct 10, 2026
…aseline (#12116) ## Summary - Adds the model-regression harness (`agents/evals/matrix.py`): every selected eval case runs through the real `/v1` flow in process (same plugins, run queue, judge, pipeline and gates as a user's request) with the models the flags name and the render backend of `--renderer` (`remote`, `local`, `fake`); per case it records status, attempts, gate failures, validator rejections, edit-apply failures, the reviewer's verdict, LLM calls, tokens by kind, `model_version`, cost at list price and latencies, writes a JSON report and a Markdown summary, diffs against the pinned model's baseline over the shared cases, and exits non-zero on a regression. A preflight render, an outage stop, `--gcloud-token` renewal, `--budget-usd` and a 30-render review page with Accept/Reject export make one command answer spike X and every later model or prompt comparison; `report blind` merges two arms into one unnamed gallery and `report score` reads the judgements back. - Adds the spike-X fixtures: 120 synthetic cases (10 specs from 10 plot families, eligible in matplotlib and seaborn, times six datasets: renamed headers, values times 10, 12 rows, 5,000 rows, `DD.MM.YYYY` and ISO dates, semicolon text with decimal commas) from the seeded `make_fixtures.py`, which a test keeps in sync with the committed files; plus `agents/evals/eligibility.py`, a model-free sweep that classifies every matplotlib and seaborn file of the catalogue; plus the first committed baseline, `agents/evals/baselines/claude-haiku-5-5.json`, from the spike-X run below. - The runtime now writes content-free attribution fields for every pipeline step (adapter outcome, validator rule ids, gate ids with render times, reviewer defect ids, tokens by kind, judge tokens) so the harness can price and count from the real flow; `.dockerignore` keeps `agents/evals/` out of every image. ## Spike X, Claude Haiku 5.5 arm (2026-10-10) Run from a scratch merge of this branch with #12115 against the throwaway renderer of #12115, `--provider anthropic-vertex --location eu --cases full --renderer remote --gcloud-token --budget-usd 4 --no-baseline --save-baseline`. The baseline in this PR is this run. | Metric | Value | |---|---| | Runs | 122, no harness or pipeline errors, no outage | | Pass rate (host and probe gates) | 83.6 % (102 of 122), 93 of them after the one repair round | | Statuses | `ok` 14 (11.5 %), `needs_attention` 98, `failed` 10 (all `validation`) | | Cost | $0.52 at list price for the matrix, $0.0051 per passed plot | | End to end p50 / p95 | 18.5 s / 25.1 s (render p50 4.0 s, time to first event p50 0.5 s) | | LLM calls per run | 4.77 | By spec: `line-basic` and `line-timeseries` 100 %, `scatter-basic` 92 %, `area-basic` and `histogram-basic` 92 %, `box-basic` 50 % (5 of 6 matplotlib runs end in `validation`: the adapter replaces the `df = load_user_data()` placeholder, rule `placeholder-count`), `pie-basic` 58 %. By perturbation: 12 rows weakest at 70 %, values times 10 best at 95 %. matplotlib 78.7 %, seaborn 88.5 %. What keeps a passed run at `needs_attention` instead of `ok`, in order: the probe gate G3 (text beyond the canvas) fails in 60 runs and reports in 56 more, and the reviewer's AR-09 lines agree in most of them; the reviewer's answer could not be read in 18 runs; the adapter's answer missed its schema in 23 of 141 calls, each costing the repair round; the IMPRINT palette rule (VQ-07) stands in 11 residual lines. These are the first items for the prompt and validator work on the Claude arm, not for this PR. The exit criterion of phase 0 asks for at least 80 % passing after at most one repair (met) and the owner's acceptance of at least 75 % on the 30-render sample (open: the gallery is with the owner). Eligibility sweep (no model call, about 9 s): matplotlib 285 of 325 files eligible, seaborn 287 of 325; blocked by SECURITY findings 29 / 29, the map specs 13 / 13, a missing `THEME` block 16 / 11, a savefig target other than `f"plot-{THEME}.png"` 16 / 11. ## Plan Design: `docs/concepts/agent-network.md` ("Model-version regression harness", spike X in "Phase 0"). The 120-case matrix needs the `remote` backend of #12115 to render on Cloud Run; the harness itself merges independently (its tests use scripted models and the fake renderer). ## Test plan - [x] `uv run ruff check . && ruff format --check .`, `mypy api core agents`, `pytest tests/unit tests/integration` (7888 passed before the baseline; with it the baseline test runs: `tests/unit/agents/evals` 81 passed), changelog check, `uv lock --check` - [x] Eligibility sweep on the current catalogue (numbers above) - [x] Spike X on the Claude arm through the throwaway renderer (numbers above); the committed baseline is that run - [ ] Owner: open the review page (`gallery.html` of the run, 30 renders that passed the gates), judge each one, export the judgements; the acceptance rate is the phase-0 criterion (75 %) - [ ] The Gemini arm (`--provider gemini --budget-usd 18`) and the blind two-arm gallery wait for the owner's go: about $10 to $15 at list price --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Summary
anyplot-renderer(agents/renderer/): a small FastAPI service withPOST /render,POST /render/{job_id}/cancelandGET /statusbehind Cloud Run IAM and the ID-token claims check, that runs each theme of a job in its own Cloud Run sandbox (sandbox do, never--writeor egress, explicitPATH), one at a time, with the limits spikes S and S2 measured: run directories on a 512 MiB in-memory volume the service requires on Cloud Run (503 volume_missingwithout it), a watchdog that kills a run past 64 MiB, 10,000 files or a 512 MiBMemAvailablefloor, a kill that counts only once the launcher has exited, one retry when the launcher fails before the harness starts, bounded output, a cleaned probe, rlimits, a stop on client disconnect or cancel, JSON500 internalwithout the message, and a self-restart when a launcher outlives its kill. It only runs code; the host gates stay in the agents service. The image installs the plotting libraries without ADK;agents/renderer/cloudbuild.yamlbuilds, deploys acandidate, smokes and promotes it like the API, and creates the service on the first build only whendescribereports it missing.remoterender backend, now the default (AGENT_RENDERER=remote,AGENT_RENDER_URL), so a localadk weband the deployed agents service render in the same sandboxes: an ID token from the metadata server when deployed, or fromAGENT_RENDER_TOKENor the developer's Application Default Credentials in development; retries over a cold start (1, 3, 8, 15 s within 30 s) for connection errors and front-end 429/5xx only, never for the renderer's own answers; a cancel call for abandoned renders; limit and measured value in the repair feedback.X-Serverless-Authorizationwins whenever it is present, because Cloud Run then checks only that header and passesAuthorizationthrough unverified. The harness prints a start and an end line and sets soft and hard limits alike. CI gains arenderer-imagejob (ci-image.yml) that builds the image and renders a seaborn plot through the harness with no network.Plan
Design:
docs/concepts/agent-network.md("Render", "Serving and infrastructure"; status line and service table updated). The renderer split for phase 1 was decided on 2026-10-10 (one renderer for the deployed agents and foradk web). Follows #12111 and #12112.Test plan
uv run ruff check . && ruff format --check .,mypy api core agents(101 files),pytest tests/unit tests/integration(7935 passed; executor tests run the real harness against a fake launcher: kill path, cancel, disconnect, stuck launcher, retry once and twice, byte, file and memory watchdog, bounded output, refused output files; service tests cover the caller check, the slot, low memory, the volume check and/status; theremotebackend is tested against the in-process app and once end to end through the real executor and host gates), changelog check,uv lock --checkanyplot-renderer-spike, a role-less service account, Cloud Build5b86cc7afrom this head with_SMOKE_RENDER=false, 3 min 51 s): the first-build path created the service;gcloud beta run deploy --sandbox-launcherand the repeated--add-volumeflags are accepted by the Cloud Build gcloud image (587.0.0); an anonymous call gets 403;/statuswith the owner's gcloud token reportssandbox: true,volume: true(the in-memory volume reports its 512 MiB limit), 3,524 MiB available; a real matplotlib render answered in 2.9 s with a 108 KB PNG; spike X (122 cases) renders through it from a laptop with--gcloud-token_SMOKE_RENDER=true(Cloud Builda89a46b2from09f50b62d, 3 min 47 s): the--no-trafficcandidate path, the smoke as the service's own account (impersonated; a build cannot mint an ID token for itself, which buildsda6f53f5anda9d3ac7fshowed and the two fix commits address),/statuswith the volume in place, a real render in 3.2 s, promote. Therenderer-imageCI job ran and passed on every push./tmpfill through the promoted revision (the review's open question): code writing 40 MiB files into the sandbox's private/tmpends after 4.6 s withOSError: [Errno 28] No space left on device; the instance stays up (3,451 MiB available before, 3,328 after), the memory watchdog never fires. Decided in the design doc: 4 GiB with one sandbox at a time, noRENDERER_BIND_TMP, no 8 GiB instance. The throwaway service, its service account and its images are deleted..github/workflows/has no pre-merge loopanyplot-rendererservice account, the first build and thedeploy-renderertrigger are owner tasks (design doc, task 14)