Skip to content

feat(agents): anyplot-renderer service and the remote render backend - #12115

Merged
MarkusNeusinger merged 8 commits into
mainfrom
feat/agents-renderer
Oct 10, 2026
Merged

MarkusNeusinger merged 8 commits into
mainfrom
feat/agents-renderer

Conversation

@MarkusNeusinger

@MarkusNeusinger MarkusNeusinger commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Adds anyplot-renderer (agents/renderer/): a small FastAPI service with POST /render, POST /render/{job_id}/cancel and GET /status behind Cloud Run IAM and the ID-token claims check, that runs each theme of a job in its own Cloud Run sandbox (sandbox do, never --write or egress, explicit PATH), one at a time, with the limits spikes S and S2 measured: run directories on a 512 MiB in-memory volume the service requires on Cloud Run (503 volume_missing without it), a watchdog that kills a run past 64 MiB, 10,000 files or a 512 MiB MemAvailable floor, a kill that counts only once the launcher has exited, one retry when the launcher fails before the harness starts, bounded output, a cleaned probe, rlimits, a stop on client disconnect or cancel, JSON 500 internal without the message, and a self-restart when a launcher outlives its kill. It only runs code; the host gates stay in the agents service. The image installs the plotting libraries without ADK; agents/renderer/cloudbuild.yaml builds, deploys a candidate, smokes and promotes it like the API, and creates the service on the first build only when describe reports it missing.
  • Adds the remote render backend, now the default (AGENT_RENDERER=remote, AGENT_RENDER_URL), so a local adk web and the deployed agents service render in the same sandboxes: an ID token from the metadata server when deployed, or from AGENT_RENDER_TOKEN or the developer's Application Default Credentials in development; retries over a cold start (1, 3, 8, 15 s within 30 s) for connection errors and front-end 429/5xx only, never for the renderer's own answers; a cancel call for abandoned renders; limit and measured value in the repair feedback.
  • Hardens the caller check in both services: X-Serverless-Authorization wins whenever it is present, because Cloud Run then checks only that header and passes Authorization through unverified. The harness prints a start and an end line and sets soft and hard limits alike. CI gains a renderer-image job (ci-image.yml) that builds the image and renders a seaborn plot through the harness with no network.

Plan

Design: docs/concepts/agent-network.md ("Render", "Serving and infrastructure"; status line and service table updated). The renderer split for phase 1 was decided on 2026-10-10 (one renderer for the deployed agents and for adk web). Follows #12111 and #12112.

Test plan

  • uv run ruff check . && ruff format --check ., mypy api core agents (101 files), pytest tests/unit tests/integration (7935 passed; executor tests run the real harness against a fake launcher: kill path, cancel, disconnect, stuck launcher, retry once and twice, byte, file and memory watchdog, bounded output, refused output files; service tests cover the caller check, the slot, low memory, the volume check and /status; the remote backend is tested against the in-process app and once end to end through the real executor and host gates), changelog check, uv lock --check
  • Deployed to Cloud Run as a throwaway service (anyplot-renderer-spike, a role-less service account, Cloud Build 5b86cc7a from this head with _SMOKE_RENDER=false, 3 min 51 s): the first-build path created the service; gcloud beta run deploy --sandbox-launcher and the repeated --add-volume flags are accepted by the Cloud Build gcloud image (587.0.0); an anonymous call gets 403; /status with the owner's gcloud token reports sandbox: true, volume: true (the in-memory volume reports its 512 MiB limit), 3,524 MiB available; a real matplotlib render answered in 2.9 s with a 108 KB PNG; spike X (122 cases) renders through it from a laptop with --gcloud-token
  • Configured build on the throwaway with _SMOKE_RENDER=true (Cloud Build a89a46b2 from 09f50b62d, 3 min 47 s): the --no-traffic candidate path, the smoke as the service's own account (impersonated; a build cannot mint an ID token for itself, which builds da6f53f5 and a9d3ac7f showed and the two fix commits address), /status with the volume in place, a real render in 3.2 s, promote. The renderer-image CI job ran and passed on every push.
  • /tmp fill through the promoted revision (the review's open question): code writing 40 MiB files into the sandbox's private /tmp ends after 4.6 s with OSError: [Errno 28] No space left on device; the instance stays up (3,451 MiB available before, 3,328 after), the memory watchdog never fires. Decided in the design doc: 4 GiB with one sandbox at a time, no RENDERER_BIND_TMP, no 8 GiB instance. The throwaway service, its service account and its images are deleted.
  • Not verifiable before merge: the Application Default Credentials token path against Cloud Run IAM; .github/workflows/ has no pre-merge loop
  • After merge nothing deploys: the real anyplot-renderer service account, the first build and the deploy-renderer trigger are owner tasks (design doc, task 14)

MarkusNeusinger and others added 2 commits October 10, 2026 01:20
The renderer runs adapted plot code in Cloud Run sandboxes as a service of
its own (owner decision of 2026-10-10), so a local `adk web` and the deployed
agents service share one renderer.

- agents/renderer/: FastAPI POST /render and GET /status behind the IAM
  caller check (copied from agents/main.py; RENDERER_AUDIENCES and
  RENDERER_ALLOWED_CALLERS), one render slot with a bounded wait, a
  MemAvailable floor, and the sandbox executor built on spikes S and S2:
  no --write or --allow-egress, an explicit PATH, run directories under
  /tmp/runs with a byte and entry watchdog, killpg + `sandbox delete
  --force` confirmed by the launcher's exit (stuck launchers block renders),
  one retry with a new sandbox name when the launcher fails before the
  harness starts, bounded stdout/stderr, O_NOFOLLOW + fstat output reads.
  wire.py is the JSON contract; no ADK, agents.anyplot or core import, and
  no __init__.py so `adk web agents` does not list it as an agent.
- Dockerfile (plotting venv without ADK, baked font cache, harness at
  /opt/anyplot/harness.py, runs as root for the launcher) and
  cloudbuild.yaml (candidate deploy that bootstraps the service on the
  first build, IAM 403 smoke and an optional real render smoke, promote).
- agents/anyplot/render/backends/remote.py: the remote backend, now the
  AGENT_RENDERER default, with an ID-token source (metadata server when
  deployed; AGENT_RENDER_TOKEN or the user's ADC id_token in development),
  one retry on connection errors and Cloud Run front-end 429/5xx, and every
  other failure mapped to RendererUnavailable. The in-process sandbox
  backend stays as an unused stub.
- harness.py: HARNESS start/end lines, soft = hard rlimits, opt-in
  RLIMIT_AS and RLIMIT_NPROC, the baked font cache copied into MPLCONFIGDIR.
- Tests for the executor (fake launcher, real harness), the service, the
  contract and the remote backend in process; docs and changelog fragment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Apply the adversarial review of the renderer service and the remote backend:

- The deploy smoke mints its token for the service URL (Cloud Run accepts
  no tag URL as an audience); the deploy adds status.url to
  RENDERER_AUDIENCES itself and creates the service only on NOT_FOUND.
- Plot output can no longer crash /render: a forged harness end line, a
  probe nested too deep or holding an oversized number, and host OSErrors
  (503 io) are handled; any other error answers a JSON 500 internal, which
  the remote backend never retries.
- The remote backend rides out a cold start (1, 3, 8, 15 s within 30 s)
  and posts /render/{job_id}/cancel when its render is cancelled or times
  out; the service races each render against http.disconnect.
- The watchdog also kills a run below a MemAvailable floor (memory), counts
  an unmeasurable tree as over budget (file_budget), and the service
  requires the run volume on Cloud Run and exits on a stuck launcher.
- The caller check reads X-Serverless-Authorization when present, in the
  renderer and in agents/main.py.
- Renderer reasons reach R1 with the measured value and the limit; a
  launcher failure is RendererUnavailable. Failures are logged.
- CI builds the renderer image and renders through its harness.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
Copilot AI balanced review requested due to automatic review settings October 10, 2026 00:35
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Duplicate render execution and a short-lived watchdog-budget bypass must be resolved before approval.

6 open findings
What changed in this PR

Adds a dedicated Cloud Run sandbox renderer and makes it the default agent rendering backend.

Changes:

  • Adds the authenticated renderer service, sandbox executor, image, and deployment pipeline.
  • Adds remote rendering with ID-token authentication, retries, cancellation, and limit feedback.
  • Expands CI, tests, documentation, and caller-header hardening.
File Description
agents/​renderer/​wire.py Defines renderer wire contracts.
agents/​renderer/​settings.py Adds renderer configuration and limits.
agents/​renderer/​main.py Implements renderer API and scheduling.
agents/​renderer/​executor.py Runs and monitors sandbox processes.
agents/​renderer/​auth.py Validates IAM-forwarded caller claims.
agents/​renderer/​Dockerfile Builds the renderer runtime image.
agents/​renderer/​cloudbuild.yaml Deploys, smokes, and promotes revisions.
agents/​main.py Hardens caller-header selection.
agents/​README.md Documents renderer development and deployment.
agents/​anyplot/​settings.py Makes remote rendering configurable and default.
agents/​anyplot/​render/​__init__.py Constructs the remote backend.
agents/​anyplot/​render/​contract.py Carries renderer limit details.
agents/​anyplot/​render/​gates.py Adds limit-specific repair feedback.
agents/​anyplot/​render/​harness.py Adds limits, cache seeding, and usage reporting.
agents/​anyplot/​render/​runtimes/​python.py Configures the matplotlib cache seed.
agents/​anyplot/​render/​backends/​remote.py Implements remote calls, tokens, retries, and cancellation.
agents/​anyplot/​render/​backends/​sandbox.py Reframes the in-process backend as a stub.
agents/​anyplot/​render/​backends/​local.py Targets the renderer image locally.
agents/​anyplot/​render/​backends/​__init__.py Updates backend documentation.
.github/​workflows/​ci-image.yml Adds renderer image smoke testing.
tests/​unit/​agents/​test_settings.py Tests remote-renderer settings.
tests/​unit/​agents/​runtime/​test_stream.py Tests serverless authorization precedence.
tests/​unit/​agents/​runtime/​test_render.py Tests remote backend construction.
tests/​unit/​agents/​renderer/​test_wire.py Tests contracts and package isolation.
tests/​unit/​agents/​renderer/​test_service.py Tests renderer routes and cancellation.
tests/​unit/​agents/​renderer/​test_remote_backend.py Tests remote rendering and authentication.
tests/​unit/​agents/​renderer/​test_executor.py Tests sandbox execution and limits.
tests/​unit/​agents/​renderer/​helpers.py Provides renderer test doubles.
tests/​unit/​agents/​renderer/​conftest.py Provides hermetic executor fixtures.
tests/​unit/​agents/​renderer/​__init__.py Defines the renderer test package.
docs/​workflows/​overview.md Documents renderer image CI.
docs/​reference/​repository.md Adds renderer repository structure.
docs/​index.md Updates agent-network status.
changelog.d/​agents-renderer.md Records the new service and backend.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread agents/renderer/executor.py Outdated
Comment thread agents/renderer/main.py Outdated
Comment thread agents/anyplot/render/backends/remote.py Outdated
Comment thread agents/renderer/executor.py Outdated
Comment thread changelog.d/agents-renderer.md
Comment thread docs/workflows/overview.md Outdated
@codecov

codecov Bot commented Oct 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…r cannot slip past the watchdog

Copilot's findings on #12115:

- `job_id` and the requested themes are now an idempotency key in the
  renderer. The remote backend and the deploy smoke replay a POST /render
  after an ambiguous failure, and the first request may have arrived: a
  replay with the same payload joins the render in flight or gets the stored
  answer of a finished one (the last 16 successes), another payload under
  the same key is refused with 409 job_conflict, and only the last waiter to
  leave cancels a render. The key includes the themes because the backend
  sends one request per theme under the job's id.
- The watchdog samples between sleeps, so a run that wrote past the byte or
  entry budget and exited before the next sample was accepted. The run
  directory is checked once more after the code exits; such a run is refused
  like a live one, with nothing collected and its sandbox deleted.
- The backend's "failed twice" message now says once or twice by `attempts`.
- The executor heading, the fragment and the workflow overview describe the
  run volume and the sandbox's private /tmp as they are.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
MarkusNeusinger and others added 4 commits October 10, 2026 03:11
…ata server

`gcloud auth print-identity-token --include-email` refuses Cloud Build's own
identity ("Invalid account type for --include-email. Requires an impersonate
service account", build da6f53f5 on 2026-10-10), so the smoke could never
pass with `_SMOKE_RENDER=true`. The token now comes from the metadata
server with `format=full`, which carries the `email` claim the caller check
reads. The same build verified the rest of the configured path: the
`--no-traffic` candidate deploy and the repeated `--add-volume` flags.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…vice's own account

Cloud Build's metadata server has no `identity` endpoint (build a9d3ac7f
answered 404), and gcloud refuses `--include-email` for the build's own
identity, so a build cannot mint an ID token for itself. The smoke now
impersonates the renderer's service account (`gcloud auth
print-identity-token --impersonate-service-account`), which needs
roles/iam.serviceAccountTokenCreator on that account for the build's account,
roles/run.invoker on the service for the service account, and the service
account in _ALLOWED_CALLERS. The design doc's IAM list and owner task 14
carry the three grants.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
…ox stays

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdMSqLR5ww4ji74EUSmk9M
@MarkusNeusinger
MarkusNeusinger merged commit 4cc6993 into main Oct 10, 2026
11 checks passed
@MarkusNeusinger
MarkusNeusinger deleted the feat/agents-renderer branch October 10, 2026 07:07
MarkusNeusinger added a commit that referenced this pull request Oct 10, 2026
…aseline (#12116)

## Summary
- Adds the model-regression harness (`agents/evals/matrix.py`): every
selected eval case runs through the real `/v1` flow in process (same
plugins, run queue, judge, pipeline and gates as a user's request) with
the models the flags name and the render backend of `--renderer`
(`remote`, `local`, `fake`); per case it records status, attempts, gate
failures, validator rejections, edit-apply failures, the reviewer's
verdict, LLM calls, tokens by kind, `model_version`, cost at list price
and latencies, writes a JSON report and a Markdown summary, diffs
against the pinned model's baseline over the shared cases, and exits
non-zero on a regression. A preflight render, an outage stop,
`--gcloud-token` renewal, `--budget-usd` and a 30-render review page
with Accept/Reject export make one command answer spike X and every
later model or prompt comparison; `report blind` merges two arms into
one unnamed gallery and `report score` reads the judgements back.
- Adds the spike-X fixtures: 120 synthetic cases (10 specs from 10 plot
families, eligible in matplotlib and seaborn, times six datasets:
renamed headers, values times 10, 12 rows, 5,000 rows, `DD.MM.YYYY` and
ISO dates, semicolon text with decimal commas) from the seeded
`make_fixtures.py`, which a test keeps in sync with the committed files;
plus `agents/evals/eligibility.py`, a model-free sweep that classifies
every matplotlib and seaborn file of the catalogue; plus the first
committed baseline, `agents/evals/baselines/claude-haiku-5-5.json`, from
the spike-X run below.
- The runtime now writes content-free attribution fields for every
pipeline step (adapter outcome, validator rule ids, gate ids with render
times, reviewer defect ids, tokens by kind, judge tokens) so the harness
can price and count from the real flow; `.dockerignore` keeps
`agents/evals/` out of every image.

## Spike X, Claude Haiku 5.5 arm (2026-10-10)
Run from a scratch merge of this branch with #12115 against the
throwaway renderer of #12115, `--provider anthropic-vertex --location eu
--cases full --renderer remote --gcloud-token --budget-usd 4
--no-baseline --save-baseline`. The baseline in this PR is this run.

| Metric | Value |
|---|---|
| Runs | 122, no harness or pipeline errors, no outage |
| Pass rate (host and probe gates) | 83.6 % (102 of 122), 93 of them
after the one repair round |
| Statuses | `ok` 14 (11.5 %), `needs_attention` 98, `failed` 10 (all
`validation`) |
| Cost | $0.52 at list price for the matrix, $0.0051 per passed plot |
| End to end p50 / p95 | 18.5 s / 25.1 s (render p50 4.0 s, time to
first event p50 0.5 s) |
| LLM calls per run | 4.77 |

By spec: `line-basic` and `line-timeseries` 100 %, `scatter-basic` 92 %,
`area-basic` and `histogram-basic` 92 %, `box-basic` 50 % (5 of 6
matplotlib runs end in `validation`: the adapter replaces the `df =
load_user_data()` placeholder, rule `placeholder-count`), `pie-basic` 58
%. By perturbation: 12 rows weakest at 70 %, values times 10 best at 95
%. matplotlib 78.7 %, seaborn 88.5 %.

What keeps a passed run at `needs_attention` instead of `ok`, in order:
the probe gate G3 (text beyond the canvas) fails in 60 runs and reports
in 56 more, and the reviewer's AR-09 lines agree in most of them; the
reviewer's answer could not be read in 18 runs; the adapter's answer
missed its schema in 23 of 141 calls, each costing the repair round; the
IMPRINT palette rule (VQ-07) stands in 11 residual lines. These are the
first items for the prompt and validator work on the Claude arm, not for
this PR. The exit criterion of phase 0 asks for at least 80 % passing
after at most one repair (met) and the owner's acceptance of at least 75
% on the 30-render sample (open: the gallery is with the owner).

Eligibility sweep (no model call, about 9 s): matplotlib 285 of 325
files eligible, seaborn 287 of 325; blocked by SECURITY findings 29 /
29, the map specs 13 / 13, a missing `THEME` block 16 / 11, a savefig
target other than `f"plot-{THEME}.png"` 16 / 11.

## Plan
Design: `docs/concepts/agent-network.md` ("Model-version regression
harness", spike X in "Phase 0"). The 120-case matrix needs the `remote`
backend of #12115 to render on Cloud Run; the harness itself merges
independently (its tests use scripted models and the fake renderer).

## Test plan
- [x] `uv run ruff check . && ruff format --check .`, `mypy api core
agents`, `pytest tests/unit tests/integration` (7888 passed before the
baseline; with it the baseline test runs: `tests/unit/agents/evals` 81
passed), changelog check, `uv lock --check`
- [x] Eligibility sweep on the current catalogue (numbers above)
- [x] Spike X on the Claude arm through the throwaway renderer (numbers
above); the committed baseline is that run
- [ ] Owner: open the review page (`gallery.html` of the run, 30 renders
that passed the gates), judge each one, export the judgements; the
acceptance rate is the phase-0 criterion (75 %)
- [ ] The Gemini arm (`--provider gemini --budget-usd 18`) and the blind
two-arm gallery wait for the owner's go: about $10 to $15 at list price

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants