Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 46 additions & 5 deletions agents/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ This directory holds the anyplot agent network: the service that lets an admin p

## What is built

The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer, the private `/v1` service with its run queue, the theme toggle, and the footer strip on every PNG the service serves ("made with any.plot()" and "anyplot.ai/<spec-id>", see [Footer strip](../docs/concepts/agent-network.md#footer-strip)). The renderer service `anyplot-renderer` (`renderer/`), which runs the adapted code in Cloud Run sandboxes, and the `remote` backend that calls it are built, with the renderer's image and Cloud Build config, but not deployed yet (status of 2026-10-10). The model-regression harness, the 120 synthetic spike-X cases, the blind two-run review gallery and the catalogue eligibility sweep are built (see [Run the regression harness](#run-the-regression-harness)), and the Claude Haiku 5.5 baseline from the rerun of spike X on main in the evening of 2026-10-10 is committed (`evals/baselines/claude-haiku-5-5.json`: 244 runs, 90.6 % passed the gates, 23.4 % ended `ok`); the Gemini baseline is not. The chat page that uses the service through the API's `/debug/agent` routes is in the app (`app/src/pages/AgentChatPage.tsx`, built with `VITE_ENABLE_AGENT_CHAT=true`). The service's own image (`Dockerfile`) and Cloud Build config (`cloudbuild.yaml`) are built, with a CI job that builds and smokes the image before merge, but the service is not deployed yet (see [Build and deploy the service](#build-and-deploy-the-service)). Not built yet: the deploy of either service, the scope evalset, and the `agents-eval.yml` workflow.
The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer, the private `/v1` service with its run queue, the theme toggle, and the footer strip on every PNG the service serves ("made with any.plot()" and "anyplot.ai/<spec-id>", see [Footer strip](../docs/concepts/agent-network.md#footer-strip)). The renderer service `anyplot-renderer` (`renderer/`), which runs the adapted code in Cloud Run sandboxes, and the `remote` backend that calls it are built, with the renderer's image and Cloud Build config, but not deployed yet (status of 2026-10-10). The model-regression harness, the 120 synthetic spike-X cases, the 23 stress cases with their model-free static stage, the blind two-run review gallery and the catalogue eligibility sweep are built (see [Run the regression harness](#run-the-regression-harness)), and the Claude Haiku 5.5 baseline from the rerun of spike X on main in the evening of 2026-10-10 is committed (`evals/baselines/claude-haiku-5-5.json`: 244 runs, 90.6 % passed the gates, 23.4 % ended `ok`); the Gemini baseline is not. The chat page that uses the service through the API's `/debug/agent` routes is in the app (`app/src/pages/AgentChatPage.tsx`, built with `VITE_ENABLE_AGENT_CHAT=true`). The service's own image (`Dockerfile`) and Cloud Build config (`cloudbuild.yaml`) are built, with a CI job that builds and smokes the image before merge, but the service is not deployed yet (see [Build and deploy the service](#build-and-deploy-the-service)). Not built yet: the deploy of either service, the scope evalset, and the `agents-eval.yml` workflow.

The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default. **Gemini 3.8 Flash** is the second arm: set `AGENT_PROVIDER=gemini` together with Gemini model ids, so the two can be compared on price and quality later. Every agent and the scope judge run on the configured provider.

Expand All @@ -25,7 +25,7 @@ The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default.
| `anyplot/sub_agents/` | One single-turn adapter per enabled library and the tool-less reviewer |
| `anyplot/tools/session.py` | The root's tools: `get_dataset_profile`, `get_spec_brief`, `get_current_code`, `set_bindings` and the `plot_pipeline` workflow |
| `anyplot/plugins/` | `ScopeGuardPlugin`, `BudgetPlugin`, `ToolSafetyPlugin` and the request ledger they share |
| `anyplot/render/` | The render contract, the probe harness, the host gates (R1-R3, advisory G3/G5/G7/G8), PNG hardening, the render store, the `remote` (the renderer service, phase 1), `local`, `fake` and `sandbox` (an unused in-process stub) backends, `serial.py`, the one render semaphore in front of every backend, and `watermark.py`, the footer strip the artifact route and the feedback bundle add to the raw render |
| `anyplot/render/` | The render contract, the probe harness, the host gates (R1-R3, advisory G3/G5/G7/G8/G9), PNG hardening, the render store, the `remote` (the renderer service, phase 1), `local`, `fake` and `sandbox` (an unused in-process stub) backends, `serial.py`, the one render semaphore in front of every backend, and `watermark.py`, the footer strip the artifact route and the feedback bundle add to the raw render |
| `anyplot/render/fonts/` | JetBrains Mono 2.304 Regular and Bold for the footer strip, unmodified, with `OFL.txt` (SIL Open Font License 1.1) and a notice with their SHA-256 values ([fonts README](anyplot/render/fonts/README.md)) |
| `renderer/` | The `anyplot-renderer` service: `main.py` (`POST /render`, `GET /status`, one render slot), `executor.py` (one `sandbox do` per theme with the kill path, the byte watchdog and the launcher retry), `auth.py` (the caller check), `settings.py` (`RENDERER_*`), `wire.py` (the JSON contract the `remote` backend shares), `Dockerfile` and `cloudbuild.yaml`. No ADK import and no `__init__.py`, so `adk web agents` does not list it as an agent |
| `anyplot/opening.py`, `session_state.py`, `services.py`, `briefs.py` | Opening a session and taking in a dataset, the server-set session state, the process-wide stores, and the spec and dataset briefs |
Expand All @@ -35,9 +35,10 @@ The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default.
| `anyplot/data/`, `anyplot/code/` | The deterministic data and code layers, without ADK |
| `evals/matrix.py` | The regression harness: runs eval cases through `/v1` in process and writes the report, the Markdown summary, the diff against a baseline and the review galleries |
| `evals/make_fixtures.py` | The seeded generator of the 120 synthetic spike-X cases |
| `evals/make_stress.py`, `evals/stress.py` | The seeded generator of the 23 stress cases, and the model-free static stage that `--static` runs over them ([Run the stress cases](#run-the-stress-cases)) |
| `evals/cases.py`, `evals/pricing.py`, `evals/report.py` | Case loading and selection, list prices, and the summary, diff and gallery code; `report.py` also merges two runs into one blind gallery and scores its export |
| `evals/eligibility.py` | The model-free sweep of the catalogue's matplotlib and seaborn files: eligible, coupled and blocked pairs, and why |
| `evals/fixtures/cases/` | The eval cases: 120 generated (`<spec>-<library>-<perturbation>`) and the hand-written `scatter-basic-matplotlib` and `bar-grouped-seaborn` ([fixtures README](evals/fixtures/README.md)) |
| `evals/fixtures/cases/` | The eval cases: 120 generated (`<spec>-<library>-<perturbation>`), the hand-written `scatter-basic-matplotlib` and `bar-grouped-seaborn`, and 23 stress cases (`stress-<name>`) ([fixtures README](evals/fixtures/README.md)) |
| `evals/baselines/` | Committed baseline reports, one per model ([baselines README](evals/baselines/README.md)) |

The unit tests are in `tests/unit/agents/`; `tests/unit/agents/runtime/` drives a full "Create plot" through `/v1` with a fake renderer and scripted models for both providers, and `tests/unit/agents/evals/` runs the harness the same way.
Expand Down Expand Up @@ -340,7 +341,7 @@ Before the first case the harness renders the catalogue file of the first case o
--renderer remote --render-url https://RENDERER_URL --gcloud-token
```

- The full matrix is all 122 cases. Cap the spend with `--budget-usd`:
- The full matrix is the 122 cases that are not stress cases. Cap the spend with `--budget-usd`:

```bash
uv run --extra agents python -m agents.evals.matrix --cases full --budget-usd 4 \
Expand All @@ -356,7 +357,8 @@ Before the first case the harness renders the catalogue file of the first case o
| `--renderer`, `--render-url` | The render backend (`remote`, `local` or `fake`) and the remote renderer's URL |
| `--gcloud-token` | Mint the renderer token with `gcloud auth print-identity-token` and renew it before it expires (`remote` only) |
| `--no-preflight` | Skip the preflight render before the first case |
| `--cases` | `smoke` (default), `full`, or comma-separated globs over case ids |
| `--cases` | `smoke` (default), `full` (every case but the stress cases), `stress`, or comma-separated globs over case ids |
| `--static` | Run only the model-free static stage and write `<date>-static.json` instead of a matrix report (see [Run the stress cases](#run-the-stress-cases)); the renderer is `fake` unless `--renderer` or `AGENT_RENDERER` names another |
| `--repeats N` | Runs per case, to see how stable a result is |
| `--max-attempts N` | Adapter attempts per run, 1 to 3 (`AGENT_MAX_ATTEMPTS`, 2 by default), so 1 measures the first attempt alone and 3 a second repair round; the report's stamp records the value as `max_attempts` |
| `--budget-usd X` | Stop before a case that could take the estimated list-price cost past X (the spend so far plus the costliest case so far), or once the cost has passed X (exit code 3) |
Expand Down Expand Up @@ -397,6 +399,45 @@ The exit code is one of the following:

Ctrl-C, exit codes 3 and 4 all write the partial report before the harness stops.

### Run the stress cases

The 23 stress cases (`evals/fixtures/cases/stress-*`, from `evals/make_stress.py`) push every limit before real use: data exactly at and one past the parser's caps, 100 bars and 30 pie slices, degenerate data, prompt injection, and requests that make the code compute mass data. Each `case.json` names the gate, limit or error that must bound the case (`bound_expected`); the [fixtures README](evals/fixtures/README.md#stress-cases-what-bounds-them) lists every bound with the numbers measured on 2026-10-10.

1. Run the static stage. It calls no model and spends nothing: for every case it runs the dataset route's size check and parse, eligibility, bindings, the loader's `pd.read_csv`, measures what the adapter, the dataset judge and the reviewer would read, traces injected text through those inputs, and renders the five hand-written stand-ins (`standin.txt`) for the adapter's answer:

```bash
uv run --extra agents python -m agents.evals.matrix --cases stress --static
```

It writes `<date>-static.json` and a Markdown table to `--out` and exits with 1 when a case's expectations (`static` in its `case.json`) did not hold. The default `fake` renderer runs no code, so the stand-ins' gates stay unchecked.

2. To see the stand-ins hit the real limits (G7, G9 and the sandbox's address space), render them on the deployed renderer, still without a model:

```bash
uv run --extra agents python -m agents.evals.matrix --cases stress --static \
--renderer remote --render-url https://RENDERER_URL --gcloud-token
```

With a `remote` or `local` renderer the static stage also checks each case's `render` expectations and reports the render time. Only the `remote` renderer reports the harness's peak memory; the `local` backend leaves that column empty. When the renderer becomes unavailable during a stand-in render, the run stops with exit code 4, as an outage stops the matrix.

3. Run the cases whose bound, or the outcome their bound leaves open, only a model run shows. The output names them as `needs a model run`:
- `stress-compute-lorenz`, `stress-compute-oversample` and `stress-compute-distances`: does the adapter write the mass-data code, and does G9 or the renderer stop it.
- `stress-pie-30-slices`: only the reviewer sees its piled labels.
- `stress-single-row` and `stress-constant-series`: how the plot copes with one point or a flat series.
- `stress-inject-long-cell`: whether the plot draws the injected cell as a tick label the reviewer reads.

This is a normal matrix run and spends tokens:

```bash
uv run --extra agents python -m agents.evals.matrix \
--cases "stress-compute-*,stress-pie-*,stress-single-row,stress-constant-series,stress-inject-long-cell" \
--no-baseline --budget-usd 1 --renderer remote --render-url https://RENDERER_URL --gcloud-token
```

`--cases stress` runs all 23 the same way. The four over-cap cases end at the dataset route before any model call, and the two bindings refusals after one judge call. The full matrix's measured cost per case puts the whole set at about $0.15 to $0.20 on the Claude arm and $3 to $4 on the Gemini arm; the computing cases send their change request as a second turn and cost about twice a normal case.

After you change `evals/make_stress.py`, regenerate the cases with `uv run --extra agents python -m agents.evals.make_stress` (`--check` exits with 1 when a committed file is stale or a stray file sits in a `stress-*` directory).

### Run spike X

Spike X is the go or no-go for phase 1: 10 specs times matplotlib and seaborn times 6 datasets (renamed, x10, n=12, n=5000, date, decimal comma) on both arms, under $25 in total. The full matrix is those 120 generated cases plus the two hand-written ones. At list price a full run is estimated at about $1 on the Claude arm and $10 to $15 on the Gemini arm. The budgets below leave $3 of the $25 for headroom, because a budget is checked between cases.
Expand Down
25 changes: 25 additions & 0 deletions agents/anyplot/briefs.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
from collections.abc import Iterator

from .data.roles import DataRole
from .policy import fence
from .schemas import DatasetProfile
from .session_state import SessionView

Expand All @@ -18,6 +19,12 @@
MAX_DESCRIPTION_CHARS = 1_500
MAX_NOTE_CHARS = 300
MAX_SUMMARY_CHARS = 2_000
MAX_ADAPTER_PROFILE_BYTES = 16 * 1024
"""The fenced dataset profile in an adapter request, at most, in UTF-8 bytes like the root's `get_dataset_profile` limit.

Bytes, not characters: a CJK or emoji cell is three or four bytes, so a character
count would let a profile of such cells grow the request several times past the cap.
"""


def describe_role(role: DataRole) -> str:
Expand Down Expand Up @@ -51,6 +58,24 @@ def profile_json(profile: DatasetProfile) -> str:
return profile.model_dump_json(exclude={"warnings"})


def adapter_profile(profile: DatasetProfile, limit: int = MAX_ADAPTER_PROFILE_BYTES) -> tuple[str, bool]:
"""The adapter's profile JSON: the first of `trimmed_profiles` within `limit` UTF-8 bytes, and whether it was trimmed.

The limit applies to the version as the adapter request carries it, inside its
`<user_data>` fence: `fence` escapes every fence tag in the data (`<user_data>`
becomes `&lt;user_data>`), so cells full of tag text grow the profile after a check
on the raw JSON. The last version (every column's name, type and counts) goes out
even over the limit, because the adapter cannot bind a column it never saw; at the
parser's 50 columns of at most 64 characters it stays far below it.
"""
last = ("", False)
for version in trimmed_profiles(profile):
if len(fence("user_data", version[0]).encode("utf-8")) <= limit:
return version
last = version
return last


def trimmed_profiles(profile: DatasetProfile) -> Iterator[tuple[str, bool]]:
"""The profile JSON, then ever smaller versions of it, each with whether it was trimmed.

Expand Down
10 changes: 10 additions & 0 deletions agents/anyplot/code/loader.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,16 @@ def to_run_form(working: str, *, columns: list[str], dtypes: dict[str, str], par
return _ensure_pandas(code)


def loader_arguments(
columns: list[str], dtypes: dict[str, str], parse_dates: list[str]
) -> tuple[dict[str, str], list[str]]:
"""The `dtype=` mapping and the `parse_dates=` list of the run form's `pd.read_csv`, validated.

The eval harness's static stage reads `data.csv` with exactly these arguments.
"""
return _loader_arguments(columns, dtypes, parse_dates)


def _loader_arguments(
columns: list[str], dtypes: dict[str, str], parse_dates: list[str]
) -> tuple[dict[str, str], list[str]]:
Expand Down
12 changes: 8 additions & 4 deletions agents/anyplot/pipeline.py
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@
from google.genai import types
from pydantic import ValidationError

from .briefs import profile_summary, spec_brief
from .briefs import adapter_profile, profile_summary, spec_brief
from .code.edits import MAX_NEW_LITERAL_CHARS, apply_plan, new_literal_chars
from .code.export import export_code
from .code.regions import find_regions
Expand Down Expand Up @@ -353,17 +353,21 @@ def render_adapt_request(request: AdaptRequest, view: SessionView) -> str:
the way the `plot_pipeline` result reaches the root: the readiness hints quote
catalogue lines, the feedback quotes model-written code (failed edits, validator
findings) or is model-written (reviewer lines), and the previous plan is
model-written and can copy catalogue or dataset text.
model-written and can copy catalogue or dataset text. The profile is the first
version of `trimmed_profiles` within `MAX_ADAPTER_PROFILE_BYTES` (`adapter_profile`):
50 text columns of long values make it about 28,000 characters in full (stress case
`stress-cap-profile`), more than the rest of the request.
"""
profile_text, trimmed = adapter_profile(request.profile)
lines = [
f"Library: {view.library}",
"",
DATA_PREAMBLE,
"Spec brief:",
fence("spec_text", spec_brief(view)),
fence("catalogue_code", request.code),
"Dataset profile:",
fence("user_data", request.profile.model_dump_json(exclude={"warnings"})),
"Dataset profile (shortened: fewer sample rows or top values):" if trimmed else "Dataset profile:",
fence("user_data", profile_text),
"Bindings and loader columns:",
fence("user_data", _columns_json(request.bindings, request.loader_columns)),
]
Expand Down
2 changes: 2 additions & 0 deletions agents/anyplot/render/backends/remote.py
Original file line number Diff line number Diff line change
Expand Up @@ -351,5 +351,7 @@ def _result(self, job: RenderJob, response: httpx.Response) -> RenderResult:
reason=run.reason,
measured=run.measured,
limit=run.limit,
max_rss_mb=run.max_rss_mb,
cpu_s=run.cpu_s,
)
return result
5 changes: 5 additions & 0 deletions agents/anyplot/render/contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,11 @@ class ThemeOutput:
"""The value that broke the limit (bytes, entries or MiB of `MemAvailable`, by `reason`)."""
limit: int | None = None
"""The limit `measured` broke, in the same unit."""
max_rss_mb: float | None = None
"""The harness's peak memory from its closing `HARNESS` line, when the backend reports it (`remote`).
Advisory: the code under test shares stdout and can forge the line."""
cpu_s: float | None = None
"""The harness's CPU seconds from the same line, equally advisory."""

def __post_init__(self) -> None:
self.stderr_tail = self.stderr_tail[-MAX_STDERR_CHARS:]
Expand Down
Loading
Loading