Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
40 changes: 38 additions & 2 deletions agents/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ This directory holds the anyplot agent network: the service that lets an admin p

## What is built

The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer, the private `/v1` service with its run queue, and the theme toggle. The renderer service `anyplot-renderer` (`renderer/`), which runs the adapted code in Cloud Run sandboxes, and the `remote` backend that calls it are built, with the renderer's image and Cloud Build config, but not deployed yet (status of 2026-10-10). The model-regression harness, the 120 synthetic spike-X cases, the blind two-run review gallery and the catalogue eligibility sweep are built (see [Run the regression harness](#run-the-regression-harness)), and the Claude Haiku 5.5 baseline from the first spike-X run of 2026-10-10 is committed (`evals/baselines/claude-haiku-5-5.json`); the Gemini baseline is not. Not built yet: the agents image, the deploy of either service, the scope evalset, and the `agents-eval.yml` workflow.
The runtime core runs locally: the agents, the plot pipeline, the guardrail plugins, the render layer, the private `/v1` service with its run queue, and the theme toggle. The renderer service `anyplot-renderer` (`renderer/`), which runs the adapted code in Cloud Run sandboxes, and the `remote` backend that calls it are built, with the renderer's image and Cloud Build config, but not deployed yet (status of 2026-10-10). The model-regression harness, the 120 synthetic spike-X cases, the blind two-run review gallery and the catalogue eligibility sweep are built (see [Run the regression harness](#run-the-regression-harness)), and the Claude Haiku 5.5 baseline from the first spike-X run of 2026-10-10 is committed (`evals/baselines/claude-haiku-5-5.json`); the Gemini baseline is not. The chat page that uses the service through the API's `/debug/agent` routes is in the app (`app/src/pages/AgentChatPage.tsx`, built with `VITE_ENABLE_AGENT_CHAT=true`). Not built yet: the agents image, the deploy of either service, the scope evalset, and the `agents-eval.yml` workflow.

The model is **Claude Haiku 5.5 on Vertex AI** (`claude-haiku-5-5`) by default. **Gemini 3.8 Flash** is the second arm: set `AGENT_PROVIDER=gemini` together with Gemini model ids, so the two can be compared on price and quality later. Every agent and the scope judge run on the configured provider.

Expand Down Expand Up @@ -141,10 +141,46 @@ Three rules hold for everything here:
uv run uvicorn agents.main:app --port 8001
```

The service needs the header `X-Anyplot-User` on every `/v1` route; outside `ENVIRONMENT=development` it also requires the IAM-forwarded ID token (`AGENT_SERVICE_URLS`, `AGENT_ALLOWED_CALLERS`). To drive it from the plot page, run the API with `AGENT_ENABLED=true AGENT_SERVICE_URL=http://localhost:8001` (see `api/routers/agent.py`).
The service needs the header `X-Anyplot-User` on every `/v1` route; outside `ENVIRONMENT=development` it also requires the IAM-forwarded ID token (`AGENT_SERVICE_URLS`, `AGENT_ALLOWED_CALLERS`). To drive it from the plot page, run the API in front of it as described under [Drive the chat page](#drive-the-chat-page).

At the defaults only one run may start a minute, so a second "Create plot" within a minute waits in the run queue and the stream shows `status` events with `step: "queued"`. To iterate faster on your own machine, export `AGENT_RUNS_PER_MINUTE=60`. The theme toggle (`POST /v1/sessions/{sid}/versions/{version}/render {"theme": "dark"}`) never waits in the queue.

### Drive the chat page

The chat page is the app's `/debug/agent?spec=&library=&language=`. Build or serve the app with `VITE_ENABLE_AGENT_CHAT=true`; local development shows it without an admin sign-in, and the plot page then shows the `.adapt()` button for eligible pairs.

- **Against this service:** the API's agent routes need all three of `AGENT_ENABLED`, `AGENT_SERVICE_URL` and `AGENT_USER_ID_KEY`, or every route answers `404 not_enabled`; and the admin gate has no development bypass, so you sign in with an admin token:

1. Run the service as above.
2. In a second terminal, run the API with the agent routes switched on and an admin token of your choice:

```bash
AGENT_ENABLED=true AGENT_SERVICE_URL=http://localhost:8001 AGENT_USER_ID_KEY=dev-key \
ADMIN_TOKEN=<token> uv run uvicorn api.main:app --reload --port 8000
```

3. In a third terminal, start the app with `cd app && VITE_ENABLE_AGENT_CHAT=true yarn dev`.
4. Open `http://localhost:3000/debug`, enter `<token>`, and open the chat page in the same tab: the token lives in that tab's session storage.

- **Without any backend:** the mock BFF serves the documented `/debug/agent/*` routes, a scripted stream (two queue positions, the pipeline steps, a plot and a reply) and PNGs drawn from the pasted data, plus the catalogue routes the plot page needs for `scatter-basic` and `line-multi` (a series family `y1, y2, ...`). It listens on the loopback interface only, needs no sign-in, touches no database and calls no model:

1. Start the mock:

```bash
node app/scripts/agent-bff-mock.mjs
```

2. In a second terminal, start the app against it:

```bash
cd app && VITE_ENABLE_AGENT_CHAT=true VITE_API_URL=http://localhost:8010 \
VITE_DEBUG_API_URL=http://localhost:8010 yarn dev
```

3. Open `http://localhost:3000/scatter-basic/python/matplotlib` and select the `.adapt()` button, or open `http://localhost:3000/debug/agent?spec=scatter-basic&library=matplotlib&language=python` directly. Paste `agents/evals/fixtures/cases/scatter-basic-matplotlib/data.csv` as your data.

The header of `app/scripts/agent-bff-mock.mjs` lists the scripted replies (a refusal, a capacity error, a question, a refinement with a repair round) and the timing variables.

### Render through the deployed renderer

`adk web` and the local `/v1` service can render through the deployed `anyplot-renderer`, the same sandboxes the deployed agents service uses. The renderer accepts a token only when Cloud Run IAM lets it through (`roles/run.invoker` on the service) and its claims pass the renderer's own check: the token's `aud` in `RENDERER_AUDIENCES` and its `email` in `RENDERER_ALLOWED_CALLERS`.
Expand Down
20 changes: 15 additions & 5 deletions agents/anyplot/pipeline.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,9 @@
and the reviewer passed exactly that render; `needs_attention` ships a render with
residual defect lines (reviewer lines that no reviewed render fixed, adaptation
findings, a padded canvas, or no review because the budget or the deadline ran out);
`failed` names its reason; `not_ready` comes before any model call. Progress goes out
`failed` names its reason; `not_ready` comes before any model call. An `ok` or
`needs_attention` result is stored as the session's next version before it is
yielded, and carries that number as `version`. Progress goes out
as content-free events with `custom_metadata={"anyplot_status": {"step", "attempt"}}`,
which the stream translator turns into `status` events and no model ever reads.

Expand Down Expand Up @@ -351,17 +353,24 @@ def finish(run: Run) -> PlotResult:
)


def _store_version(ctx: Context, services: Services, run: Run, result: PlotResult) -> None:
def _store_version(ctx: Context, services: Services, run: Run, result: PlotResult) -> PlotResult:
"""Store a shipped result as the session's next version; the result comes back with its `version`.

The number travels in the `plot` event, so a client addresses the artifact and theme
toggle routes by the server's number instead of counting the events it received.
"""
shipped = run.best or run.padded
if shipped is None or result.status not in ("ok", "needs_attention"):
return
return result
session_id = ctx.session.id
render_id = shipped.render_id or services.renders.put(session_id, shipped.pngs)
snapshot = run.view.snapshot
number = services.versions.next_number(session_id)
result = PlotResult.model_validate({**result.model_dump(), "version": number})
services.versions.add(
session_id,
CodeVersion(
number=services.versions.next_number(session_id),
number=number,
working=shipped.working,
run_form=shipped.run_form,
export=export_code(
Expand All @@ -381,6 +390,7 @@ def _store_version(ctx: Context, services: Services, run: Run, result: PlotResul
themes={run.theme: ThemeRender("needs_attention", PADDED_REASON) if shipped.padded else ThemeRender("ok")},
),
)
return result


async def _adapt(ctx: Context, scope: str, request_text: str, library: str) -> AdaptPlan | str:
Expand Down Expand Up @@ -455,7 +465,7 @@ async def run_pipeline(ctx: Context, node_input: PipelineArgs) -> AsyncGenerator
ledger.review_render_id = None
result = finish(run)
try:
_store_version(ctx, services, run, result)
result = _store_version(ctx, services, run, result)
except Exception as exc: # a result whose artifacts cannot be stored is not shippable
logger.warning("storing the version failed: %s", type(exc).__name__)
result = PlotResult(status="failed", reason="error", attempts=run.attempts)
Expand Down
7 changes: 7 additions & 0 deletions agents/anyplot/schemas.py
Original file line number Diff line number Diff line change
Expand Up @@ -292,6 +292,10 @@ class PlotResult(_ServerContract):
padded. `needs_attention`: a render shipped with residual defect lines.
`failed`: no render passed the blocking host gates; `reason` says why.
`not_ready`: no dataset or incomplete bindings, decided before any LLM call.

`version` is the number the session's version store gave the shipped result
(`ok` or `needs_attention`), the one the artifact and theme toggle routes take;
it is set once the version is stored, and a `failed` or `not_ready` result has none.
"""

status: PlotStatus
Expand All @@ -300,9 +304,12 @@ class PlotResult(_ServerContract):
artifacts: list[ArtifactName] = Field(default_factory=list, max_length=4)
changes: list[ChangeNote] = Field(default_factory=list, max_length=MAX_CHANGES)
residual_defects: list[Line] = Field(default_factory=list, max_length=MAX_RESIDUAL_DEFECTS)
version: int | None = Field(default=None, ge=1)

@model_validator(mode="after")
def _status_rules(self) -> Self:
if self.status in ("failed", "not_ready") and self.version is not None:
raise ValueError(f"a {self.status!r} result stores no version")
if self.status == "failed":
if self.reason not in FAILURE_REASONS:
raise ValueError(f"a failed result needs a reason in {sorted(FAILURE_REASONS)}")
Expand Down
28 changes: 26 additions & 2 deletions agents/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@
| `GET /v1/eligibility?spec=&library=` | | `{eligible, status, reasons}` | |
| `POST /v1/sessions` | `{user, spec_id, library, locale, snapshot}` | `{session_id, eligibility}` | `422 not_eligible` |
| `POST /v1/sessions/{sid}/library` | `{library, snapshot}` | `{session_id, eligibility}` | `422 not_eligible`, `409 run_active` |
| `POST /v1/sessions/{sid}/dataset` | `{text}` | `{preview, profile, bindings, warnings}` | `413 too_long`, `422 unparseable`, `403 data_refused`, `503 guard_unavailable` |
| `PUT /v1/sessions/{sid}/bindings` | `[{role, column}]` | `{bindings, complete, missing_roles}` | `409 run_active`, `422 invalid` |
| `POST /v1/sessions/{sid}/dataset` | `{text}` | `{preview, profile, bindings, warnings, roles}` | `413 too_long`, `422 unparseable`, `403 data_refused`, `503 guard_unavailable` |
| `PUT /v1/sessions/{sid}/bindings` | `[{role, column}]` | `{bindings, complete, missing_roles}` | `409 run_active`, `422 invalid` (with `errors`, at most 20 lines) |
| `POST /v1/sessions/{sid}/messages` | `{text}` or `{action}` | SSE `anyplot/1` | `413 too_long`, `409 run_active`, `503 capacity` |
| `POST /v1/sessions/{sid}/cancel` | | `204` | |
| `POST /v1/sessions/{sid}/versions/{version}/render` | `{theme}` | `{status, reason?, artifacts}` | `404 not_found`, `409 run_active`, `503 capacity` (no render slot in time, or the render store is full) |
Expand Down Expand Up @@ -43,6 +43,11 @@
in flight at a time. `adk web` runs the agents without this service, so its runs
bypass the queue; its renders still go through the one render slot.

The dataset answer lists the spec's data roles as `roles`, each
`{name, kinds, required, variadic, description}`, so a client can offer a column
choice for every role: a single role binds under its own name, a variadic family
`y` binds its members `y1`, `y2`, ... (`data/bindings.py`).

Run locally with `uv run uvicorn agents.main:app --port 8001`.
"""

Expand Down Expand Up @@ -77,6 +82,7 @@
from agents.anyplot import agent as agent_module
from agents.anyplot.code.readiness import MAP_SPECS
from agents.anyplot.data.parse import MAX_INPUT_BYTES, ParseError, parse_dataset
from agents.anyplot.data.roles import DataRole
from agents.anyplot.data.store import StoreFull
from agents.anyplot.dev_fixture import FixtureError, load_case
from agents.anyplot.models import JudgeUnavailable
Expand Down Expand Up @@ -589,11 +595,29 @@ async def upload_dataset(
"profile": parsed.profile.model_dump(mode="json"),
"bindings": [binding.model_dump() for binding in ingested.bindings],
"warnings": parsed.warnings,
"roles": [_role_body(role) for role in view.snapshot.roles()],
},
headers=_NO_STORE,
)


MAX_ROLE_DESCRIPTION_CHARS = 200


def _role_body(role: DataRole) -> dict[str, Any]:
"""One spec data role for the binding controls; the description is the spec's own text, capped."""
description = role.description
if len(description) > MAX_ROLE_DESCRIPTION_CHARS:
description = description[: MAX_ROLE_DESCRIPTION_CHARS - 1].rstrip() + "…"
return {
"name": role.name,
"kinds": list(role.kinds),
"required": role.required,
"variadic": role.variadic,
"description": description,
}


@app.put("/v1/sessions/{sid}/bindings", dependencies=v1_dependencies)
async def put_bindings(
sid: SessionId,
Expand Down
11 changes: 8 additions & 3 deletions agents/stream.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
| `status` | `step: "queued"`, `position`, `waiting` | the run queue, while the run waits: at once, on every change, and every 15 s unchanged (`position` 1 runs next; `waiting` counts every queued entry, this one included) |
| `status` | `step`, `attempt` | the pipeline's content-free `custom_metadata` progress events |
| `message` | `text` | a final, non-partial text response authored by the root (`anyplot`) |
| `plot` | `status`, `reason`, `attempts`, `artifacts`, `changes`, `residual_defects` | the pipeline's `PlotResult` output event |
| `plot` | `status`, `reason`, `attempts`, `artifacts`, `changes`, `residual_defects`, `version` | the pipeline's `PlotResult` output event; `version`, the stored version's number, only on `ok` and `needs_attention` |
| `refusal` | `code`, `text` | the request ledger's refusal (scope guard or budget), in place of the message; a user already over the daily budget gets it right after `ready`, without waiting in the queue |
| `error` | `code`, `ref` | `guard_unavailable`, `capacity` (also when the run waited the queue's maximum), `deadline` or `internal` |
| `done` | `llm_calls`, `tokens` | the end of every run, always last |
Expand Down Expand Up @@ -45,7 +45,7 @@
STEPS = frozenset({"adapting", "checking", "rendering", "reviewing", "repairing"})
QUEUED_STEP = "queued"
"""The status step the route sends while the run waits in the run queue; no pipeline event carries it."""
PLOT_FIELDS = ("status", "reason", "attempts", "artifacts", "changes", "residual_defects")
PLOT_FIELDS = ("status", "reason", "attempts", "artifacts", "changes", "residual_defects", "version")
MODEL_WRITTEN_FIELDS = ("changes", "residual_defects")
MAX_MESSAGE_CHARS = 3_000
MAX_CODE_LINES = 10
Expand Down Expand Up @@ -151,8 +151,13 @@ def translate(self, event: Event) -> list[str]:
return out

def _plot(self, output: dict[str, Any]) -> dict[str, Any]:
"""The plot event: the allowlisted fields, the model-written lines sanitised to one plain line each."""
"""The plot event: the allowlisted fields, the model-written lines sanitised to one plain line each.

`version` goes out only when the result was stored as a version.
"""
data = {key: output[key] for key in PLOT_FIELDS if key in output}
if not isinstance(data.get("version"), int):
data.pop("version", None)
for key in MODEL_WRITTEN_FIELDS:
if isinstance(data.get(key), list):
lines = (plain_line(item, spec_id=self.spec_id) for item in data[key] if isinstance(item, str))
Expand Down
Loading
Loading